UART baud rate error: how much you can get away with
Why 2 % per end is the limit, where 8.5 % errors come from, what the odd crystal values buy, and what a fractional baud divider costs you in return.
The short answer is about 2 % per end for an ordinary 8N1 frame, and it is not a rule of thumb — it is in the datasheet. The ATmega48A/PA/88A/PA/168A/PA/328/P datasheet (DS40002061B) publishes a receiver that tolerates a total error of +4.58 % / −4.64 % on an 8-bit frame, and then recommends ±2.0 % per end, because:
The recommendations of the maximum receiver baud rate error was made under the assumption that the Receiver and Transmitter equally divides the maximum total error.
The baud rate error calculator works out what your clock and divider actually deliver. What follows is why the limit is where it is, why some clock frequencies are exact and others are hopeless, and what the fractional divider on a modern part costs in exchange for fixing that.
The error accumulates, which is the whole problem
An asynchronous link has no clock line. The receiver resynchronises exactly once per character, on the falling edge of the start bit, and then runs open-loop for the rest of the frame. If its bit period is wrong by a fraction , its sample point for bit lands late by
and that drift never recovers. The first data bit is sampled a twentieth of a bit off at 5 % error; the stop bit is sampled nearly half a bit off. The naive limit is where the last sample falls out of its own bit:
for a 10-bit frame. Two things then make the real limit tighter: the receiver takes several samples and votes, so it needs a window rather than an instant, and the start edge itself is only located to within one sample period.
What the receiver actually does
The ATmega samples at 16× the bit rate and takes samples 8, 9 and 10 of each bit, deciding by majority vote. Its published tolerance follows from that directly, and the datasheet gives the two expressions:
where is the number of data plus parity bits, is samples per bit (16 normal, 8 double speed), is the first majority-vote sample and the middle one. For a plain 8N1 frame, , , , :
R_slow = (8+1)(16) / (16 - 1 + 8(16) + 8) = 144 / 151 = 95.36 %
R_fast = (8+2)(16) / ((8+1)(16) + 9) = 160 / 153 = 104.58 %
which are the 95.36 % and 104.58 % the datasheet’s table 20-2 prints for D = 8. One cell of that table disagrees with itself: its “max total error” column gives −4.54 % where 100 − 95.36 is −4.64 %. Every other row of the table is self-consistent, so that entry is a slip in the datasheet, and −4.64 % is the figure used here. Evaluating the two expressions across every frame size gives back the whole of tables 20-2 and 20-3, which is worth doing once: it is the cheapest available check that the formulas have been transcribed correctly, and it is what catches that slip.
Two properties fall out of the plot. Longer frames tolerate less, because the drift has more bits to accumulate over — a 10-bit payload gets ±1.5 % recommended where a 5-bit one gets ±3.0 %. And double-speed mode tolerates less at every frame size, because halving the oversampling halves the resolution with which the receiver can place its samples. Doubling the maximum rate is not free.
Where 8.5 % and −18.6 % come from
The other half of the error is the divider, and it is pure arithmetic. An integer baud generator produces
so the achievable rates are a harmonic series, and the error is whatever rounding the nearest whole divisor costs. At 1 MHz aiming for 38 400 baud the ideal divisor is 1.63; round it to 2 and the port runs at 31 250 baud, which is 18.6 % low. Nothing is broken and no setting fixes it — that rate does not exist on that clock.
Recomputing the datasheet’s own table 20-4 from the equation reproduces every published value, including the ones that look like typos:
1.0000 MHz 1.8432 MHz 2.0000 MHz
9 600 UBRR 6 −7.0% UBRR 11 0.0% UBRR 12 +0.2%
19 200 UBRR 2 +8.5% UBRR 5 0.0% UBRR 6 −7.0%
38 400 UBRR 1 −18.6% UBRR 2 0.0% UBRR 2 +8.5%
57 600 UBRR 0 +8.5% UBRR 1 0.0% UBRR 1 +8.5%
115 200 — UBRR 0 0.0% UBRR 0 +8.5%
The middle column is exact at every rate. That is the entire reason the value exists.
Why 1.8432, 11.0592 and 14.7456 MHz
They are the clocks for which the divider comes out whole. At 16× oversampling, 115 200 baud needs to be an integer, so the exact clocks are the multiples of 1.8432 MHz: 3.6864, 5.5296, 7.3728, 9.216, 11.0592, 12.9024, 14.7456, 18.432. Every one of those is a stock crystal value, and they look arbitrary only if you have not divided them by 1.8432.
Because every standard rate below 115 200 divides it without remainder — 57 600, 38 400, 19 200, 9 600, 4 800 and 2 400 are 115 200 over 2, 3, 6, 12, 24 and 48, and 28 800 and 14 400 are 115 200 over 4 and 8 — a clock exact at 115 200 is exact at all of them. That is why the 1.8432 MHz column has no error anywhere. The one common rate that is not on the list is 76 800, which is 115 200 × 2/3: at 1.8432 MHz and 16× oversampling it needs a divisor of 1.5, and the datasheet’s table prints −25 % for it.
The common round-numbered clocks are not on that list. At 115 200 baud a 16 MHz part is 3.5 % off and a 12 MHz part is 7.0 % off — both past the per-end recommendation, and the 12 MHz case past the total the receiver can absorb even with a perfect partner. This is why so many AVR boards run from an awkward-looking crystal, and why so many that do not have a serial port that works at 9 600 and fails at 115 200.
The fractional divider, and what it costs
Modern parts solve this with resolution rather than crystal selection. ST’s RM0008 (STM32F10x reference manual) describes a divider held as a fixed-point number:
USARTDIV is an unsigned fixed point number that is coded on the USART_BRR register.
with a 12-bit mantissa and a 4-bit fraction, so the divisor moves in steps of 1/16 rather than 1. The effect on the arithmetic is dramatic: at 12 MHz and 115 200 baud, an integer divisor gives −7.0 % and a 1/16 divisor gives +0.16 %. Across 8 to 20 MHz the worst case at that rate falls from 12.5 % to 0.72 %.
Then the part almost nobody quotes. RM0008 publishes two receiver tolerance tables, and which one applies depends on whether the fraction is in use:
| Frame | Noise flag | DIV_Fraction = 0 | DIV_Fraction ≠ 0 |
|---|---|---|---|
| M = 0 (10-bit) | NF is an error | 3.75 % | 3.33 % |
| M = 0 (10-bit) | NF is don’t care | 4.375 % | 3.88 % |
| M = 1 (11-bit) | NF is an error | 3.41 % | 3.03 % |
| M = 1 (11-bit) | NF is don’t care | 3.97 % | 3.53 % |
Using a non-zero fraction costs roughly half a point of receiver tolerance in every case, because the sample clock is now alternating between two divider counts rather than being uniform. The trade is still overwhelmingly worth taking — half a point of tolerance against several points of divider error — but it means a fractional divider is not simply free precision, and on a link that is already marginal it moves the wrong way.
The budget has four terms, and the divider is the smallest
RM0008 sets the constraint out as a sum, which is the right way to think about it:
DTRA: Deviation due to the transmitter error (which also includes the deviation of the transmitter local oscillator)
DQUANT: Error due to the baud rate quantization of the receiver
DREC: Deviation of the receiver’s local oscillator
DTCL: Deviation due to the transmission line (generally due to the transceivers that can introduce an asymmetry between the low-to-high transition timing and the high-to-low transition timing)
DTRA + DQUANT + DREC + DTCL < USART receiver tolerance
Written out like that, the usual failure becomes obvious. The quantisation term — the only one most people compute — is a fraction of a per cent on any modern part. The oscillator terms are where the budget goes, and the ATmega datasheet says so plainly:
The Receiver’s system clock (XTAL) will always have some minor instability over the supply voltage range and the temperature range. When using a crystal to generate the system clock, this is rarely a problem, but for a resonator the system clock may differ more than 2% depending of the resonators tolerance.
A ±20 ppm crystal spends 0.002 % of a 3.75 % budget. A ceramic resonator can spend all of it on its own. Two devices each running from an untrimmed internal RC oscillator cannot reliably talk to each other at all, which is the real reason a board with a serial bootloader has a crystal on it — see the crystal ppm calculator for how the tolerance, the load mismatch and the temperature drift stack up, and why the crystal won’t start for what happens when it is fitted badly.
The fourth term, DTCL, is the one that catches differential buses. A transceiver whose low-to-high and high-to-low propagation delays differ shortens some bits and lengthens others, spending budget that has nothing to do with either clock — which is part of why an RS-485 link has its own termination and biasing rules.
Why 2 % per end and not 4 %
The receiver’s constraint is on the sum of the two ends’ errors, but neither end knows the other’s. Splitting the budget equally is what makes a device interoperable without negotiation: any two devices each inside ±2 % are guaranteed to be inside the ±4.58 % total, whatever they are.
The consequence is asymmetric in a useful way. A device with a crystal contributes essentially nothing, so it can talk to a partner using nearly the whole budget — which is why a USB-serial adapter with a crystal will happily receive from a sloppy microcontroller that cannot receive from anything. Two sloppy devices, each individually “within spec” at 3 %, fail against each other while working perfectly against a good one. That failure pattern is diagnostic: if a link works against a PC and not against another board, suspect the sum, not either end.
The ceiling, and the trap next to it
The maximum rate is the divisor at its minimum: , or in double-speed mode. A 16 MHz AVR tops out at 1 Mbaud normally and 2 Mbaud doubled, which the datasheet’s own table confirms.
The trap is that the divisor is also the resolution. With a divisor of 1 the next available rate is half the current one; with a divisor of 2 it is a third lower. Near the ceiling there is nothing to round to, so the errors are largest exactly where the rate is highest — which is the mechanism behind the +8.5 % entries in the table above, every one of which occurs at a divisor of one, two or three.
The practical rule is to keep the divisor above about 16, which caps the rounding error at roughly 3 % in the worst case and well under 1 % typically. Below that, either change the clock or use a part with a fractional divider.
The checklist
- Compute the error for both ends, not one. The constraint is on the sum.
- Aim for under 2 % per end on an 8N1 frame, under 1.5 % if you are using parity or 9 data bits.
- Check the divisor, not just the error. A divisor under about 16 means the next available rate is far away and temperature drift has nowhere to go.
- Use a crystal if either end has to interoperate with something you did not build. A resonator spends the whole per-end budget by itself.
- If the clock is fixed and awkward, a part with a fractional divider will fix the arithmetic — at the cost of about half a point of receiver tolerance.
- If the link works against a PC and not against your own second board, the two errors are adding. Measure the actual bit time on a scope rather than trusting either datasheet.
Work the numbers with the baud rate error calculator, which takes the clock, the rate, the oversampling factor and the number of fraction bits, and reports the divisor, the realised rate and the error — including whether the divisor will fit the register at all.