100nF

UART baud rate error: how much you can get away with

Why 2 % per end is the limit, where 8.5 % errors come from, what the odd crystal values buy, and what a fractional baud divider costs you in return.

The short answer is about 2 % per end for an ordinary 8N1 frame, and it is not a rule of thumb — it is in the datasheet. The ATmega48A/PA/88A/PA/168A/PA/328/P datasheet (DS40002061B) publishes a receiver that tolerates a total error of +4.58 % / −4.64 % on an 8-bit frame, and then recommends ±2.0 % per end, because:

The recommendations of the maximum receiver baud rate error was made under the assumption that the Receiver and Transmitter equally divides the maximum total error.

The baud rate error calculator works out what your clock and divider actually deliver. What follows is why the limit is where it is, why some clock frequencies are exact and others are hopeless, and what the fractional divider on a modern part costs in exchange for fixing that.

One asynchronous character drawn as a start bit, eight data bits and a stop bit, with sixteen clock ticks inside each bit and three of them — samples eight, nine and ten — marked as the majority vote that decides the bit.
Fig 1 — What the receiver actually does. It has no clock from the transmitter, so it starts a counter on the falling edge of the start bit and samples each bit near its middle. The ATmega takes samples 8, 9 and 10 of 16 and votes; everything about baud error comes down to whether those three ticks still land inside the right bit ten bits later.

The error accumulates, which is the whole problem

An asynchronous link has no clock line. The receiver resynchronises exactly once per character, on the falling edge of the start bit, and then runs open-loop for the rest of the frame. If its bit period is wrong by a fraction ee, its sample point for bit kk lands late by

Δtk=(k+12)ebit-times\Delta t_k = \left(k + \tfrac{1}{2}\right) e \quad \text{bit-times}

and that drift never recovers. The first data bit is sampled a twentieth of a bit off at 5 % error; the stop bit is sampled nearly half a bit off. The naive limit is where the last sample falls out of its own bit:

∣e∣<0.5k+12=0.59.5=5.26 %|e| < \frac{0.5}{k + \tfrac{1}{2}} = \frac{0.5}{9.5} = 5.26\ \%

for a 10-bit frame. Two things then make the real limit tighter: the receiver takes several samples and votes, so it needs a window rather than an instant, and the start edge itself is only located to within one sample period.

Two rows of sample instants drawn against the transmitter bit boundaries. At two per cent error every sample stays near the middle of its own bit; at six per cent the samples creep steadily right until the last one falls into the following bit.
Fig 2 — Why the limit is a few per cent and not a few tenths. The error does not average out: it accumulates from the start edge, so the sample for bit k lands (k + ½)·e bit-times late. The first bits are unaffected and the last one decides everything, which is why longer frames tolerate less error.

What the receiver actually does

The ATmega samples at 16× the bit rate and takes samples 8, 9 and 10 of each bit, deciding by majority vote. Its published tolerance follows from that directly, and the datasheet gives the two expressions:

Rslow=(D+1)SS−1+D⋅S+SF,Rfast=(D+2)S(D+1)S+SMR_{slow} = \frac{(D+1)S}{S - 1 + D\cdot S + S_F}, \qquad R_{fast} = \frac{(D+2)S}{(D+1)S + S_M}

where DD is the number of data plus parity bits, SS is samples per bit (16 normal, 8 double speed), SFS_F is the first majority-vote sample and SMS_M the middle one. For a plain 8N1 frame, D=8D = 8, S=16S = 16, SF=8S_F = 8, SM=9S_M = 9:

R_slow  = (8+1)(16) / (16 - 1 + 8(16) + 8)   = 144 / 151 = 95.36 %
R_fast  = (8+2)(16) / ((8+1)(16) + 9)        = 160 / 153 = 104.58 %

which are the 95.36 % and 104.58 % the datasheet’s table 20-2 prints for D = 8. One cell of that table disagrees with itself: its “max total error” column gives −4.54 % where 100 − 95.36 is −4.64 %. Every other row of the table is self-consistent, so that entry is a slip in the datasheet, and −4.64 % is the figure used here. Evaluating the two expressions across every frame size gives back the whole of tables 20-2 and 20-3, which is worth doing once: it is the cheapest available check that the formulas have been transcribed correctly, and it is what catches that slip.

Two properties fall out of the plot. Longer frames tolerate less, because the drift has more bits to accumulate over — a 10-bit payload gets ±1.5 % recommended where a 5-bit one gets ±3.0 %. And double-speed mode tolerates less at every frame size, because halving the oversampling halves the resolution with which the receiver can place its samples. Doubling the maximum rate is not free.

Maximum tolerable baud-rate error plotted against the number of data and parity bits in a frame, for both the sixteen-times and eight-times oversampling modes. Tolerance falls as frames get longer, and the eight-times mode tolerates less at every frame size.
Fig 3 — The ATmega's own Rslow and Rfast equations evaluated across the frame sizes they cover. Longer frames tolerate less because the drift has more bits to accumulate over, and double-speed mode tolerates less at every size because halving the oversampling halves the resolution of the sample position.

Where 8.5 % and −18.6 % come from

The other half of the error is the divider, and it is pure arithmetic. An integer baud generator produces

BAUD=fosc16 (UBRR+1)\text{BAUD} = \frac{f_{osc}}{16\,(\text{UBRR}+1)}

so the achievable rates are a harmonic series, and the error is whatever rounding the nearest whole divisor costs. At 1 MHz aiming for 38 400 baud the ideal divisor is 1.63; round it to 2 and the port runs at 31 250 baud, which is 18.6 % low. Nothing is broken and no setting fixes it — that rate does not exist on that clock.

Recomputing the datasheet’s own table 20-4 from the equation reproduces every published value, including the ones that look like typos:

            1.0000 MHz        1.8432 MHz        2.0000 MHz
   9 600    UBRR 6   −7.0%    UBRR 11   0.0%    UBRR 12  +0.2%
  19 200    UBRR 2   +8.5%    UBRR  5   0.0%    UBRR  6  −7.0%
  38 400    UBRR 1  −18.6%    UBRR  2   0.0%    UBRR  2  +8.5%
  57 600    UBRR 0   +8.5%    UBRR  1   0.0%    UBRR  1  +8.5%
 115 200         —            UBRR  0   0.0%    UBRR  0  +8.5%

The middle column is exact at every rate. That is the entire reason the value exists.

A table of baud-rate errors for three system clock frequencies at several standard baud rates, recomputed from the integer divider equation. The 1.8432 megahertz column is exact everywhere while the one and two megahertz columns reach errors of eight and eighteen per cent.
Fig 4 — The ATmega datasheet's table 20-4, recomputed from BAUD = fosc/(16(UBRR+1)) with the divisor rounded to the nearest integer. Every value matches the published table, including the −18.6 % that makes 38 400 baud impossible from a 1 MHz clock, and the dash where 115 200 is above the 62.5 kbaud ceiling of that clock. The middle column is exact at every rate, and that is not a coincidence.

Why 1.8432, 11.0592 and 14.7456 MHz

They are the clocks for which the divider comes out whole. At 16× oversampling, 115 200 baud needs fosc/1 843 200f_{osc}/1\,843\,200 to be an integer, so the exact clocks are the multiples of 1.8432 MHz: 3.6864, 5.5296, 7.3728, 9.216, 11.0592, 12.9024, 14.7456, 18.432. Every one of those is a stock crystal value, and they look arbitrary only if you have not divided them by 1.8432.

Because every standard rate below 115 200 divides it without remainder — 57 600, 38 400, 19 200, 9 600, 4 800 and 2 400 are 115 200 over 2, 3, 6, 12, 24 and 48, and 28 800 and 14 400 are 115 200 over 4 and 8 — a clock exact at 115 200 is exact at all of them. That is why the 1.8432 MHz column has no error anywhere. The one common rate that is not on the list is 76 800, which is 115 200 × 2/3: at 1.8432 MHz and 16× oversampling it needs a divisor of 1.5, and the datasheet’s table prints −25 % for it.

The common round-numbered clocks are not on that list. At 115 200 baud a 16 MHz part is 3.5 % off and a 12 MHz part is 7.0 % off — both past the per-end recommendation, and the 12 MHz case past the total the receiver can absorb even with a perfect partner. This is why so many AVR boards run from an awkward-looking crystal, and why so many that do not have a serial port that works at 9 600 and fails at 115 200.

Baud-rate error at 115200 baud plotted against system clock frequency, showing deep notches to zero at multiples of 1.8432 megahertz and large errors between them.
Fig 5 — Where the odd-looking crystal values come from. At 115 200 baud with 16× oversampling the divider needs fosc/1 843 200 to be a whole number, so the exact clocks are the multiples of 1.8432 MHz: 3.6864, 7.3728, 11.0592, 14.7456, 18.4320. A 16 MHz part sits between two of them and pays 3.5 %; a 12 MHz part pays 7.0 %.
A grid of system clock frequencies against standard baud rates, each cell shaded by the size of the resulting baud-rate error with an integer divider. The rows for the odd crystal values are entirely clean while the round-numbered clocks are patchy.
Fig 6 — Every combination of a common crystal and a standard rate, with an integer divider. The failures are not at high rates in general — they are wherever the clock is not a multiple of 16 times that rate, which is why 12 MHz manages 9 600 perfectly and 115 200 at 7 %.

The fractional divider, and what it costs

Modern parts solve this with resolution rather than crystal selection. ST’s RM0008 (STM32F10x reference manual) describes a divider held as a fixed-point number:

USARTDIV is an unsigned fixed point number that is coded on the USART_BRR register.

with a 12-bit mantissa and a 4-bit fraction, so the divisor moves in steps of 1/16 rather than 1. The effect on the arithmetic is dramatic: at 12 MHz and 115 200 baud, an integer divisor gives −7.0 % and a 1/16 divisor gives +0.16 %. Across 8 to 20 MHz the worst case at that rate falls from 12.5 % to 0.72 %.

Baud-rate error against system clock frequency for one baud rate, drawn once for an integer divider and once for a divider with a four-bit fraction. The fractional version stays inside a narrow band everywhere the integer version has large excursions.
Fig 7 — What the fractional divider is for. RM0008's USARTDIV is a fixed-point number with a 4-bit fraction, so the divider resolution is 1/16 rather than 1. Across 8 to 20 MHz that takes the worst case at 115 200 baud from 12.5 % down to 0.72 %, and it is why an STM32 hits standard rates from a 72 MHz clock at all.

Then the part almost nobody quotes. RM0008 publishes two receiver tolerance tables, and which one applies depends on whether the fraction is in use:

FrameNoise flagDIV_Fraction = 0DIV_Fraction ≠ 0
M = 0 (10-bit)NF is an error3.75 %3.33 %
M = 0 (10-bit)NF is don’t care4.375 %3.88 %
M = 1 (11-bit)NF is an error3.41 %3.03 %
M = 1 (11-bit)NF is don’t care3.97 %3.53 %

Using a non-zero fraction costs roughly half a point of receiver tolerance in every case, because the sample clock is now alternating between two divider counts rather than being uniform. The trade is still overwhelmingly worth taking — half a point of tolerance against several points of divider error — but it means a fractional divider is not simply free precision, and on a link that is already marginal it moves the wrong way.

Bars comparing the receiver tolerance an STM32 USART has with an exact integer divider against the tolerance it has when the fractional part is non-zero, for both frame lengths and both noise-flag settings. Every fractional case is smaller.
Fig 8 — The part of the fractional divider that its own reference manual states and almost nobody quotes. RM0008 publishes two tolerance tables, and using a non-zero fraction costs about 0.4 to 0.5 points of receiver tolerance in every case — because the divider is now producing a sample clock that is itself jittering between two counts.

The budget has four terms, and the divider is the smallest

RM0008 sets the constraint out as a sum, which is the right way to think about it:

DTRA: Deviation due to the transmitter error (which also includes the deviation of the transmitter local oscillator)

DQUANT: Error due to the baud rate quantization of the receiver

DREC: Deviation of the receiver’s local oscillator

DTCL: Deviation due to the transmission line (generally due to the transceivers that can introduce an asymmetry between the low-to-high transition timing and the high-to-low transition timing)

DTRA + DQUANT + DREC + DTCL < USART receiver tolerance

Written out like that, the usual failure becomes obvious. The quantisation term — the only one most people compute — is a fraction of a per cent on any modern part. The oscillator terms are where the budget goes, and the ATmega datasheet says so plainly:

The Receiver’s system clock (XTAL) will always have some minor instability over the supply voltage range and the temperature range. When using a crystal to generate the system clock, this is rarely a problem, but for a resonator the system clock may differ more than 2% depending of the resonators tolerance.

A ±20 ppm crystal spends 0.002 % of a 3.75 % budget. A ceramic resonator can spend all of it on its own. Two devices each running from an untrimmed internal RC oscillator cannot reliably talk to each other at all, which is the real reason a board with a serial bootloader has a crystal on it — see the crystal ppm calculator for how the tolerance, the load mismatch and the temperature drift stack up, and why the crystal won’t start for what happens when it is fitted badly.

The fourth term, DTCL, is the one that catches differential buses. A transceiver whose low-to-high and high-to-low propagation delays differ shortens some bits and lengthens others, spending budget that has nothing to do with either clock — which is part of why an RS-485 link has its own termination and biasing rules.

Four stacked bars showing how the four contributors to total clock deviation add up in four illustrative scenarios, against the receiver tolerance line. Crystals at both ends leave large margin, a resonator at one end spends most of it, resonators at both ends exceed it, and two factory-calibrated internal RC oscillators are off the scale.
Fig 9 — RM0008 lists four contributors and requires their sum to stay under the tolerance: transmitter error, receiver quantisation, receiver oscillator, and the transmission line's own asymmetry. The bar lengths are illustrative assumptions — a 50 ppm crystal, the 2 % the ATmega datasheet warns a resonator can exceed, its own ±10 % factory RC calibration — only the four terms, the sum rule and the 3.75 % line are RM0008's. Written out this way the usual failure is obvious: it is almost never the divider.
A scale of clock source accuracies from tens of parts per million for a crystal, through a ceramic resonator that can exceed two per cent, up to the ten per cent factory calibration of an internal RC oscillator, with the two-per-cent per-end recommendation marked.
Fig 10 — Where clock error comes from, on the scale the UART cares about. The datasheet's own warning is that "for a resonator the system clock may differ more than 2 % depending of the resonators tolerance" — which by itself spends the entire per-end budget for an 8-bit frame.

Why 2 % per end and not 4 %

The receiver’s constraint is on the sum of the two ends’ errors, but neither end knows the other’s. Splitting the budget equally is what makes a device interoperable without negotiation: any two devices each inside ±2 % are guaranteed to be inside the ±4.58 % total, whatever they are.

The consequence is asymmetric in a useful way. A device with a crystal contributes essentially nothing, so it can talk to a partner using nearly the whole budget — which is why a USB-serial adapter with a crystal will happily receive from a sloppy microcontroller that cannot receive from anything. Two sloppy devices, each individually “within spec” at 3 %, fail against each other while working perfectly against a good one. That failure pattern is diagnostic: if a link works against a PC and not against another board, suspect the sum, not either end.

A plane of transmitter error against receiver error, with a diagonal band marking the combinations whose sum stays inside the receiver tolerance, and a square marking the region both ends satisfy if each keeps to the recommended per-end limit.
Fig 11 — Why the datasheet recommends 2 % per end when the total it can absorb is 4.58 %. The constraint is on the sum, so the recommendation is that square: the largest region in which neither end has to know anything about the other. Two devices each inside it always work; two devices each at 3 % might not.

The ceiling, and the trap next to it

The maximum rate is the divisor at its minimum: fosc/16f_{osc}/16, or fosc/8f_{osc}/8 in double-speed mode. A 16 MHz AVR tops out at 1 Mbaud normally and 2 Mbaud doubled, which the datasheet’s own table confirms.

The trap is that the divisor is also the resolution. With a divisor of 1 the next available rate is half the current one; with a divisor of 2 it is a third lower. Near the ceiling there is nothing to round to, so the errors are largest exactly where the rate is highest — which is the mechanism behind the +8.5 % entries in the table above, every one of which occurs at a divisor of one, two or three.

The practical rule is to keep the divisor above about 16, which caps the rounding error at roughly 3 % in the worst case and well under 1 % typically. Below that, either change the clock or use a part with a fractional divider.

Maximum achievable baud rate plotted against system clock for the sixteen-times and eight-times oversampling modes, both straight lines, with common standard baud rates drawn across them.
Fig 12 — The ceiling, and the reason double-speed mode exists. With a divisor of one the rate is fosc/16, or fosc/8 in double-speed mode — the ATmega datasheet's own maximum row. But the divisor is also the resolution: at a divisor of 1 the next available rate is half, so rates near the ceiling are the ones with no neighbours to round to.

The checklist

  • Compute the error for both ends, not one. The constraint is on the sum.
  • Aim for under 2 % per end on an 8N1 frame, under 1.5 % if you are using parity or 9 data bits.
  • Check the divisor, not just the error. A divisor under about 16 means the next available rate is far away and temperature drift has nowhere to go.
  • Use a crystal if either end has to interoperate with something you did not build. A resonator spends the whole per-end budget by itself.
  • If the clock is fixed and awkward, a part with a fractional divider will fix the arithmetic — at the cost of about half a point of receiver tolerance.
  • If the link works against a PC and not against your own second board, the two errors are adding. Measure the actual bit time on a scope rather than trusting either datasheet.

Work the numbers with the baud rate error calculator, which takes the clock, the rate, the oversampling factor and the number of fraction bits, and reports the divisor, the realised rate and the error — including whether the divisor will fit the register at all.