Guide
How airline on-time statistics work (and how airlines game them)
Every US carrier above a revenue threshold files flight-level data with the DOT each month. The filing is honest. The incentives around it are not always.

Every month, each US carrier above the DOT's revenue threshold files a record for every single domestic flight it operated: scheduled and actual gate times, tail number, cancellation and diversion flags, and a breakdown of delay minutes by cause. Roughly 7.05 million flights per twelve months end up in that database, and it is published in full, for free, with no registration.
The data is unusually good. The behaviour it produces is more complicated.
What is in the filing
The reporting requirement covers carriers accounting for at least 0.5% of domestic scheduled passenger revenue. That captures the mainline carriers and, since a rule change, requires them to report their regional partners' flights as well, which closed a large hole. Before that, a flight sold as a mainline product but operated by a regional under a capacity purchase agreement could vanish from the mainline's numbers.
The fields are flight-level rather than aggregated. That matters: it means an outside analyst can recompute any published statistic from scratch rather than taking a carrier's word for the summary. Everything on this site is recomputed that way each month.
Publication runs roughly two months behind. Data for a given month appears in the middle of the month after next. This site says which period it covers on every page, because a delay statistic without a date attached is not a statistic.
Three ways the number gets managed
None of the following is fraud. All of it is rational behaviour by people measured on a published metric.
Schedule padding. Block times are set by the airline. Extending a scheduled flight time improves the on-time rate with no operational change at all. Across the industry, scheduled block times on many routes have grown by twenty to forty minutes over three decades on aircraft that fly no slower. Some of that is real growth in taxi time at saturated airports. Some of it is buying a better number.
Cancelling to protect the average. Cancelled flights appear in no on-time calculation. An operation that cancels early and decisively on a bad-weather morning will post a better punctuality figure than one that tries to fly everything and delivers a day of four-hour delays. Whether that is good or bad for passengers depends entirely on what happens to the rebooking, which the data does not capture.
Choosing where to compete. A carrier concentrated in Salt Lake City and Minneapolis is operating in easier airspace than one concentrated in Newark and LaGuardia. Its national on-time average will look better without any difference in operational quality. This is why comparing carriers at the same airport, or on the same route, is worth far more than comparing their headline national figures.
What the data genuinely proves
The comparisons that survive all three problems are the narrow ones.
Two carriers flying the same route in the same months face the same airspace, the same weather, and similar padding pressures. A ten-point gap there is real. Every route page on this site leads with exactly that table.
The same is true of a carrier's performance at a single airport against its competitors at that airport. That comparison controls for the hardest confounder, which is where the airline chose to build its network.
Cause breakdowns hold up too, because the categories are defined by the DOT rather than by the carrier, and a carrier that files a delay as NAS when it was a maintenance problem is misreporting rather than optimising.
What it cannot tell you
The database records what the aircraft did. It records nothing about what happened to the passenger.
It does not know whether a cancelled passenger was rebooked in two hours or two days. It does not know whether a delay was announced four hours in advance or discovered at the gate. It does not know about missed connections at all: a flight arriving 90 minutes late is one row, regardless of whether every passenger made their onward flight or none did.
It also says nothing about seat comfort, fees, or how the airline behaves when things go wrong, which for many travellers is the thing that actually determines whether they book again.
On-time data answers one question well: how often does this operation deliver the schedule it sold. That is worth knowing precisely because it is measurable, comparable, and published by law. It is not the whole picture, and this site does not pretend otherwise.
Reading a carrier's number honestly
Three habits make the figures much harder to mislead you with.
Compare at the route or airport level, not nationally. Look at cancellation rate and on-time rate in the same glance. And check the sample size: a rate built from 400 flights moves by several points on one bad week, while one built from 40,000 does not.
Every table on this site prints the flight count next to the percentage for that reason.
Why an independent recomputation matters
The DOT publishes summary tables alongside the raw records, and most reporting quotes the summaries. This site does not use them.
The reason is not distrust of the arithmetic. It is that a summary embeds someone else's choices about what to group, what to exclude and how to weight, and those choices are rarely visible in the published number. Recomputing from the flight-level records means every threshold on this site is one you can read on the methodology page and check against the source yourself.
It also means corrections propagate. BTS occasionally revises a published month; because the pipeline here recomputes the full twelve-month window each time rather than appending to a stored total, a revision lands automatically at the next refresh rather than persisting until someone notices.
Compare carriers where it counts: Southwest's on-time record · JetBlue's on-time record · Boston (BOS) delay statistics