How often buses actually report, and what our poll interval can do about it
Status: measured. Evening baseline taken 2026-08-18; peak re-run scheduled for 2026-08-19 08:00 BST.
The question behind this: our poller asks BODS for vehicle positions every ten seconds — until 2026-08-18 it was fifteen — and buses report on their own unsynchronised clocks. So does polling faster get fresher data, or is the limit somewhere we don't control?
It is somewhere we don't control. Buses in service report every ~19 seconds on average, so at a ten-second poll we already sample about twice as fast as the data changes. Going to five seconds would collect the same fixes about 2.5 seconds sooner — latency, not information — for double the requests and double the write volume.
Where the staleness actually comes from
Three delays stack between a bus moving and a rider seeing it move:
| typical | who controls it | |
|---|---|---|
| Bus reports its position | ~19s between fixes | the operator's telematics |
| BODS ingests and republishes | seconds | DfT |
| Our poll picks it up | ≤10s, mean 5s | us |
Measured at ingest — the gap between a fix's own timestamp and our write of it — the median is 19.9s and the p90 is 1,037s (17 minutes). Only 807 of 3,812 vehicles were carrying a fix fresher than one poll interval. The long tail is BODS's own stale tail, not our lag: about 14% of the feed at any moment is vehicles whose last fix is old (see the SIRI feed inventory).
Our contribution to that median is the poll wait, which averages half the interval: 5 seconds at a 10s poll, 7.5s at 15s.
The measurement, and the trap in it
There is no history to read. vehicle_positions holds one row per bus,
overwritten every poll, so the gaps between one bus's fixes exist only as they
happen. frontend/operator_cadence.py watches the table for a window and
records them. It adds no load on BODS — it reads what the poller already
fetched — and the gaps come from the feed's own timestamps, so they are exact
whenever they are observed at all.
Two figures per operator, because either alone misleads:
- mean gap = observed vehicle-time ÷ fixes seen. Unbiased: a bus that goes quiet contributes its whole silence to the denominator.
- median gap = the middle observed gap. Biased low and cannot be otherwise, since a gap longer than the window is never seen. Read it as a floor.
The trap is the population, and it is easy to fall into. Diffing consecutive BODS fetches and counting changed timestamps gives about 6% of vehicles per 5 seconds, which implies a 75–85 second cadence. That is wrong. At night the feed is mostly husks — vehicles whose last fix is hours old and which will not change again before morning. Measured across the raw feed, the median fix age was 10,640 seconds. Those entries sit in the denominator and halve the apparent rate.
Three ways of asking the same minute of feed data, 2026-08-18:
| method | population | answer |
|---|---|---|
| % changed over a 5s step | whole feed | 5.9% → implies ~85s |
| % changed over a 20s step | whole feed | 17.7% → implies ~113s |
| exposure ÷ fixes | whole feed | 84s |
| exposure ÷ fixes | fix newer than 20 min | 19s |
Only the last is a cadence. The others are a measure of how many buses are parked. Any future claim about reporting frequency has to say which population it measured.
Per-operator, evening baseline
12 minutes from about 22:05 BST on Tuesday 2026-08-18, whole polled area (northern England), 1,727 buses across 38 operators. Slowest first.
| operator | buses | mean gap | median | silent | name |
|---|---|---|---|---|---|
| BPTR | 19 | 32s | 20s | 10.5% | The Burnley Bus Company |
| GAWY | 7 | 31s | 20s | 14.3% | East Yorkshire |
| DAGC | 6 | 26s | 24s | 16.7% | D & G Bus |
| EYMS | 26 | 25s | 23s | 11.5% | East Yorkshire |
| KDTR | 13 | 25s | 21s | 7.7% | The Keighley Bus Company |
| TMTL | 7 | 25s | 16s | 14.3% | TM Travel |
| ANWE | 17 | 25s | 21s | 11.8% | Arriva North West |
| HRGT | 12 | 25s | 22s | 16.7% | The Harrogate Bus Company |
| AMSY | 159 | 24s | 19s | 5.0% | Arriva North West |
| ANEA | 54 | 24s | 22s | 5.6% | Arriva North East |
| TLCT | 5 | 24s | 28s | 0.0% | TLC Travel |
| TPEN | 9 | 24s | 16s | 11.1% | Team Pennine |
| ACYM | 8 | 23s | 24s | 0.0% | Arriva Wales |
| FYOR | 23 | 23s | 23s | 4.3% | First York |
| PBLT | 16 | 23s | 22s | 0.0% | Preston Bus |
| SCCU | 82 | 23s | 20s | 6.1% | Stagecoach Cumbria and North Lancashire |
| GNEL | 143 | 23s | 21s | 5.6% | Go North East |
| SCMY | 57 | 23s | 20s | 7.0% | Stagecoach Merseyside and South Lancashire |
| ANUM | 39 | 23s | 22s | 5.1% | Arriva North East |
| LNUD | 20 | 23s | 17s | 5.0% | The Blackburn Bus Company |
| SCEM | 40 | 22s | 20s | 7.5% | Stagecoach East Midlands |
| FBRA | 30 | 22s | 15s | 13.3% | First Bradford |
| WRAY | 56 | 22s | 18s | 3.6% | Arriva Yorkshire |
| SCNE | 110 | 21s | 20s | 0.0% | Stagecoach North East |
| FHUD | 20 | 21s | 18s | 5.0% | First Halifax, Calder Valley & Huddersfield |
| SYRK | 65 | 20s | 20s | 1.5% | Stagecoach Yorkshire |
| FLDS | 98 | 20s | 17s | 5.1% | First Leeds |
| FHAL | 18 | 20s | 16s | 5.6% | First Halifax |
| FSYO | 69 | 20s | 17s | 4.3% | First South Yorkshire |
| WBTR | 16 | 18s | 19s | 6.2% | Warrington's Own Buses |
| FLYE | 6 | 18s | 15s | 0.0% | Flyer |
| HUYT | 10 | 16s | 17s | 0.0% | Huyton Travel |
| BLAC | 27 | 15s | 15s | 3.7% | Blackpool Transport |
| BNSM | 150 | 15s | 12s | 9.3% | Bee Network |
| BNML | 150 | 14s | 12s | 12.0% | Bee Network |
| BNGN | 101 | 13s | 12s | 6.9% | Bee Network |
| BNDB | 27 | 13s | 12s | 0.0% | Bee Network |
| BNFM | 12 | 12s | 12s | 0.0% | Bee Network |
Fleet-weighted mean gap 19s. A further 12 operators had fewer than 5 buses out and are not listed; 338 buses carried no resolved journey, so no operator could be attributed to them.
The South Yorkshire finding is a negative one, and worth stating plainly: no South Yorkshire operator is a slow-reporting outlier. First South Yorkshire and Stagecoach Yorkshire both sit at 20s, at or slightly better than the fleet-weighted mean, with low silent shares. TM Travel's 25s is on 7 buses, which is too thin to put in front of anyone.
The comparison that does carry an argument is Bee Network's 12–13s across all five of its NOCs: Greater Manchester's buses report roughly twice as often as ours, on the same national feed. That is what an authority specifying the telematics looks like from the outside.
What follows for the poll interval
AVL_POLL_SECONDS=10 since 2026-08-18 17:27. Fifteen was never a decision —
it was the default in avl_poller.py. The change bought about 2.5 seconds of
median freshness.
Halving again, to five seconds:
| 15s | 10s (now) | 5s | |
|---|---|---|---|
| Polls per day | 5,760 | 8,640 | 17,280 |
| Row updates per day | ~44M | ~66M | ~133M |
| Mean poll wait added | 7.5s | 5.0s | 2.5s |
| Fixes captured | ~all | ~all | ~all |
The last row is the argument. At 10s we sample twice as fast as buses report, so we are already catching essentially every fix; five seconds collects the same fixes sooner rather than collecting more of them. Not worth doing.
The same 19s cadence bounds departure-time interpolation at timing points: a bus at 30 km/h covers about 165 metres between fixes, and no poll rate of ours improves on that. If we want genuinely fresher data, the lever is upstream — which operators' equipment reports slowly — and the table above is the list to take to them.
Re-running it
cd /opt/journey-planner/frontend
venv/bin/python operator_cadence.py --minutes 30 # whole polled area
venv/bin/python operator_cadence.py --minutes 30 --box 370 # South Yorkshire
--box takes a named area (370 South Yorkshire, 450 West Yorkshire, 320
East Riding & Hull) or an explicit minLon,minLat,maxLon,maxLat ordered as
AVL_BOUNDING_BOX is. It is a rough rectangle on purpose: this measures
telematics cadence, not which authority a journey serves.
A peak run is scheduled for 2026-08-19 09:00 CEST / 08:00 BST via a
transient systemd timer (cadence-peak.timer, running
/usr/local/bin/cadence-peak-run), taking both views in parallel for 30
minutes into data/cadence-peak-{all,sy}-2026-08-19.txt. The timer is
transient and will not survive a reboot of the host.
What this does not tell you
- One evening. Several operators had fewer than ten buses out. Nobody's number should be quoted at them before the peak run.
- Only buses that report at all. The population is vehicles whose last fix is under twenty minutes old. An operator whose equipment reports every half hour is excluded entirely rather than appearing as slow; the "silent" column — buses that reported nothing new for the whole window — only partly catches this.
- Only resolved buses carry an operator. 338 vehicles in the baseline had no journey we could resolve, so their cadence is measured but unattributable.
- Gaps longer than the window are invisible. Which is why the mean, not the median, is the figure to quote.