Network Pulse

← All notes

How often buses actually report, and what our poll interval can do about it

Status: measured. Evening baseline taken 2026-08-18; peak re-run scheduled for 2026-08-19 08:00 BST.

The question behind this: our poller asks BODS for vehicle positions every ten seconds — until 2026-08-18 it was fifteen — and buses report on their own unsynchronised clocks. So does polling faster get fresher data, or is the limit somewhere we don't control?

It is somewhere we don't control. Buses in service report every ~19 seconds on average, so at a ten-second poll we already sample about twice as fast as the data changes. Going to five seconds would collect the same fixes about 2.5 seconds sooner — latency, not information — for double the requests and double the write volume.

Where the staleness actually comes from

Three delays stack between a bus moving and a rider seeing it move:

typical who controls it
Bus reports its position ~19s between fixes the operator's telematics
BODS ingests and republishes seconds DfT
Our poll picks it up ≤10s, mean 5s us

Measured at ingest — the gap between a fix's own timestamp and our write of it — the median is 19.9s and the p90 is 1,037s (17 minutes). Only 807 of 3,812 vehicles were carrying a fix fresher than one poll interval. The long tail is BODS's own stale tail, not our lag: about 14% of the feed at any moment is vehicles whose last fix is old (see the SIRI feed inventory).

Our contribution to that median is the poll wait, which averages half the interval: 5 seconds at a 10s poll, 7.5s at 15s.

The measurement, and the trap in it

There is no history to read. vehicle_positions holds one row per bus, overwritten every poll, so the gaps between one bus's fixes exist only as they happen. frontend/operator_cadence.py watches the table for a window and records them. It adds no load on BODS — it reads what the poller already fetched — and the gaps come from the feed's own timestamps, so they are exact whenever they are observed at all.

Two figures per operator, because either alone misleads:

  • mean gap = observed vehicle-time ÷ fixes seen. Unbiased: a bus that goes quiet contributes its whole silence to the denominator.
  • median gap = the middle observed gap. Biased low and cannot be otherwise, since a gap longer than the window is never seen. Read it as a floor.

The trap is the population, and it is easy to fall into. Diffing consecutive BODS fetches and counting changed timestamps gives about 6% of vehicles per 5 seconds, which implies a 75–85 second cadence. That is wrong. At night the feed is mostly husks — vehicles whose last fix is hours old and which will not change again before morning. Measured across the raw feed, the median fix age was 10,640 seconds. Those entries sit in the denominator and halve the apparent rate.

Three ways of asking the same minute of feed data, 2026-08-18:

method population answer
% changed over a 5s step whole feed 5.9% → implies ~85s
% changed over a 20s step whole feed 17.7% → implies ~113s
exposure ÷ fixes whole feed 84s
exposure ÷ fixes fix newer than 20 min 19s

Only the last is a cadence. The others are a measure of how many buses are parked. Any future claim about reporting frequency has to say which population it measured.

Per-operator, evening baseline

12 minutes from about 22:05 BST on Tuesday 2026-08-18, whole polled area (northern England), 1,727 buses across 38 operators. Slowest first.

operator buses mean gap median silent name
BPTR 19 32s 20s 10.5% The Burnley Bus Company
GAWY 7 31s 20s 14.3% East Yorkshire
DAGC 6 26s 24s 16.7% D & G Bus
EYMS 26 25s 23s 11.5% East Yorkshire
KDTR 13 25s 21s 7.7% The Keighley Bus Company
TMTL 7 25s 16s 14.3% TM Travel
ANWE 17 25s 21s 11.8% Arriva North West
HRGT 12 25s 22s 16.7% The Harrogate Bus Company
AMSY 159 24s 19s 5.0% Arriva North West
ANEA 54 24s 22s 5.6% Arriva North East
TLCT 5 24s 28s 0.0% TLC Travel
TPEN 9 24s 16s 11.1% Team Pennine
ACYM 8 23s 24s 0.0% Arriva Wales
FYOR 23 23s 23s 4.3% First York
PBLT 16 23s 22s 0.0% Preston Bus
SCCU 82 23s 20s 6.1% Stagecoach Cumbria and North Lancashire
GNEL 143 23s 21s 5.6% Go North East
SCMY 57 23s 20s 7.0% Stagecoach Merseyside and South Lancashire
ANUM 39 23s 22s 5.1% Arriva North East
LNUD 20 23s 17s 5.0% The Blackburn Bus Company
SCEM 40 22s 20s 7.5% Stagecoach East Midlands
FBRA 30 22s 15s 13.3% First Bradford
WRAY 56 22s 18s 3.6% Arriva Yorkshire
SCNE 110 21s 20s 0.0% Stagecoach North East
FHUD 20 21s 18s 5.0% First Halifax, Calder Valley & Huddersfield
SYRK 65 20s 20s 1.5% Stagecoach Yorkshire
FLDS 98 20s 17s 5.1% First Leeds
FHAL 18 20s 16s 5.6% First Halifax
FSYO 69 20s 17s 4.3% First South Yorkshire
WBTR 16 18s 19s 6.2% Warrington's Own Buses
FLYE 6 18s 15s 0.0% Flyer
HUYT 10 16s 17s 0.0% Huyton Travel
BLAC 27 15s 15s 3.7% Blackpool Transport
BNSM 150 15s 12s 9.3% Bee Network
BNML 150 14s 12s 12.0% Bee Network
BNGN 101 13s 12s 6.9% Bee Network
BNDB 27 13s 12s 0.0% Bee Network
BNFM 12 12s 12s 0.0% Bee Network

Fleet-weighted mean gap 19s. A further 12 operators had fewer than 5 buses out and are not listed; 338 buses carried no resolved journey, so no operator could be attributed to them.

The South Yorkshire finding is a negative one, and worth stating plainly: no South Yorkshire operator is a slow-reporting outlier. First South Yorkshire and Stagecoach Yorkshire both sit at 20s, at or slightly better than the fleet-weighted mean, with low silent shares. TM Travel's 25s is on 7 buses, which is too thin to put in front of anyone.

The comparison that does carry an argument is Bee Network's 12–13s across all five of its NOCs: Greater Manchester's buses report roughly twice as often as ours, on the same national feed. That is what an authority specifying the telematics looks like from the outside.

What follows for the poll interval

AVL_POLL_SECONDS=10 since 2026-08-18 17:27. Fifteen was never a decision — it was the default in avl_poller.py. The change bought about 2.5 seconds of median freshness.

Halving again, to five seconds:

15s 10s (now) 5s
Polls per day 5,760 8,640 17,280
Row updates per day ~44M ~66M ~133M
Mean poll wait added 7.5s 5.0s 2.5s
Fixes captured ~all ~all ~all

The last row is the argument. At 10s we sample twice as fast as buses report, so we are already catching essentially every fix; five seconds collects the same fixes sooner rather than collecting more of them. Not worth doing.

The same 19s cadence bounds departure-time interpolation at timing points: a bus at 30 km/h covers about 165 metres between fixes, and no poll rate of ours improves on that. If we want genuinely fresher data, the lever is upstream — which operators' equipment reports slowly — and the table above is the list to take to them.

Re-running it

cd /opt/journey-planner/frontend
venv/bin/python operator_cadence.py --minutes 30            # whole polled area
venv/bin/python operator_cadence.py --minutes 30 --box 370  # South Yorkshire

--box takes a named area (370 South Yorkshire, 450 West Yorkshire, 320 East Riding & Hull) or an explicit minLon,minLat,maxLon,maxLat ordered as AVL_BOUNDING_BOX is. It is a rough rectangle on purpose: this measures telematics cadence, not which authority a journey serves.

A peak run is scheduled for 2026-08-19 09:00 CEST / 08:00 BST via a transient systemd timer (cadence-peak.timer, running /usr/local/bin/cadence-peak-run), taking both views in parallel for 30 minutes into data/cadence-peak-{all,sy}-2026-08-19.txt. The timer is transient and will not survive a reboot of the host.

What this does not tell you

  • One evening. Several operators had fewer than ten buses out. Nobody's number should be quoted at them before the peak run.
  • Only buses that report at all. The population is vehicles whose last fix is under twenty minutes old. An operator whose equipment reports every half hour is excluded entirely rather than appearing as slow; the "silent" column — buses that reported nothing new for the whole window — only partly catches this.
  • Only resolved buses carry an operator. 338 vehicles in the baseline had no journey we could resolve, so their cadence is measured but unattributable.
  • Gaps longer than the window are invisible. Which is why the mean, not the median, is the figure to quote.