Network Pulse

Network Pulse  /  how the data gets in

Feed to Report

Every bus in the region broadcasts where it is. None of them broadcasts whether it is late. That figure is made here — by matching eight thousand positions a second-hand at a time against the timetable they were supposed to keep.

Real capture 25 Aug 2026 · 20:53–20:59 BST 36 consecutive polls 8,400 vehicles Nothing on this page is illustrative
Northern England bbox −3.80,52.90 → −0.10,55.90
Loading six minutes of positions…
Poll1of 36
Vehicles in the feed8,379
Matched to a journey7,58290.5%
Median delay, matched+63 s
Moving right now2,401
20:53:24
more than 1 min early on time 6–15 min late over 15 min late parked or unmatched
South coast bbox −1.05,50.70 → 0.10,51.25

There is no map underneath this. Every mark is one bus, drawn where it said it was — the road network is what the buses trace out on their own. Six minutes replays in about ten seconds.

At nine in the evening most of the fleet has finished. A bus keeps its place here for fifteen minutes after its last report, so the still marks are real buses parked up, not gaps in the feed.

36
polls, one every 10 s
8,400
distinct vehicles seen
3,358
departures recorded
1,342
of them at timing points
3,112
stops involved
55
operators
BODS GTFS-RT avl-poller vehicle_positions stop_observations nightly rollups the reports tt_* timetable BODS GTFS export 2 bounding boxes one process, in memory 8,376 rows, overwritten 13.1 M rows · 4.3 GB 03:10 – 04:10 punctuality · mileage 403,136 trips whole of GB, weekly every 10 s where it is when it left aggregate read supplies the time it was due Sun 04:00
The whole system in one picture. The left-hand column is everything that arrives from outside; everything to the right of the poller is ours. Note what the live feed does not carry: the arrow that turns a position into a punctuality figure comes up from the timetable, not in from the feed.
Every 10 seconds 8,640 polls a day, two requests each

What actually arrives

Two HTTP requests to the Bus Open Data Service, one per bounding box, repeated round the clock. This is one real pair, timed while writing this page.

Bounding boxPayloadVehiclesWith a journey idResponse
Northern England
−3.80,52.90 → −0.10,55.90
1,112,372 B7,6036,826 · 89.8%0.29 s
South coast
−1.05,50.70 → 0.10,51.25
103,145 B726704 · 97.0%0.17 s

Two boxes rather than one because the two service areas are 250 km apart, and a single rectangle drawn round both is most of England — four times the vehicles to match and store for the same two areas' worth of answers.

One vehicle, exactly as it came off the wire

id: "37fa80bc-fe90-483b-bfb1-1e8aec8d3f03"
    vehicle {
      trip { trip_id: ""
             route_id: "" }
      position { latitude: 53.3651772
                 longitude: -3.06502604
                 bearing: -1 }
      timestamp: 1787681530
      vehicle { id: "A2BV-24" }
    }

That is the whole record. A fleet number, a point, a clock reading. No stop, no route in this case, and — on every vehicle in the country, in both feeds BODS publishes — no delay. Nothing here says whether this bus is early or late. Everything else on this page is built from that absence.

Note the heading, too: bearing: -1 is GTFS-RT for don't know, and a quarter of the fleet sends it — some publishers send 0 for the same thing, which is why due north turns up fourteen times more often than any other single degree. Both are dropped, and a heading is worked out instead from the ground the bus has covered since its last fix. A bus that hasn't moved is drawn without an arrow rather than with one pointing north.

Why GTFS-RT, not SIRI-VM

BODS publishes both. SIRI-VM names a journey only by the operator's own ticket-machine code, which we would then have to match to the timetable ourselves. The GTFS-RT feed has already done that work and quotes a real journey id for around 90% of vehicles — the one number this whole pipeline turns on.

Continuously One process, holding every bus's progress in memory

Where a delay figure comes from

The poller keeps each bus's place along its scheduled journey between polls. When a bus is seen far enough past a stop, it writes down that the bus left — and what time the timetable said it should have.

Below is one real journey from the capture window: Stagecoach Yorkshire 26030, service 28, Pontefract Bus Station 19:45 → Barnsley Interchange 20:55, on 25 August. The timetable gives it 81 stops, ten of them timing points. Sixty-six departures were recorded. Each bar is one of them: how far from its due time the bus actually left.

Loading the journey…
Show these 66 departures as a table
#StopDueLeftDifference

The shaded band is "on time" as the UK convention defines it: one minute early to five minutes late. Every one of these 66 departures sits inside it. The journey lost three and a half minutes and still scores 100% — which is what the convention is for, and also what it hides. Timing points, the stops an operator publishes times for and is held to, are marked below the axis; times at the stops between them are interpolated by the publisher rather than promised, which is why the reporting scores timing points only.

Read left to right, this is a journey losing three and a half minutes through Pontefract in the first ten stops, holding that gap along the A628, and pulling it back to arrive half a minute early. One row of one report, from a bus that never reported a delay at any point.

Timed from the departure, not the arrival

A bus that reaches a timing point early waits there until it is due. Timing the arrival therefore recorded correct behaviour as early running: on 64,479 paired observations, moving to departure-based timing moved "early" from 19.9% to 6.7% and "on time" from 67.4% to 78.6%. Two thirds of the early running we were reporting was an artefact of measuring the wrong moment.

As it happens Two tables, doing opposite jobs

What gets written down

vehicle_positions — the present tense

One row per bus, overwritten every poll. 8,376 rows, never more. This is what the live map and the arrival predictions read; it has no memory at all.

stop_observations — the record

One row per departure, kept for 35 days. 13,055,887 rows, 4.29 GB. Yesterday alone added 1,246,347 of them. Everything a report says about last week is counted out of this table.

The split matters when someone asks why a figure moved. The live board can be wrong for ninety seconds and self-correct; the record cannot be re-derived, because the feed is a snapshot with no history. If the poller is down for an hour, that hour is simply not evidence, and every report that touches it has to say so rather than count it as failure.

The window you have just watched, in the record

early 42 · 3.1% on time 1,096 · 81.7% late 177 · 13.2% over 15 min 27 · 2.0%

1,342 timing-point departures, from 1,361 journeys and 55 operators, in five minutes and fifty seconds. Bands are the UK convention: on time is one minute early to five minutes late.

In code The same account, with the numbers that decide it

The mechanism, in detail

Everything above, restated as it is actually implemented. The values below are the ones the running poller uses; where a figure is a threshold somebody chose, the reason it is that number is given with it.

One poll, in full

GET https://data.bus-data.dft.gov.uk/api/v1/gtfsrtdatafeed/
    ?api_key=<key>&boundingBox=-3.80,52.90,-0.10,55.90

feed = gtfs_realtime_pb2.FeedMessage()
feed.ParseFromString(response.content)      # protobuf, not JSON

Two of those, one per bounding box, every ten seconds — so 17,280 requests a day. The feed is a complete snapshot each time rather than a stream of changes, which is why a missed poll costs resolution and nothing else, and why the poller backs off to a maximum of five minutes rather than retrying hard against an API that is already struggling.

The same bus can arrive twice in one pass. Boxes are not meant to overlap, but the feed also repeats vehicles inside a single box — 8 of 7,838 northern vehicles on 21 August, all of them untracked, where two operators happened to use the same fleet number. Vehicles are therefore keyed on (vehicle_id, route_id, trip_id) and the fresher fix wins: position deliberately plays no part in identity, because two requests a moment apart legitimately carry different fixes for the same bus. Untracked vehicles report 0.0, 0.0 rather than omitting the position, so null-island fixes are dropped before anything else looks at them.

From a point to a place on the journey

A latitude and longitude are not progress. Every question the reports ask — how far along, how late, which stops have been served — comes from one derived number: the distance the vehicle has travelled along its trip. That is obtained by projecting the position onto the chain of the trip's stops and reading the schedule off at that point, interpolating between the stops either side so that a bus halfway between two of them is not counted late by the whole hop.

The chain is straight lines, deliberately

Stop to stop, great-circle, not the road. The export's shapes.txt gives real road geometry and is imported — the lost-mileage figures are measured along it — but the matcher does not use it: 7 million shape points would refine a projection whose answer is already "three stops away, four minutes". Distances along the chain are therefore slight underestimates on curved roads, which matters for mileage and not for matching. The two use different geometry on purpose.

A projection can be wrong, and most of the work is in refusing bad ones. Circular services pass the same road twice; a bus can be legitimately far from the straight line between two stops; equipment goes quiet mid-journey. These are the limits that decide what is accepted:

LimitValueWhat it decides
off route750 mHow far off the stop chain a vehicle can sit and still count as running the trip. Generous, because the chain cuts corners the road does not.
ambiguous snap100 mTwo candidate points this close are treated as equally good and the timetable breaks the tie — sized to cover a circular route's two passes along one road.
progress speed30 m/s~108 km/h. Faster than this is the projection jumping to another part of the route, not a bus moving. Treating it as travel would invent an arrival at every stop in between.
re-snap improvement250 mHow much better a free fit must be before the vehicle's recorded progress is abandoned and restarted. Wider than a dual carriageway, so the two directions of one road cannot trigger it.
observation gap180 sBeyond this between two fixes, the crossing time cannot be interpolated and no departure is timed. The bus went quiet for too long to say when it passed anything.
departure clearance40 mHow far past a stop counts as having left, capped at half the gap to the next stop so closely spaced stops cannot report out of order.
at origin60 mInside this, a bus is on the stand rather than running, and "early" is reported as zero: it has not set off yet.
implausible delay7,200 sTwo hours out is a matching failure, not a late bus — usually the vehicle is running a different journey from the one the feed claims. Discarded rather than recorded.
stale vehicle15 minSilence this long and the bus leaves the live table. Well above the poll interval, because operators' equipment goes quiet for a few minutes at a time.
service-day slack3 hHow far outside its scheduled window a journey can still be the one happening now. A run that starts at 23:50 keeps yesterday's service date, with times past 24:00:00.

The moment a bus counts as having left

At a ten-second poll a bus doing 30 mph moves about 130 metres between fixes, so it can clear a stop without ever being seen near it. Nothing is detected by proximity. A stop is served when the vehicle's progress crosses it — when the distance travelled goes from below the stop's position on the chain to above it — and the time it did so is interpolated between the two fixes either side.

previous_distance < mark <= current_distance   # the whole test

mark = cumulative[i] + min(40 m, gap_to_next / 2)   # departure
mark = cumulative[i]                              # arrival

Two marks per stop, forty metres apart, and which one is used is the single most consequential choice on this page. Arrival is when the bus reaches the stop; departure is when it has cleared it. A bus that arrives early at a timing point waits there until it is due, so timing arrivals records correct behaviour as early running — which is what moved "early" from 19.9% to 6.7% when it was corrected. The terminus is the exception in both directions: a bus that finishes its journey never departs, so the arrival is all there is, and it is compared against the scheduled arrival rather than a departure that does not exist.

Two writes, two shapes

Every poll ends in exactly two statements, and their conflict clauses are where the difference between the live table and the record actually lives.

vehicle_positionsthe present tense
INSERT INTO vehicle_positions (…)
VALUES %s
ON CONFLICT (vehicle_id)
  DO UPDATE SET …
stop_observationsthe record
INSERT INTO stop_observations (…)
VALUES %s
ON CONFLICT (service_date, trip_id,
             stop_sequence)
  DO NOTHING

One row per vehicle, overwritten forever; one row per departure, written once and never revised. DO NOTHING is what makes the record idempotent: the poller can restart mid-journey, re-observe stops it has already recorded, and change nothing. It also means the first observation of a departure is the one that stands, which is the right way round — a later re-match is a guess made with less information, not more.

The record is indexed for the three questions the reports ask of it — by line and date, by stop and date, by operator and date — and pruned to 35 days by the same nightly job that aggregates it. Departures are kept at timing points only, except in two ATCO areas (370 South Yorkshire and 450 West Yorkshire) where the journey-times reports need a run time and a dwell at every stop. That exception is scoped by area because that is what the cost scales with: every stop in South Yorkshire is ~290k rows a weekday, every stop across everything polled would be ~2.3M, and ~5 GB a month becomes ~40 GB. is_timepoint marks which rows are promised times and which are the publisher's interpolations.

What is joined to what

"Combined" is four joins, and only one of them is on anything the live feed provides. Everything else is keyed on identifiers that come out of the weekly timetable export or are already encoded in the stop code itself.

JoinKeyComes fromFails when
vehicle → journeytrip_idquoted by the live feed, resolved against tt_tripsthe operator republishes its timetable — the id is a content hash, so a re-export keeps it and a republish does not
journey → its stopstrip_id, stop_sequencett_stop_times, from the weekly exportnever, once the journey is matched — this is the join that supplies scheduled times
departure → areaLEFT(stop_id, 3)the ATCO code itself; every stop code begins with its administrative areanever — it needs no lookup and no extra storage
departure → operator, lineoperator_noc, line_namecopied onto the observation at write timerarely; denormalised on purpose so a report is never wrong about last month because an operator was recoded this month
report → reportjourney_keybuilt by us from date, operator, line, direction, origin stop and aimed departure — the fields a re-issued timetable leaves aloneit does not, which is the point: 22.8% of one Monday's journeys were given a new trip_id by a republication a week later, and anything joined across two weeks on that id loses them silently

The first row is the whole fragility of the system in one line. A journey id that no longer resolves does not error — it produces a bus that is on the road, in the feed, and absent from every report.

The last row is what we do about it downstream. Matching a live bus has to use the id the feed quotes; a report does not, and since 17 September every journey we schedule is also given a key of our own that a republished timetable cannot move. It is what the exported files ask people to join on.

And then it is arithmetic

The nightly aggregation is one INSERT … SELECT … GROUP BY per job. Punctuality groups a day's departures by service date, area, operator, line and hour, counts each band with a filtered aggregate, and keeps two extra columns — the sum of delays and the sum of their squares — so that variance can be recovered later without going back to rows that will have been pruned:

GROUP BY service_date, LEFT(stop_id, 3), operator_noc, line_name,
         route_id,
         EXTRACT(HOUR FROM scheduled_at AT TIME ZONE 'Europe/London'),
         EXTRACT(MINUTE FROM scheduled_at AT TIME ZONE 'Europe/London') / 30

The route is in that list because a line number is not a service: Stagecoach Yorkshire run two services called 21, Barnsley to Penistone and Rotherham to Harthill, with not one stop in common, and a mean across both describes neither. The half-hour is there because the bands this is reported against do not all turn on the hour — the AM peak ends at 09:30 — and hour keeps its meaning, so anything grouping by it alone still sums both halves. Between them they are why this table holds 67,785 rows for a day rather than 34,279.

The time zone in that last line is not decoration. The server runs on CEST; bucketing on its clock would label the morning peak 09:00 for buses that ran it at 08:00, and that is the sort of number which gets quoted in a meeting long before anyone notices it is an hour out.

03:10 → 04:10 nightly Five jobs, in an order that matters

Turning a night's departures into a day's figures

Nothing on the reports queries thirteen million raw rows. Each night, five jobs reduce the day to something a page can read in milliseconds. These are the counts for Wednesday 16 September 2026.

TimeJobWhat it writesRows, last night
03:10journey timesstop-by-stop run times and dwell, per journey18,659
03:20punctualityon time / early / late per operator, line, route, half-hour, area — and prunes raw departures past 35 days67,785
03:35first and lastwhether each service's first and last journey ran, per route and direction1,020
03:50lost mileagescheduled against operated distance — and keeps the day's schedule, journey by journey, with our own key on each295 + 6,710
04:10pinch pointslink-level congestion scoring6,366 links
The honest caveat we keep on the page

Three of these jobs read the departures table and the one that deletes from it runs second. The order is held by the clock alone — fifteen-minute gaps, no hard dependency — so a badly overrunning job could read a table another had already pruned. It has not happened; it is written down because it could.

Sunday 04:00 The half of the pipeline that isn't live

Keeping the timetable in step

A position only means something against the journey it was supposed to be running. That comes from the national GTFS export, and it goes stale faster than anyone expects.

Currently loaded: export 20260822_024148, imported 23 August — 403,136 journeys, 17,381,506 scheduled stop times, 312,501 stops, 634 operators. One unattended job downloads the national export, rebuilds the journey-planning graph, imports the same export into the timetable tables, and restarts the poller, so the two can never drift apart.

Why weekly, and not "when it changes"

Journey ids are content hashes, so they survive a re-export unchanged — unless the operator republishes, and Stagecoach republishes all thirteen of its datasets daily. Measured a fortnight after an import: 20% of live buses were quoting a journey id we no longer held, and eight Stagecoach North East lines had gone from ~150 journeys matching to zero in a week. Match rate decays visibly within a fortnight, which is what sets the cadence.

The failure this guards against is silent. A re-issued timetable does not produce an error; it produces an operator whose buses are on the road, in the feed, and absent from every report. That is why the reporting carries a coverage panel comparing scheduled journeys against departures actually recorded, per operator.

On demand What the rooms actually ask for

What comes out

Everything above, combined. Every figure below comes out of the two tables on this page and the timetable they are matched against — nothing else feeds them.

ReportCombinesWhere it stands
Punctualitydepartures × timetable × administrative areas83.4% on time yesterday, on 476,061 timed departures across 100 operators, 920 lines, 43 areas. 82.4% over seven days, on 2.8 million.
First and last busdepartures × scheduled journeys × watched hours87.5% of 4,875 watched first and last journeys seen to run, last seven days.
Lost mileagedepartures × route geometry from the export's shapesSouth Yorkshire, last seven days: 420,700 of 485,584 scheduled miles operated — 86.6%, with 94.7% of the distance measured along the road rather than crow-flies.
Journey timesevery stop, not just timing points, in two areasReplaces the operator's own commercial reporting, in the same layout. 467 lines last week.
Pinch pointsdepartures × road geometry × traffic samples6,366 scored road links, ranked by where buses lose time repeatedly.
Live boardpositions × timetable, no historyService status by area, updated as the feed arrives.
Not at all Stated because it will be asked

What this data cannot tell you

The gaps are properties of the feed, not of the build. Each was measured nationally rather than assumed.

How many people were on the bus

Nothing. SIRI 2.0 has no element that can express a passenger number, so no operator could publish one through BODS even if they wanted to. The one crowding field that exists is a three-value word, present on 1.4% of vehicles. Counts have to come from on-board ticketing or APC equipment.

Whether a bus was accessible

The accessibility fields exist in the schema and are unpopulated across the national feed.

Trams, trains, coaches and ferries

No operator of any of these publishes vehicle positions to BODS — verified against the national feed, where every resolvable vehicle is a bus. They are recorded as untrackable rather than as failures, which is why a coach operator never appears at the top of a missed-journeys table.

Anything before 4 August 2026

The record starts when the poller did. It cannot be backfilled: the feed is a live snapshot and BODS keeps no history to go back for.

Buses that never report

An unreported journey is a ceiling, never a count. "Not seen" and "did not run" are the same evidence here, and every delivery figure on the reports is phrased that way. A contract requiring the vehicle to report is what closes the gap.