Network Pulse

← All notes

How a live bus gets matched to a scheduled journey

Status: describes what runs today. Written 2026-08-19; block figures re-measured against the 2026-08-18 timetable import; the last section added 2026-09-17, when reports stopped being keyed on BODS's id.

The short answer: on the GTFS trip_id in the BODS realtime feed, and on nothing else. Blocks are imported and measured but not wired into anything live; journey codes we never see, because they belong to a feed we don't read. That is the match. What a finished report joins on is a separate question with a separate answer, and it is the last section here.

The only key for the match: trip_id

frontend/avl_poller.py reads vehicle.trip.trip_id out of the BODS GTFS-RT vehicle-positions feed and looks it up in tt_trips through avl_timetable.TimetableCache. That lookup is the match. There is no fuzzy fallback — no "route 120, heading north, 09:14, probably this journey". A trip id we don't hold means an unmatched bus that contributes nothing to punctuality, the departure board or the live map.

Everything downstream is geometry rather than identity. avl_match.locate() projects the position onto the trip's chain of stops to get a distance along the trip, disambiguating snaps closer than 100 m to each other — the two laps of a circular route — by the schedule, then rejecting apparent progress above 30 m/s as a projection jump rather than a bus. Vehicle identity for the vehicle_positions primary key is operator:fleet_id, because fleet numbers collide between operators.

Why not SIRI-VM, and why journey codes never come up

BODS publishes SIRI-VM alongside GTFS-RT. SIRI identifies a journey by the operator's own ticket-machine code, which has to be matched back to the timetable before it means anything — and that matching is exactly what BODS has already done in the GTFS-RT feed, where a real trip_id resolves for about 97% of vehicles.

So JourneyCode, TicketMachineServiceCode, BlockRef, DriverRef and VehicleUniqueId appear nowhere in this repo. They are SIRI fields; GTFS-RT carries none of them. Adopting them would mean ingesting a second feed and rebuilding the journey match we currently get for free.

The one thing SIRI has that GTFS-RT does not is BlockRef, on about 66% of vehicles. That is the only reason to keep it on the table at all.

Blocks: imported, measured, not yet used

import_timetable.py loads tt_trips.block_id from the GTFS export (blockBasedInterlining was on when the graph was built). As of the 2026-08-18 import:

trips in the timetable 313,544
trips carrying a block 164,088 (52%)
distinct blocks 28,261
blocks spanning two agencies 0

Nothing in the serving path reads any of it, and the decision not to use it is written up in block chaining, considered. The only consumer is frontend/block_coverage.py (systemd block-coverage.timer), which measures the ceiling for block chaining: of the journeys starting in the next 90 minutes that no bus is visibly on, how many belong to a block that a tracked bus is working right now.

It reports two scopes, because one number was answering a question nobody asked. Block publication is not uniform: 43% of trips carry one nationally, with the split by operator bimodal rather than partial — Arriva, Go North East and the Bee Network agencies publish none at all — while across trips calling in South Yorkshire it is 96%, First South Yorkshire, Stagecoach Yorkshire and TM Travel publishing on every trip. So the national figure understates what could be done in the area the contract is about, by roughly half:

2026-09-22 23:04 London journeys in 90 min board live now with chaining
everywhere the poller watches 1,413 10% 26%
South Yorkshire 119 12% 50%

That is a late-evening reading and the daytime shape is what matters; the timer runs from 08:15. Region-wide daytime readings sit at 19–25% of untracked journeys, against roughly 8–9k journeys starting in each window:

untracked predictable via block
18 Aug 15:15 8,149 1,576 (19%)
18 Aug 17:15 7,007 1,736 (25%)
18 Aug 19:15 3,797 857 (23%)
18 Aug 21:15 3,057 694 (23%)
19 Aug 07:15 7,850 1,497 (19%)
19 Aug 09:15 8,701 2,135 (25%)

The 11% first seen at 23:35 was the floor, as suspected — blocks end rather than continue late on.

Two properties that matter for building on this:

  • Block coverage is bimodal by operator, never partial. First South Yorkshire and Stagecoach (North East, Yorkshire, Cumbria & North Lancashire, Merseyside) are at 100%; First Leeds 74%, First York 66%. It is zero for all three Arriva divisions (North West 16,277 trips, North East 8,492, Yorkshire 7,585), for Go North East (13,836) and for all three Bee Network agencies (37,446 trips between them). The hole is Arriva plus Go North East plus franchised Manchester.
  • Block ids are content hashes and globally unique. None of the 28,261 blocks spans two agencies, so joining on block_id alone is safe — unlike fleet ids, which needed qualifying by operator. Being hashes, they churn on every operator re-export exactly as trip ids do, so nothing derived from a block may be cached across a timetable import.

The failure mode this design leaves

Because trip id is the only key, timetable freshness is match rate. Measured 2026-08-18, fourteen days after the previous import, 20.3% of live buses (1,624 of 7,982) quoted a trip id we did not hold. Roughly half were genuinely outside the imported regions — East and West Midlands, Wales, Scotland, inside the AVL rectangle but never imported. The other half was pure id churn, and Stagecoach-specific: they republish all 13 of their datasets daily at about 04:00, where First republishes only on real changes.

Blocks would not rescue any of it. A trip id we cannot resolve has no block either, because the block travels in the same row of the same import.

Which is why the timetable wants re-importing weekly rather than occasionally; see the refresh procedure in the AVL notes.

After the match: the key a report joins on

Everything above is about the live match, where there is no choice: the id the feed quotes is the only thing on offer, and trip_id is it.

A report is a different problem, and since 2026-09-17 it has a different key. The nightly rollup writes scheduled_journeys.journey_key beside every journey it records — 2026-09-15:FSYO:120:0:370021639:0519, which is the service date, the operator, the line, the direction, the origin stop and the aimed departure. Those are the fields a re-issued timetable leaves alone, where trip_id is a content hash of the vehicle journey and moves when the journey is republished: of the 6,710 journeys scheduled in South Yorkshire on Monday 14 September, only 77.2% still carried the id the export ten days earlier had given them. A report joined across two weeks on that id loses the rest without saying so.

The same four fields are the ones this note's sibling — matching-without-bods.md — scored at 99.0% as a matcher against BODS's own trip id. That is not a coincidence: what identifies a journey well enough to match it is what identifies it well enough to key it.

The service-day export publishes journey_key on every grain and says to join on it. journey_key.py holds the one definition, and venv/bin/python journey_key.py --verify re-derives every stored key and reports any that disagree.

Re-measuring

# block coverage right now (also runs on block-coverage.timer through the day)
cd frontend && ./venv/bin/python block_coverage.py

Anchor anything you write against vehicle_positions to (NOW() AT TIME ZONE 'Europe/London')::date, never CURRENT_DATE — the box runs CEST, so after 23:00 London its current date is already tomorrow and every join silently returns nothing.