How a live bus gets matched to a scheduled journey
Status: describes what runs today. Written 2026-08-19; block figures re-measured against the 2026-08-18 timetable import; the last section added 2026-09-17, when reports stopped being keyed on BODS's id.
The short answer: on the GTFS trip_id in the BODS realtime feed, and on
nothing else. Blocks are imported and measured but not wired into anything
live; journey codes we never see, because they belong to a feed we don't read.
That is the match. What a finished report joins on is a separate question
with a separate answer, and it is the last section here.
The only key for the match: trip_id
frontend/avl_poller.py reads vehicle.trip.trip_id out of the BODS GTFS-RT
vehicle-positions feed and looks it up in tt_trips through
avl_timetable.TimetableCache. That lookup is the match. There is no fuzzy
fallback — no "route 120, heading north, 09:14, probably this journey". A
trip id we don't hold means an unmatched bus that contributes nothing to
punctuality, the departure board or the live map.
Everything downstream is geometry rather than identity. avl_match.locate()
projects the position onto the trip's chain of stops to get a distance along
the trip, disambiguating snaps closer than 100 m to each other — the two laps
of a circular route — by the schedule, then rejecting apparent progress above
30 m/s as a projection jump rather than a bus. Vehicle identity for the
vehicle_positions primary key is operator:fleet_id, because fleet numbers
collide between operators.
Why not SIRI-VM, and why journey codes never come up
BODS publishes SIRI-VM alongside GTFS-RT. SIRI identifies a journey by the
operator's own ticket-machine code, which has to be matched back to the
timetable before it means anything — and that matching is exactly what BODS
has already done in the GTFS-RT feed, where a real trip_id resolves for
about 97% of vehicles.
So JourneyCode, TicketMachineServiceCode, BlockRef, DriverRef and
VehicleUniqueId appear nowhere in this repo. They are SIRI fields; GTFS-RT
carries none of them. Adopting them would mean ingesting a second feed and
rebuilding the journey match we currently get for free.
The one thing SIRI has that GTFS-RT does not is BlockRef, on about 66% of
vehicles. That is the only reason to keep it on the table at all.
Blocks: imported, measured, not yet used
import_timetable.py loads tt_trips.block_id from the GTFS export
(blockBasedInterlining was on when the graph was built). As of the
2026-08-18 import:
| trips in the timetable | 313,544 |
| trips carrying a block | 164,088 (52%) |
| distinct blocks | 28,261 |
| blocks spanning two agencies | 0 |
Nothing in the serving path reads any of it, and the decision not to use it is
written up in block chaining, considered. The only
consumer is frontend/block_coverage.py (systemd block-coverage.timer),
which measures
the ceiling for block chaining: of the journeys starting in the next 90
minutes that no bus is visibly on, how many belong to a block that a tracked
bus is working right now.
It reports two scopes, because one number was answering a question nobody asked. Block publication is not uniform: 43% of trips carry one nationally, with the split by operator bimodal rather than partial — Arriva, Go North East and the Bee Network agencies publish none at all — while across trips calling in South Yorkshire it is 96%, First South Yorkshire, Stagecoach Yorkshire and TM Travel publishing on every trip. So the national figure understates what could be done in the area the contract is about, by roughly half:
| 2026-09-22 23:04 London | journeys in 90 min | board live now | with chaining |
|---|---|---|---|
| everywhere the poller watches | 1,413 | 10% | 26% |
| South Yorkshire | 119 | 12% | 50% |
That is a late-evening reading and the daytime shape is what matters; the timer runs from 08:15. Region-wide daytime readings sit at 19–25% of untracked journeys, against roughly 8–9k journeys starting in each window:
| untracked | predictable via block | |
|---|---|---|
| 18 Aug 15:15 | 8,149 | 1,576 (19%) |
| 18 Aug 17:15 | 7,007 | 1,736 (25%) |
| 18 Aug 19:15 | 3,797 | 857 (23%) |
| 18 Aug 21:15 | 3,057 | 694 (23%) |
| 19 Aug 07:15 | 7,850 | 1,497 (19%) |
| 19 Aug 09:15 | 8,701 | 2,135 (25%) |
The 11% first seen at 23:35 was the floor, as suspected — blocks end rather than continue late on.
Two properties that matter for building on this:
- Block coverage is bimodal by operator, never partial. First South Yorkshire and Stagecoach (North East, Yorkshire, Cumbria & North Lancashire, Merseyside) are at 100%; First Leeds 74%, First York 66%. It is zero for all three Arriva divisions (North West 16,277 trips, North East 8,492, Yorkshire 7,585), for Go North East (13,836) and for all three Bee Network agencies (37,446 trips between them). The hole is Arriva plus Go North East plus franchised Manchester.
- Block ids are content hashes and globally unique. None of the 28,261
blocks spans two agencies, so joining on
block_idalone is safe — unlike fleet ids, which needed qualifying by operator. Being hashes, they churn on every operator re-export exactly as trip ids do, so nothing derived from a block may be cached across a timetable import.
The failure mode this design leaves
Because trip id is the only key, timetable freshness is match rate. Measured 2026-08-18, fourteen days after the previous import, 20.3% of live buses (1,624 of 7,982) quoted a trip id we did not hold. Roughly half were genuinely outside the imported regions — East and West Midlands, Wales, Scotland, inside the AVL rectangle but never imported. The other half was pure id churn, and Stagecoach-specific: they republish all 13 of their datasets daily at about 04:00, where First republishes only on real changes.
Blocks would not rescue any of it. A trip id we cannot resolve has no block either, because the block travels in the same row of the same import.
Which is why the timetable wants re-importing weekly rather than occasionally; see the refresh procedure in the AVL notes.
After the match: the key a report joins on
Everything above is about the live match, where there is no choice: the id the
feed quotes is the only thing on offer, and trip_id is it.
A report is a different problem, and since 2026-09-17 it has a different key.
The nightly rollup writes scheduled_journeys.journey_key beside every journey
it records — 2026-09-15:FSYO:120:0:370021639:0519, which is the service date,
the operator, the line, the direction, the origin stop and the aimed departure.
Those are the fields a re-issued timetable leaves alone, where trip_id is a
content hash of the vehicle journey and moves when the journey is republished:
of the 6,710 journeys scheduled in South Yorkshire on Monday 14 September, only
77.2% still carried the id the export ten days earlier had given them. A report
joined across two weeks on that id loses the rest without saying so.
The same four fields are the ones this note's sibling — matching-without-bods.md
— scored at 99.0% as a matcher against BODS's own trip id. That is not a
coincidence: what identifies a journey well enough to match it is what
identifies it well enough to key it.
The service-day export publishes journey_key on every grain and says to join
on it. journey_key.py holds the one definition, and
venv/bin/python journey_key.py --verify re-derives every stored key and
reports any that disagree.
Re-measuring
# block coverage right now (also runs on block-coverage.timer through the day)
cd frontend && ./venv/bin/python block_coverage.py
Anchor anything you write against vehicle_positions to
(NOW() AT TIME ZONE 'Europe/London')::date, never CURRENT_DATE — the box
runs CEST, so after 23:00 London its current date is already tomorrow and every
join silently returns nothing.