Matching a live bus to its journey without BODS doing it for us
Status: measured, and serving nothing — a comparison rig, built 2026-08-27.
Figures below are from South Yorkshire on that date unless a paragraph says
otherwise; the RTIG comparison carries its own, re-measured 2026-09-22 against
the live feed. Nothing in the live path changed; the only thing that reads any
of it is the explanation of it on /feed, under If BODS stopped naming the
journey.
For the plain-English version of what follows, see how we know which journey a bus is on. This note is the workings: the measurements, and the traps that produced them.
Every live figure this system produces — punctuality, lost mileage, first and
last bus, the departure board, the dot on the map — hangs on one field. BODS
publishes a GTFS trip_id per vehicle in its GTFS-RT feed, we look it up, and
that lookup is the match (how a live bus gets matched).
There is no fallback.
Two things on the roadmap take that field away. Franchised on-board equipment reports the operator's own identity for the work — a duty, a block, a journey code — and knows nothing of GTFS. SIRI-VM, the other feed BODS publishes, is the same story. So the honest question behind "shall we take the feed direct" is not architectural, it is arithmetic:
If BODS stopped telling us which journey a bus is on, could we work it out ourselves, and how often would we be right?
That is measurable, because BODS hands us the answer key. GTFS-RT and SIRI-VM describe the same buses at the same moment. Take both at once, throw away the trip id BODS supplies, rebuild it from what the other feed publishes, and count the agreements.
The short answer
It is not very complicated, and the reason is that the operator's journey code turns out not to be the interesting part.
A journey is pinned down by four things almost any source publishes:
operator (NOC) + line + the stop it starts from + when it leaves there
resolved against the service calendar for the day in question. That is the
whole method. It is one SQL query, in frontend/journey_key.py, and it is the
same query whichever source the description came from.
What the sources actually give us
Measured over 2,609 vehicle observations in the South Yorkshire box, three samples across 2026-08-27:
| field | published on |
|---|---|
DatedVehicleJourneyRef |
100% |
LineRef, DirectionRef, OperatorRef, OriginRef |
100% |
OriginAimedDepartureTime |
98.0% |
BlockRef |
70.1% |
ticket-machine JourneyCode (in Extensions) |
70.1% |
So SIRI-VM already publishes the origin stop and the aimed departure time on 98% of vehicles. No TransXChange is needed to match those at all.
Where TransXChange comes in
TransXChange is the registration-level document BODS converts into the GTFS we import, and it is the only place the operator's own vocabulary is defined alongside the schedule that vocabulary refers to:
<Operational>
<Block><BlockNumber>09702</BlockNumber></Block>
<TicketMachine><JourneyCode>0750I</JourneyCode></TicketMachine>
</Operational>
<VehicleJourneyCode>VJ6418</VehicleJourneyCode>
<DepartureTime>07:50:00</DepartureTime>
None of that survives conversion to GTFS, which is exactly why the cross-walk has to be built from TransXChange rather than from what we already hold.
The important design point is that TransXChange does not need a second
matching algorithm. A journey code is an alias for "this line, from this
stop, at this time". Once the code has been resolved to that description it
goes through the same resolver as an aimed departure time does. So the whole
job of frontend/txc_crosswalk.py is:
(operator, line, journey code, date) → (origin stop, departure time)
which is a dictionary, built offline, of a few hundred thousand rows.
That matters for on-board AVL, which is the case this was really built for. Equipment will report a duty or a journey code and no aimed departure time at all. The cross-walk is the only bridge from what it says to what our timetable calls the same journey — and under franchising it is a bridge we can actually build, because the schedule that loads the ticket machine is one we specify.
Results
Scored against BODS's own answer, South Yorkshire, three samples across 2026-08-27: 2,609 vehicle observations present in both feeds that BODS named a journey for. Of those, 2,561 (98.2%) have a BODS trip id our timetable actually holds — the fair test, since no method can produce an id that is not in the table.
| strategy | agreed | wrong | not found | nothing to match on |
|---|---|---|---|---|
| aimed departure | 98.1% | 0.0% | 0.1% | 1.8% |
journey code via TransXChange (DatedVehicleJourneyRef) |
82.8% | 0.4% | 16.8% | — |
journey code via TransXChange (ticket-machine JourneyCode) |
8.3% | 0.7% | 61.7% | 29.3% |
| aimed departure, falling back to the journey code | 99.0% | 0.0% | 1.0% | — |
| running board + clock, via TransXChange | 84.8% | 9.6% | 4.0% | — |
Read the columns, not just the first one. "Wrong" is the only dangerous outcome, because a wrong trip id is written down as fact and nothing downstream can tell it was a guess; the combined strategy produced none at all, in 2,561 attempts.
The running-board row is scored on the vehicles whose block we hold a registration for, and is the odd one out in every respect: it is the only route with a material wrong-answer rate, and the only one whose ceiling is not a data-quality problem. A block is a plan. Buses are swapped between boards, run short and get reallocated during the day, and the registration cannot know — so roughly one in seven is attributed to a journey the vehicle is not on. Two rounds of work took it from 83.1% only to 84.8%, first by importing run times so a journey's end is known, then by ranking a journey under way above one merely due; what is left is the road disagreeing with the plan.
It matters because it is the only route franchised on-board equipment can offer unaided. Specifying that equipment report the journey it is working, or that journey's due time, keeps the matching in the high nineties. Taking the duty alone accepts 85%.
Three things follow.
The aimed departure time alone very nearly does it. 98.1%, no errors. Its
whole weakness is the 1.8% of vehicles publishing no OriginAimedDepartureTime
— for those there is nothing to match on, and no cleverness recovers it.
TransXChange closes exactly that gap. On its own the journey-code route is worse (82.8%), because it depends on the operator's code appearing in the registration we hold. But it is available on 100% of vehicles, so it covers the cases the aimed departure cannot, and the combination reaches 99.0% with nothing wrong. The two are complements, not rivals.
Use DatedVehicleJourneyRef, not the ticket-machine JourneyCode. They
are different fields with different values, and only the first indexes
TransXChange: 82.8% against 8.3%. This is the single most expensive thing to
get wrong here, and both are plausible-looking journey codes sitting in the
same record.
Per operator, counting only journeys our timetable holds, the combined strategy agreed on 100% of First South Yorkshire's 758 observations, 100% of Stagecoach Yorkshire's 698, 99.8% of Arriva Yorkshire's 559, and 100% of every other operator with ten or more except First Leeds, at 95.0% of 60.
Measured against RTIG's guidance
RTIG's How does bus real-time information work? (RTIGT063 v1.0, May 2025) puts data matching on a ladder on page 8, least to most effective:
| RTIG's tier | What they match | Their verdict |
|---|---|---|
| Fuzzy match | timetable DepartureTime against AVL StartTime |
"least accurate", "results can be hit-or-miss" |
| Key field match | timetable JourneyCode against AVL DatedVehicleJourneyRef |
"moderately accurate", "helps accuracy but still limited" |
| Block match | timetable BlockNumber against AVL BlockRef |
"most accurate", "BEST MATCH" |
Those are the three strategies scored above, which makes this a direct test of the ladder rather than an opinion about it. Against BODS's own answer key, on 2,561 South Yorkshire observations, the order inverts:
| RTIG's tier | Their ranking | Agreed | Wrong |
|---|---|---|---|
| Fuzzy — aimed departure | worst | 98.1% | 0.0% |
Key field — DatedVehicleJourneyRef via TxC |
middle | 82.8% | 0.4% |
| Block — running board and clock | best | 84.8% | 9.6% |
| fuzzy, falling back to key field | not offered | 99.0% | 0.0% |
The wrong column is the one that decides it. A wrong journey is written down as fact and nothing downstream can tell it was a guess; the top tier is the only one that produces them in quantity, and the bottom tier produced none at all in 2,561 attempts.
The reason is in the section above and worth repeating here, because it is a property of blocks rather than of anyone's implementation: a block is a plan. Buses are swapped between boards, run short and get reallocated during the day, and the registration cannot know. Two rounds of work moved the running-board route from 83.1% to 84.8% and no further; what is left is the road disagreeing with the plan. Matching on what the bus is doing beats matching on what it was rostered to do, and the gap is about one journey in seven.
Where the guidance is right, and it matters
The block tier's own bullet is "allows predictions, even when a bus vehicle switches to a new route" — attributing a bus to a journey it has not started yet. That is a different job from naming the journey a bus is on now, and for it a block is not merely the best method, it is the only one. Nothing else can say anything at all about a journey that has not begun, and almost every journey on a departure board has not begun — measured across a day, only 3-5% of journeys starting within 90 minutes have a bus visibly on them.
What a block buys there is coverage, not accuracy, and the distinction took a correction to see. This note said for a time that predictions here were less accurate than the incumbent at every lead time. Re-scored on identical departures, they are not: the difference is 0-9 seconds per lead bucket and a dead heat in two of five. The earlier figure scored the incumbent only on the departures it chose to answer — 36-59% of them, depending on lead — and scored us on every one, which compares two populations rather than two predictors. Emitting the scheduled time instead of extrapolating was tested at the same time and is worse at every lead under 40 minutes.
So the ladder is sound if read as prediction lead time and misleading if read as identification accuracy, which is how the page is laid out. Both things are true at once: their worst tier is the best way to know which journey a bus is on, and their best tier is the only way to say anything about a journey that has not begun.
Two traps the diagram sets up
There are two fields called JourneyCode. The diagram is right — AVL
DatedVehicleJourneyRef matches timetable JourneyCode — but SIRI-VM also
carries a field literally named JourneyCode, the ticket machine's, in
Extensions in the same record. Re-measured 2026-09-22 against TransXChange,
for operators we hold registrations for: DatedVehicleJourneyRef resolves
96.5%, the ticket-machine JourneyCode 33.1%. Reading the diagram and
reaching for the field whose name matches costs two thirds of the matches.
BlockRef does not join to BODS's timetable. The picture shows
<BlockNumber>57839</BlockNumber> matching <BlockRef>57839</BlockRef>, which
holds only if the timetable side is TransXChange — as RTIG says on the same
page. Take the timetable from BODS's GTFS conversion instead, as most
consumers do, and the two sides are in different namespaces: GTFS block_id
is a 40-character content hash, SIRI BlockRef is the operator's own number.
SIRI BlockRef 92, 1045, 82, 81, 80
GTFS block_id 0b5a3ab202d6b49f7717d5b52de1d28bd0451990
Measured 2026-09-22: 1,604 distinct BlockRef values live, 0 matching any
GTFS block_id. Against TransXChange it does work, but only where both sides
publish. Of 1,757 live operator-and-block pairs: 46% were operators absent
from our registration cache, 30% operators publishing no BlockNumber at all,
6% published but not found, 18% matched — 74% among the operators that
publish blocks. Our imported GTFS carries a block on 43.1% of 323,663 trips,
and the split by operator is bimodal rather than partial: Arriva, Go North
East and the Bee Network agencies publish none at all.
What follows for us
Nothing changes in the live path, and the guidance is not a reason to change it. We are not on this ladder: BODS resolves the journey in GTFS-RT before we see it, and an id that is published is correct by definition where a reconstruction is inference. Moving to any of the three tiers would replace a known-correct field with a 99.0%-at-best reconstruction of it, for no gain.
What the guidance does argue for is the thing already in progress. Keep the
SIRI reader running for BlockRef, which GTFS-RT does not carry at all, so
the option stays open — none of it can be backfilled.
Whether to spend blocks on cross-journey prediction is written up on its own, with the measurements and the condition that would change the answer: block chaining, considered. The short version is that it buys coverage rather than accuracy today, and becomes unavoidable the day the journey stops arriving pre-matched.
The traps
Every one of these cost real accuracy, and none of them are about parsing XML.
The service day is not the calendar day, and ignoring that costs 7 points. GTFS files a 00:20 journey under the previous service day, timed 24:20. Match a bus in the small hours against today's calendar and you find nothing. Scoring against today alone agreed on 91.5% of an 849-vehicle sample; allowing the previous service day as well took the same sample to 96.5%. Both must be tried, and the figures in the results table above already do.
Journey codes are local time; SIRI timestamps are UTC. A journey code of
0935 and an OriginAimedDepartureTime of 08:35:00+00:00 are the same
moment all summer. Mixing them puts every bus an hour out between March and
October, and looks perfectly correct in winter.
A journey code is not reliably a time at all. Most operators' codes are the
local departure time, and it is tempting to parse them. Some are not: Arriva
Yorkshire published 2022 for a journey leaving at 09:55, TransPennine 5413
for one leaving 08:07. Reading those as times puts the bus twelve hours out.
The cross-walk is what decides what a code means; the code is never parsed.
The two journey codes in a SIRI record disagree, routinely.
DatedVehicleJourneyRef and the ticket-machine JourneyCode in Extensions
are different fields with different values — in the 933-vehicle field survey
they agreed on 81 of First South Yorkshire's 259 vehicles and on 1 of Arriva's
186. Both are scored separately above rather than treated as one thing.
Loosening the key makes it worse, not better. Matching the origin as "any stop on the trip, within 90 seconds" instead of "the first stop, exactly" recovered no extra journeys at all and made 5% of the rest ambiguous. The journeys that miss are ones our timetable vintage does not hold; they are not journeys we described badly.
A byte-order mark silently emptied ten of the 28 datasets. They parsed to
zero journeys with no error — Arriva's 82 lines among them — because the sniff
that decides "is this XML" saw \xef\xbb\xbf rather than <. A silent zero is
indistinguishable from an operator who registered nothing, which is the failure
mode this whole exercise exists to prevent.
Several registrations for the same line are live at once, differing only by
OperatingPeriod/StartDate, usually with no end date. Validity is "the latest
start date not after the service date", not a range test.
OperatorRef on a journey is the file's local id, not the national one.
FSY in the document, FSYO in SIRI. The mapping is in <Operators>.
What this does not solve
Timetable vintage. When BODS names a trip id our GTFS export does not contain, no matching skill can produce it — the id is not in the table. That is already the largest single cause of unmatched buses today, and it stays exactly as large under any of these strategies. The results table separates it out for that reason: the "fair test" row excludes vehicles whose BODS trip we do not hold, because those measure our import freshness rather than our matching.
Running it
# build the dictionary (downloads ~115 MB, caches it, ~6 minutes)
python txc_crosswalk.py build --admin-area 370
python txc_crosswalk.py stats
# score the strategies against BODS's own answers
python compare_matching.py
python compare_matching.py --samples 6 --gap 300 --json out.json
txc_crosswalk.py owns its table (txc_journeys) with plain DDL rather than a
migration, deliberately: nothing in the live path reads it, and it should not
join the production schema until something does.
On the site
/feed carries the argument as a diagram and a comparison, in a band called
If BODS stopped naming the journey. Two things there are read live from the
database rather than written down: the share of vehicles BODS currently names a
journey for, which comes from the same snapshot as the rest of that page, and
the size of the cross-walk. The accuracy figures are the dated measurement
above, captioned with the sample and the command that reproduces them, because
scoring them takes minutes against the live feed and is not something to do
while someone is reading the page.
If the cross-walk table is absent — a deployment that has never run the build — the section says so rather than quoting a figure it does not have.