Which fields the match is made on, and what happens without each source
Status: measured, written 2026-09-08. The three fields that join a live bus to a scheduled journey, where they are in SIRI-VM, where the same three things live in TransXChange, and what is still measurable if either source goes away.
Companions rather than repeats: how we know which journey a bus is on is the plain-English version, matching without BODS is the scoring rig and its numbers, and what comes out of BODS is the field-level reference for the feed as a whole.
The key
line + the stop it started from + the time it was due away there
resolved against the service calendar for that day. Three fields. Everything downstream — punctuality, lost mileage, did-it-run, the delay on the map — is about a scheduled journey, and this is the only thing connecting a dot on a map to one.
Where they are in SIRI-VM
A real record from the archive, 8 September 2026, trimmed only of whitespace:
<VehicleActivity>
<RecordedAtTime>2026-09-08T15:47:15+00:00</RecordedAtTime>
<MonitoredVehicleJourney>
<LineRef>7</LineRef>
<DirectionRef>outbound</DirectionRef>
<FramedVehicleJourneyRef>
<DataFrameRef>2026-09-08</DataFrameRef>
<DatedVehicleJourneyRef>3049</DatedVehicleJourneyRef>
</FramedVehicleJourneyRef>
<PublishedLineName>7</PublishedLineName> <!-- 1 -->
<OperatorRef>SYRK</OperatorRef>
<OriginRef>370021636</OriginRef> <!-- 2 -->
<OriginName>Crystal Peaks CP5</OriginName>
<DestinationRef>370023381</DestinationRef>
<DestinationName>Wordsworth Avenue</DestinationName>
<OriginAimedDepartureTime>2026-09-08T14:48:00+00:00</OriginAimedDepartureTime> <!-- 3 -->
<Monitored>true</Monitored>
<VehicleLocation>
<Longitude>-1.478351</Longitude>
<Latitude>53.393391</Latitude>
</VehicleLocation>
<Bearing>324</Bearing>
<VehicleRef>SYRK-11702</VehicleRef>
</MonitoredVehicleJourney>
<Extensions><VehicleJourney><DriverRef>79023</DriverRef></VehicleJourney></Extensions>
</VehicleActivity>
| SIRI-VM field | Stored as | Joins to, in GTFS | |
|---|---|---|---|
| 1 | PublishedLineName, falling back to LineRef |
line |
routes.short_name |
| 2 | OriginRef |
origin_ref |
stops.stop_id — an ATCO code, so it joins straight in |
| 3 | OriginAimedDepartureTime |
origin_dep_local |
stop_times.departure_time at the journey's first stop |
All three sit inside MonitoredVehicleJourney. The poller reads them as
.//PublishedLineName, .//OriginRef and .//OriginAimedDepartureTime.
Two things about field 3 decide whether the join works at all:
- It is offset-stamped UTC and GTFS is local service time.
14:48:00+00:00against adeparture_timeof15:48:00. Compare them raw and every BST journey is an hour out, which is why it is stored converted:format(od, "%H:%M:%S", tz = "Europe/London"). - It is the field with the coverage problem. Present on 846 of 915 vehicles (92%) in one poll of the South Yorkshire box on 8 September 2026, and absent entirely from several operators, whose buses can therefore never be matched.
The trap in the same record
DatedVehicleJourneyRef is 3049. It is published on 100% of vehicles, it
looks exactly like a journey identifier, and it is the operator's own
ticket-machine code. Joined to GTFS trip_id it matches 0%. Everyone
reaching into this XML for the first time reaches for it.
Does any of this need GTFS-RT?
No. The proof of concept reads SIRI-VM and a timetable, and nothing else. The
trip_id dependency described in matching without
BODS belongs to the live journey planner, which takes
GTFS-RT and looks the id up. The proof of concept has never had that field and
was built without it.
What it does need is a timetable — and that is a "you must know what was supposed to happen" dependency rather than a BODS one. Any schedule source answers it.
If all we had were raw SIRI-VM
The feed carries two scheduled times of its own. From that same poll of 915 vehicles:
| Field | Populated |
|---|---|
OriginAimedDepartureTime |
846 (92%) |
DestinationAimedArrivalTime |
554 (61%) |
Still measurable, with no timetable at all:
- Punctuality at the origin, against the bus's own aimed departure. Stop coordinates would come from NaPTAN — a separate national dataset, not the BODS timetable.
- Lateness at the destination, for the 61% publishing an aimed arrival.
- Headways and excess waiting time. No schedule needed, and the right measure on high-frequency corridors anyway.
- Journey times between two points, speeds, coverage, vehicle counts, route deviation.
No longer possible:
- Did the service run, and lost mileage. Both need a denominator of what should have run, and no live feed can supply it — you cannot see a bus that never turned up.
- Punctuality at intermediate timing points, which is the measure official punctuality uses.
There is a sharper objection to the origin-only version. This method deliberately excludes the first and last stop of every journey, because a bus sitting at a terminus reads as early or late without saying anything about the service. Origin-aimed-departure punctuality is precisely that measurement. A SIRI-only punctuality figure would be measuring the thing we have already established is misleading.
If all we had were TransXChange
The key does not change. What changes is that TXC does not hand you two of the three — you walk to them.
| SIRI-VM | TransXChange | |
|---|---|---|
| Line | PublishedLineName |
Service/Lines/Line/LineName |
| Origin stop | OriginRef |
via the journey pattern, below |
| Departure time | OriginAimedDepartureTime |
VehicleJourney/DepartureTime |
The line is direct, with one trap: VehicleJourney/LineRef is a local id
pointing at the <Line> element, not the line name.
The origin stop is not on the journey at all:
VehicleJourney/JourneyPatternRef
→ JourneyPattern/JourneyPatternSectionRefs (several, in order)
→ JourneyPatternSection/JourneyPatternTimingLink[1]/From/StopPointRef
That StopPointRef is an ATCO code, the same namespace as SIRI's OriginRef,
so once walked to, the comparison is exact. Patterns spanning several sections
are normal, and taking the first section encountered rather than the first in
the ordered refs yields the wrong stop.
The departure time is local clock time — 07:50:00 — so SIRI's UTC is
converted to Europe/London and compared as HH:MM:SS, exactly as in the GTFS
path.
The fourth thing, which is where the work is
"Does this journey run today?" GTFS answers it with calendar.txt and
calendar_dates.txt. TXC makes you evaluate OperatingProfile per journey,
falling back to the Service's:
RegularDayType/DaysOfWeekBankHolidayOperation, days on and days offSpecialDaysOperationServicedOrganisationDayType— school terms, a dataset of its own
And above that, which registration applies. Several are published for one
line at once, differing only by OperatingPeriod/StartDate, usually with no
EndDate. Validity is "the latest start date not after the service date", not
a range test. Getting it wrong means matching against a timetable that is not
running.
And timing points have to be computed
GTFS stop_times carries an absolute time at every stop. TXC carries run
times: JourneyPatternTimingLink/RunTime as ISO durations (PT2M30S), plus
WaitTime, plus per-journey overrides in VehicleJourneyTimingLink. Every
intermediate scheduled time is DepartureTime plus a cumulative sum, and
timing-point status comes from TimingStatus on the stop usage rather than a
timepoint column.
So for punctuality specifically — measured at timing points, on departure — TXC-only means rebuilding the stop-times table before anything beyond the origin can be measured.
What it buys
frontend/txc_crosswalk.py already does this, producing
(operator, line, journey code, date) → (origin stop, departure time,
validity): a few hundred thousand rows, built offline, looked up in a
millisecond, rebuilt only when a registration changes.
TXC carries the operator's own vocabulary — TicketMachine/JourneyCode,
Block/BlockNumber, VehicleJourneyCode — all of which the conversion to GTFS
throws away. That is what closes the last gap in the scoring, and it is the
only bridge to franchised on-board equipment that reports a duty and nothing
else.
Two traps in the files themselves: OperatorRef on a journey is the file's
local operator id (FSY), not the national NOC that SIRI publishes (FSYO),
with the mapping in <Operators>; and the datasets are enormous — one First
Bus dataset is 19 MB zipped and 624 MB of XML across 75 files — so they have to
be streamed rather than parsed into memory.
Is the key always right?
It fails safe, which is the property that matters. Scored against the answer BODS itself publishes, over 2,561 attempts: 98.1% agreed, 0.0% wrong. The failure mode is "no match", not "wrong match" — and that is the direction to fail in, because a wrong journey id is written down as fact and nothing downstream can tell it was a guess.
That is not luck. It is HAVING COUNT(*) = 1 in the keying query: if line,
origin stop and departure time point at two journeys in the day, the match is
discarded rather than guessed.
Where it fails:
- The operator publishes no departure time. 92% of vehicles carry one; the rest can never match, however good the timetable. Not fixable from this end, which is why those operators are named on the published page.
- The key is not unique — two registered journeys, same line, same stop, same minute. Discarded. Costs coverage, buys correctness.
- The timetable and the day disagree — a short-notice change, or a registration newer than the weekly export. Nothing to match against.
Two numbers that are not the same measurement
The 98.1% was scored only over vehicles BODS had itself named a journey for: the fair test for "could we reconstruct it". The published page's matched to a journey figure is over every bus running, including operators who publish nothing matchable at all. Both are honest. Quoting one as the other is not.