Network Pulse

← All notes

Which fields the match is made on, and what happens without each source

Status: measured, written 2026-09-08. The three fields that join a live bus to a scheduled journey, where they are in SIRI-VM, where the same three things live in TransXChange, and what is still measurable if either source goes away.

Companions rather than repeats: how we know which journey a bus is on is the plain-English version, matching without BODS is the scoring rig and its numbers, and what comes out of BODS is the field-level reference for the feed as a whole.

The key

line  +  the stop it started from  +  the time it was due away there

resolved against the service calendar for that day. Three fields. Everything downstream — punctuality, lost mileage, did-it-run, the delay on the map — is about a scheduled journey, and this is the only thing connecting a dot on a map to one.

Where they are in SIRI-VM

A real record from the archive, 8 September 2026, trimmed only of whitespace:

<VehicleActivity>
  <RecordedAtTime>2026-09-08T15:47:15+00:00</RecordedAtTime>
  <MonitoredVehicleJourney>
    <LineRef>7</LineRef>
    <DirectionRef>outbound</DirectionRef>
    <FramedVehicleJourneyRef>
      <DataFrameRef>2026-09-08</DataFrameRef>
      <DatedVehicleJourneyRef>3049</DatedVehicleJourneyRef>
    </FramedVehicleJourneyRef>
    <PublishedLineName>7</PublishedLineName>          <!-- 1 -->
    <OperatorRef>SYRK</OperatorRef>
    <OriginRef>370021636</OriginRef>                  <!-- 2 -->
    <OriginName>Crystal Peaks CP5</OriginName>
    <DestinationRef>370023381</DestinationRef>
    <DestinationName>Wordsworth Avenue</DestinationName>
    <OriginAimedDepartureTime>2026-09-08T14:48:00+00:00</OriginAimedDepartureTime>  <!-- 3 -->
    <Monitored>true</Monitored>
    <VehicleLocation>
      <Longitude>-1.478351</Longitude>
      <Latitude>53.393391</Latitude>
    </VehicleLocation>
    <Bearing>324</Bearing>
    <VehicleRef>SYRK-11702</VehicleRef>
  </MonitoredVehicleJourney>
  <Extensions><VehicleJourney><DriverRef>79023</DriverRef></VehicleJourney></Extensions>
</VehicleActivity>
SIRI-VM field Stored as Joins to, in GTFS
1 PublishedLineName, falling back to LineRef line routes.short_name
2 OriginRef origin_ref stops.stop_id — an ATCO code, so it joins straight in
3 OriginAimedDepartureTime origin_dep_local stop_times.departure_time at the journey's first stop

All three sit inside MonitoredVehicleJourney. The poller reads them as .//PublishedLineName, .//OriginRef and .//OriginAimedDepartureTime.

Two things about field 3 decide whether the join works at all:

  • It is offset-stamped UTC and GTFS is local service time. 14:48:00+00:00 against a departure_time of 15:48:00. Compare them raw and every BST journey is an hour out, which is why it is stored converted: format(od, "%H:%M:%S", tz = "Europe/London").
  • It is the field with the coverage problem. Present on 846 of 915 vehicles (92%) in one poll of the South Yorkshire box on 8 September 2026, and absent entirely from several operators, whose buses can therefore never be matched.

The trap in the same record

DatedVehicleJourneyRef is 3049. It is published on 100% of vehicles, it looks exactly like a journey identifier, and it is the operator's own ticket-machine code. Joined to GTFS trip_id it matches 0%. Everyone reaching into this XML for the first time reaches for it.

Does any of this need GTFS-RT?

No. The proof of concept reads SIRI-VM and a timetable, and nothing else. The trip_id dependency described in matching without BODS belongs to the live journey planner, which takes GTFS-RT and looks the id up. The proof of concept has never had that field and was built without it.

What it does need is a timetable — and that is a "you must know what was supposed to happen" dependency rather than a BODS one. Any schedule source answers it.

If all we had were raw SIRI-VM

The feed carries two scheduled times of its own. From that same poll of 915 vehicles:

Field Populated
OriginAimedDepartureTime 846 (92%)
DestinationAimedArrivalTime 554 (61%)

Still measurable, with no timetable at all:

  • Punctuality at the origin, against the bus's own aimed departure. Stop coordinates would come from NaPTAN — a separate national dataset, not the BODS timetable.
  • Lateness at the destination, for the 61% publishing an aimed arrival.
  • Headways and excess waiting time. No schedule needed, and the right measure on high-frequency corridors anyway.
  • Journey times between two points, speeds, coverage, vehicle counts, route deviation.

No longer possible:

  • Did the service run, and lost mileage. Both need a denominator of what should have run, and no live feed can supply it — you cannot see a bus that never turned up.
  • Punctuality at intermediate timing points, which is the measure official punctuality uses.

There is a sharper objection to the origin-only version. This method deliberately excludes the first and last stop of every journey, because a bus sitting at a terminus reads as early or late without saying anything about the service. Origin-aimed-departure punctuality is precisely that measurement. A SIRI-only punctuality figure would be measuring the thing we have already established is misleading.

If all we had were TransXChange

The key does not change. What changes is that TXC does not hand you two of the three — you walk to them.

SIRI-VM TransXChange
Line PublishedLineName Service/Lines/Line/LineName
Origin stop OriginRef via the journey pattern, below
Departure time OriginAimedDepartureTime VehicleJourney/DepartureTime

The line is direct, with one trap: VehicleJourney/LineRef is a local id pointing at the <Line> element, not the line name.

The origin stop is not on the journey at all:

VehicleJourney/JourneyPatternRef
  → JourneyPattern/JourneyPatternSectionRefs          (several, in order)
    → JourneyPatternSection/JourneyPatternTimingLink[1]/From/StopPointRef

That StopPointRef is an ATCO code, the same namespace as SIRI's OriginRef, so once walked to, the comparison is exact. Patterns spanning several sections are normal, and taking the first section encountered rather than the first in the ordered refs yields the wrong stop.

The departure time is local clock time — 07:50:00 — so SIRI's UTC is converted to Europe/London and compared as HH:MM:SS, exactly as in the GTFS path.

The fourth thing, which is where the work is

"Does this journey run today?" GTFS answers it with calendar.txt and calendar_dates.txt. TXC makes you evaluate OperatingProfile per journey, falling back to the Service's:

  • RegularDayType/DaysOfWeek
  • BankHolidayOperation, days on and days off
  • SpecialDaysOperation
  • ServicedOrganisationDayType — school terms, a dataset of its own

And above that, which registration applies. Several are published for one line at once, differing only by OperatingPeriod/StartDate, usually with no EndDate. Validity is "the latest start date not after the service date", not a range test. Getting it wrong means matching against a timetable that is not running.

And timing points have to be computed

GTFS stop_times carries an absolute time at every stop. TXC carries run times: JourneyPatternTimingLink/RunTime as ISO durations (PT2M30S), plus WaitTime, plus per-journey overrides in VehicleJourneyTimingLink. Every intermediate scheduled time is DepartureTime plus a cumulative sum, and timing-point status comes from TimingStatus on the stop usage rather than a timepoint column.

So for punctuality specifically — measured at timing points, on departure — TXC-only means rebuilding the stop-times table before anything beyond the origin can be measured.

What it buys

frontend/txc_crosswalk.py already does this, producing (operator, line, journey code, date) → (origin stop, departure time, validity): a few hundred thousand rows, built offline, looked up in a millisecond, rebuilt only when a registration changes.

TXC carries the operator's own vocabulary — TicketMachine/JourneyCode, Block/BlockNumber, VehicleJourneyCode — all of which the conversion to GTFS throws away. That is what closes the last gap in the scoring, and it is the only bridge to franchised on-board equipment that reports a duty and nothing else.

Two traps in the files themselves: OperatorRef on a journey is the file's local operator id (FSY), not the national NOC that SIRI publishes (FSYO), with the mapping in <Operators>; and the datasets are enormous — one First Bus dataset is 19 MB zipped and 624 MB of XML across 75 files — so they have to be streamed rather than parsed into memory.

Is the key always right?

It fails safe, which is the property that matters. Scored against the answer BODS itself publishes, over 2,561 attempts: 98.1% agreed, 0.0% wrong. The failure mode is "no match", not "wrong match" — and that is the direction to fail in, because a wrong journey id is written down as fact and nothing downstream can tell it was a guess.

That is not luck. It is HAVING COUNT(*) = 1 in the keying query: if line, origin stop and departure time point at two journeys in the day, the match is discarded rather than guessed.

Where it fails:

  1. The operator publishes no departure time. 92% of vehicles carry one; the rest can never match, however good the timetable. Not fixable from this end, which is why those operators are named on the published page.
  2. The key is not unique — two registered journeys, same line, same stop, same minute. Discarded. Costs coverage, buys correctness.
  3. The timetable and the day disagree — a short-notice change, or a registration newer than the weekly export. Nothing to match against.

Two numbers that are not the same measurement

The 98.1% was scored only over vehicles BODS had itself named a journey for: the fair test for "could we reconstruct it". The published page's matched to a journey figure is over every bus running, including operators who publish nothing matchable at all. Both are honest. Quoting one as the other is not.