Network Pulse

← All notes

How we know which journey a bus is on

Status: built and measured, not serving anything yet. The live system still uses the journey id BODS publishes; this is what would replace it. Figures from South and West Yorkshire, 2026-08-27.

The short version

A bus reports where it is, and it reports what it is doing — but it says what it is doing in the operator's own words:

First South Yorkshire, vehicle 35126, route 97, running board 09703, journey 1307.

The timetable contains none of those words. What the timetable has is:

a journey on route 97 that leaves the stop outside the Northern General at 13:07, calling at 42 stops, arriving 13:54.

Both describe the same bus. Nothing links them. Matching is translating between the two, and it matters because every number we report — was the bus on time, did the journey run, how much mileage was lost — is about a scheduled journey, not about a dot on a map.

The translator is the operator's registration, the legal document they file before running a service. It is written in TransXChange, and it is the one place where "journey 1307" and "leaves at 13:07 from that stop" appear on the same line. So we read every registration, build a dictionary from it, and look the bus up.

The three ways to look a bus up

The same real bus, on 27 August, found three different ways:

1. By the journey's name. The bus says it is on journey 1307. The registration says journey 1307 on route 97 is the 13:07. Done — this is the direct route, and it needs nothing but the operator's own code.

2. By the vehicle's day's work. The bus says it is on running board 09703. A running board is one vehicle's whole day: out at 10:11, back at 11:36, out again at 13:07, and so on until it goes to the depot. That is twelve journeys, not one — so the clock chooses. At 13:20 the bus is on the 13:07.

3. By describing it. Route 97, from that stop, at 13:07. No codes at all, just what the journey is.

All three land on the same journey. They are not spare tyres for each other: each fails in a different way, and no single one is published by every operator.

Which to use is a measured question, not a matter of taste. Against the answer BODS publishes for the same bus at the same moment:

way in gets it right gets it wrong can't tell
3 — describing it 98.1% 0.0% 1.9%
1 — the journey's name 82.8% 0.4% 16.8%
3, then 1 where 3 is silent 99.0% 0.0% 1.0%

Way 2, the running board, is scored on its own because it is answering a harder question with less to go on. Where we hold the operator's registration, the board names a journey 91% of the time, and that journey is the right one 85% of the time. That is a long way short of 99%, and the gap is the point of the next section.

Why the running board matters more than it looks

Everything above works because SIRI-VM publishes a departure time. Franchised on-board equipment will not. A ticket machine knows the duty the driver signed on to; it has never heard of a GTFS journey id, and it does not know what time the journey was due — only what board it is working.

So the running board is not the weakest of the three. It is the one that will still be there when the others are gone, which is why it is worth measuring now rather than discovering its limits after the equipment is fitted.

And what the measurement says is uncomfortable: a running board and a clock identify the journey correctly about 85% of the time, against 99% for the routes available today. Not because the matching is poor, but because a board is a plan and the road is not — buses get swapped between boards, run short, and are reallocated during the day, and no amount of reading the registration more carefully will see that.

That makes it a procurement question rather than an engineering one. If franchised equipment reports the journey it is working, or the time that journey was due, and not only the duty it signed on to, the matching stays in the high nineties. If it reports the duty alone, roughly one journey in seven will be attributed to the wrong one. That is worth specifying before the first vehicle is fitted, because it cannot be recovered afterwards.

The limitations

Read this part. The matching is good enough to build on and it is not good enough to trust blindly.

Not every bus says everything. About 30% publish no running board at all, and 2% publish no departure time. Where a bus says nothing that identifies the journey, nothing can be recovered — there is no cleverness available.

A code is only unique inside one registration. Several registrations for a route are live at once — the one running now and the one starting next month. First South Yorkshire's 0758C on the X2 leaves at 07:51 in one and 07:58 in the other. Both are real. The code alone cannot choose between them.

A code looks like a time and often isn't one. Most operators use the departure time — 1307 above. Arriva Yorkshire published 2022 for a journey leaving 09:55; TransPennine 5413 for one leaving 08:07. Reading the code as a clock works fine until it puts a bus twelve hours out.

There are two journey codes in every record and they disagree. They agreed on 81 of First South Yorkshire's 259 buses and on 1 of Arriva's 186. Only one of them is the code the registration uses: matched through the dictionary they score 82.8% and 8.3%. Picking the wrong field is the single most expensive mistake available here, and both look equally plausible.

A running board is a plan, and the road is not. Buses are swapped between boards, run short and are reallocated during the day. The registration cannot know, so a board that says "this vehicle works the 13:07" is describing an intention. This is the floor under the 85%, and it is not a software problem.

A running board says what the day's work is, not where in it the bus is. The registration gives a departure and how long the journey takes; without adding those up, "which journey is this board on now" answers "the last one that departed" — which puts a bus resting between journeys onto the one it has just finished, and drags a late bus forward onto one it has not started. Those two accounted for 89% of the wrong answers this route gave before journey lengths were imported.

We only hold registrations we have downloaded. Of the journey-code route's misses, 13.3% were buses crossing into Sheffield on West Yorkshire registrations we had not fetched. That is a scope decision, not a limit of the method — but it looks identical to a failure until you go and check.

It can be wrong quietly. This is the important one. Today BODS tells us the journey and cannot really be wrong: the id it publishes is the id it means. A translation can land on a real, plausible, different journey, and nothing downstream can tell that it was a guess. That is why the tables above separate "gets it wrong" from "can't tell": a gap is visible and costs coverage, while a wrong answer is invisible and corrupts a figure someone reports.

And the timetable underneath has to be current. None of this beats the oldest problem: if the schedule we hold is a fortnight stale, the journey the bus is running may not be in it at all. That limit is shared by every method, including the one running today.

Where it runs

txc_crosswalk.py builds the dictionary from the registrations, weekly, as a step in the network refresh. siri_poller.py reads the live feed so running boards and journey codes are recorded rather than merely available. compare_matching.py scores every route against BODS's own answer, which is how every figure on this page was produced. None of it is on the path that serves a passenger or produces a reported number.

The detail — field names, formats, the timezone traps, and how the scoring was done — is in matching without BODS, and in the comments of those three files.