Network Pulse

← All notes

Block chaining: considered, and the condition that would change the answer

Status: considered, not adopted — 2026-09-23. Nothing in the serving path does this. The measurements behind it are kept running so the decision can be revisited on evidence rather than re-argued.

A block, or running board, is a vehicle's day of work: the chain of journeys one bus is rostered to operate in turn. Block chaining means using it to say something about a journey that has not started — the bus due to work the 15:40 is currently finishing the 15:05 on the same block, eight minutes down, so the 15:40 will leave late and the board can say so before the bus arrives.

RTIG's guidance ranks this the best of three matching methods (RTIGT063, page 8, "BEST MATCH"). This note is why we do not use it, written down because "we considered it" is worth nothing without the reasons.

What it would buy

Coverage, and a lot of it. Almost every journey on a departure board has not started: measured across a full day, only 3–5% of journeys starting within the next 90 minutes have a bus visibly on them. A bus only takes on its next journey reference at the terminus, so it is invisible on the journey it is about to run until moments before it runs it.

Sampled by frontend/block_coverage.py (systemd block-coverage.timer):

2026-09-22 23:05 London journeys in 90 min board live now with chaining
everywhere the poller watches 1,351 8% 25%
South Yorkshire 113 10% 49%

South Yorkshire roughly doubles the national rate, and the reason is supply. Nationally 43% of trips carry a block, the split by operator bimodal rather than partial — Arriva, Go North East and the Bee Network agencies publish none at all. Across trips calling in South Yorkshire it is 96%: First South Yorkshire, Stagecoach Yorkshire and TM Travel publish on every trip, and what is left without one is FlixBus, National Express and a handful of coach and community operators.

So in the area the contract is about, the data is there.

Why we do not use it

It is not needed for matching. BODS resolves the journey before we see it and publishes a GTFS trip_id in the real-time feed. That id is correct by definition; anything we reconstruct is inference. Scored against BODS's own answers, running board and clock agreed on 84.8% where the four-field key reached 99.0% — so adopting it for identification would mean replacing a known-correct field with a worse guess at the same thing.

It would not make predictions better. Block chaining was pursued for months on the basis that the incumbent predicted long-lead departures far better than us and that blocks were the reason. Re-scored on 73,251 predictions instead of 239, and on identical departures rather than each system's own chosen subset, the gap is 0–9 seconds per lead bucket and a dead heat at two of five. There is no accuracy gap for blocks to close. Emitting the scheduled time instead of extrapolating was tested at the same time and is worse at every lead under 40 minutes.

It is the only method with a material wrong-answer rate: 9.6%. Every other strategy scored produced almost none. The reason is not implementation and cannot be engineered away — a block is a plan. Buses are swapped between boards, run short and get reallocated during the day, and the registration cannot know. Two rounds of work moved it from 83.1% to 84.8% and no further; what is left is the road disagreeing with the plan. On a departure board a confident wrong time is worse than an honest scheduled one, because nothing downstream can tell it was a guess.

The two feeds do not share a block identifier. RTIG's diagram shows <BlockNumber>57839</BlockNumber> matching <BlockRef>57839</BlockRef>, which holds only if the timetable side is TransXChange. Take the timetable from BODS's GTFS conversion, as we do, and block_id is a 40-character content hash while SIRI sends the operator's own number. Measured: 1,604 distinct BlockRef values live, 0 matching any GTFS block_id.

That last one does not block chaining itself — chaining runs entirely inside our own GTFS import, trip to block to next trip, and the hash only has to be internally consistent. It matters for anything that starts from the AVL feed.

The condition that changes the answer

If BODS is not the source — if the timetable is not BODS GTFS and the positions come from SIRI taken direct from the bus — every reason above except the 9.6% falls away, and block chaining stops being optional.

That is not hypothetical. It is what franchising brings: on-board equipment reports the operator's own identity for the work it is doing, knows nothing of GTFS, and there is no trip_id to inherit. In that world:

  • There is no published match to be better than. The comparison is no longer "block chaining against a correct id", it is "block chaining against nothing".
  • Blocks are what the equipment has. SIRI-VM already carries BlockRef on about 79% of vehicles where GTFS-RT carries no block at all, and franchised equipment reporting a duty is the same shape of data.
  • 84.8% becomes the floor rather than the ceiling. Running board and clock is the only route that equipment can offer unaided, and 85% of journeys identified is a great deal better than none.

The lever, and it is a procurement one

The 9.6% is not a fact about blocks, it is a fact about having only a block. Scored against the same answer key: specifying that on-board equipment report the journey it is working, or that journey's due time — not merely the duty it is signed on to — keeps matching in the high nineties. Taking the duty alone accepts 85% and a one-in-ten chance of naming the wrong journey.

That is a line in a specification, it costs nothing to ask for at the point the equipment is being bought, and it cannot be retrofitted by any amount of work at this end. It is the single most valuable thing on this page.

What is kept running, and why

Neither of these serves anything. They exist so that the decision above can be re-taken on current evidence rather than re-argued from first principles.

  • siri_poller.py banks BlockRef, DatedVehicleJourneyRef, the ticket machine's JourneyCode and DriverRef into siri_vehicles. None of it can be backfilled — a feed not read is a day not recorded — and it is the only place the operator's own journey vocabulary is visible.
  • block_coverage.py samples the ceiling through the day, nationally and for South Yorkshire, so the coverage figures above stay current.

If it is ever built

Two rules, both from the measurements rather than from taste.

Attribute only from a bus actually working the preceding journey on the block, near enough its end to plausibly continue — never from a bus merely rostered to the block. That is the difference between reading the road and reading the plan.

Say which it is. A departure predicted from a bus on a different journey is a weaker claim than one predicted from the bus itself, and a board that shows both identically is overstating what it knows. The honest presentation of a chained prediction is not a time, it is a time with its reason attached.