Network Pulse

← All notes

Bus open data: how the proof of concept works

Status: plain description, written 2026-09-08. Where the bus data comes from, what happens to it, and what comes out at the end. No jargon, and nothing here needs the code to be understood.

Download this as a Word document — the same explanation, for sending on or printing.

The technical notes are elsewhere: what comes out of BODS and what it turns into is the field-by-field reference, and running the proof of concept is the operating manual.

The shape of it, in one line

BODS → laptop → database → CSV files → web pages

Everything below is that sentence in more detail. Five steps, each of which can be re-run without redoing the ones before it.

1. Where the data comes from

Everything starts at the Bus Open Data Service, the national service every operator in England must publish to. We take two things from it:

  • Live vehicle positions. Where every bus is right now, as a snapshot we ask for every 20 seconds. About 895 buses answer inside the South Yorkshire area at the moment.
  • The timetable. What was supposed to run today: every journey, every stop, every scheduled time. A weekly file.

What matters most is what the live feed does not contain. It does not say whether a bus is late. It does not say which stop a bus is heading for. It carries no passenger numbers, and its accessibility fields are empty. Every figure we produce about punctuality is worked out by comparing positions against the timetable ourselves.

2. Onto the laptop

A small program asks BODS for the South Yorkshire area every 20 seconds and saves the answer to disk exactly as it arrived — roughly 4,300 files a day, about a third of a gigabyte compressed.

This is the only step that cannot be repeated. BODS serves the present moment and keeps no history of its own: a minute we fail to capture is gone for good, at any price. Everything after this point reads what was saved, and can be re-run over any past day as often as we like.

Because a laptop is closed at night and reboots for updates, a second copy of the same feed is now kept on our own server, which runs continuously. It holds 14 days, and the laptop can pull back any minutes it missed.

3. Into a database

The saved files are loaded into DuckDB, a database that runs inside the laptop with nothing to install and no server behind it. Each poll becomes one row per bus: where it was, when, which route, which operator, and the handful of other fields the feed carries.

4. Combining the live data with the timetable

This is the hard part, and it is where most of the work has gone.

The live feed never names the journey a bus is running — there is no identifier we can look up. So each bus is matched to a timetabled journey by three things together: the route number, the stop it started from, and the time it was scheduled to leave there. When those three point at exactly one journey in the timetable, we know what that bus was supposed to be doing, and every measure follows from that.

  • Most buses match — typically around nine in ten of those running.
  • Some can never match. A number of operators do not send a scheduled departure time at all. Their buses appear on the map and in no punctuality figure, and the page names them rather than quietly leaving them out.

One bus, all the way through

A real vehicle, on the afternoon of 8 September 2026.

What the feed said. Stagecoach Yorkshire, vehicle 11702, on line 7, starting from stop 370023381 — High Street/Wordsworth Avenue — scheduled to leave there at 14:11. Those three in bold are the key. Everything else in the record is the position itself: 198 separate fixes between 14:07 and 15:14.

What the timetable said. Exactly one journey in the day fits that combination: line 7 leaving High Street/Wordsworth Avenue at 14:11, seventy-five stops, due into Crystal Peaks at 15:27. One match, so we now know what this bus was supposed to be doing all afternoon — and the feed never told us.

What we then measured. Every time the bus is seen at a timing point on that journey, one row is written: where it was due, when it was actually there, and the difference.

Timing point Due Seen Delay
Monteney Road/Monteney Crescent 14:14 14:14:31 +0.5 min
Halifax Road/Southey Green Road 14:25 14:25:35 +0.6 min
Penistone Road/Beulah Road 14:31 14:31:18 +0.3 min
Sheffield Interchange/A4 14:48 14:50:10 +2.2 min
City Road/Elm Tree 15:04 15:09:04 +5.1 min

Those five rows are five of the tens of thousands that make up a day, and every punctuality figure on the page is counted from rows like them. Read on its own, this one says something a percentage never could: the bus kept time across the north of the city and lost five minutes getting through the centre.

The exact fields this is done with — where those three things sit in the live feed, what happens without the timetable, and how the same join would be made from TransXChange instead — are in which fields the match is made on.

Two details that decide whether such a row is fair. The time taken is the last position within 100 metres of the stop, not the nearest one — a bus that arrives early and waits is keeping time, and it is departure that the standard measures. And the first and last stop of a journey are excluded entirely, because a bus sitting at a terminus reads as early or late without saying anything about the service.

5. What comes out: the measurement files

Every minute, the combined data is written out as three plain CSV files that anyone can open in Excel or point Power BI at:

File One row per
positions.csv bus running now, with how late it is
observations.csv departure measured at a timing point today — what the punctuality figures are counted from
delivery.csv scheduled journey today, whether we saw it run, and how many miles it was

6. The pages people read

Those files are turned into web pages and dropped into a shared OneDrive folder every couple of minutes. Colleagues open them by double-clicking: no server to log into, no licence to buy, nothing for IT to install — the data is inside the page itself.

The main page answers, in order:

  1. How out of date is this, and when does it refresh
  2. How many buses are running right now, and how many we can measure
  3. Where they are, and where the delay is, on a map
  4. Punctuality today — on time, early, late
  5. Did the scheduled service actually run
  6. Lost mileage — scheduled miles that did not run, with the arithmetic shown
  7. Average speed
  8. Which operators to count — tick boxes that move every figure on the page
  9. Data quality — what is wrong with the data behind all of the above

A second page gives a departure board for any stop.

What this can and cannot tell us

It can:

  • Whether the scheduled service ran, and what mileage was lost when it did not
  • How punctual it was, measured at timing points on departure
  • Which routes and which operators are worst affected, and where
  • How complete and how honest the underlying data is

It cannot:

  • Count passengers — nothing in the national data carries them
  • Say anything about wheelchair spaces or vehicle accessibility — those fields exist in the standard and are unpopulated
  • Measure operators who do not publish the fields needed to match a journey
  • Reach back before we started capturing — there is no history to request

What it would take to do this properly

Nothing here needs new data. It needs somewhere better than a laptop to run:

  • Live rather than a snapshot. The page is as current as the last publish, a couple of minutes old. On a server with a database behind it, a page would be as current as the last thing BODS published — about twenty seconds.
  • Capture that never stops. A laptop sleeps, travels and reboots. A server does not.
  • The whole region, and more history. The area covered is currently a rectangle sized to fit one machine.
  • Nobody depending on one person being logged in.

The proof of concept has already answered the question it was built for: the national open data is good enough to measure the network with, provided its gaps are stated rather than hidden.