Network Pulse

← All notes

Passenger counts: what the open data cannot say, and what an APC feed would have to carry

Status: assessed, not procured. The absence in BODS was measured GB-wide on 2026-08-12; the supplier side is one sample from SURE Solutions -- one vehicle, one day, 168 rows -- read on 2026-08-24 against the national stop register.

Every question a network planner actually asks is a question about people. Which journeys are full. Which stop the load builds at. What a ten-minute delay costs in passenger-hours rather than bus-hours. Whether the buses that were cut were the empty ones.

None of those can be answered from the open data, and the reason is not that operators are withholding it. It is that the standards BODS publishes cannot express a passenger count at all. So the count has to be bought, and the interesting work is not the buying -- it is the joining.

The gap, measured

Checked across the whole of Great Britain on 2026-08-12, from the SIRI bulk archive (26,429 vehicles in one file) and a GB-wide GTFS-RT bounding box.

Feed Field that could carry a count What is actually there
SIRI-VM none exists 43 distinct elements in the whole national file. Zero PassengerCount. <Extensions> only ever holds ticket-machine fields, VehicleUniqueId and DriverRef.
SIRI-VM <Occupancy> 361 of 26,429 vehicles (1.4%). Three permitted values -- full, seatsAvailable, standingAvailable -- and standingAvailable was never observed, so it is binary in practice.
GTFS-RT occupancy_percentage 0 vehicles.
GTFS-RT multi_carriage_details 0 vehicles.
GTFS-RT occupancy_status 360 of 26,306, of which 359 say MANY_SEATS_AVAILABLE.
Static GTFS trips.wheelchair_accessible 0 populated of 129,641.
Static GTFS stops.wheelchair_boarding 0 populated of 32,088.

The 1.4% that do report Occupancy are reporting honestly -- vehicles were observed changing value in directions that read as real loading over a peak -- so the limits are coverage and coarseness, not truthfulness. But no numeric count is expressible in SIRI 2.0 anywhere, which means no amount of operator goodwill would put one in BODS. Wheelchair and pushchair fields are the same story from the other direction: the schema has them, nobody populates them, and in any case they describe a passenger's requirements or a vehicle's fixed specification, never whether the bay is free right now.

So: automatic passenger counting is a purchase, not an integration.

What a supplier feed actually looks like

The SURE Solutions sample is a flat event file, one row per door-opening, 21 columns:

BUS_NAME  FLEET  TRIP_ID  EVENT_DATE  STOP_TIME_FROM  STOP_TIME_TO
STOP_NAME  ROUTE_NAME  People_In  People_Out  UPPER_DECK_IN  UPPER_DECK_OUT
WHEELCHAIR_IN  WHEELCHAIR_OUT  PRAM_IN  PRAM_OUT  OCCUPANCY  LATITUDE  LONGITUDE

One vehicle (OXBC-789), one day, eleven trips plus a depot row, 539 boardings and 539 alightings. Peak occupancy 65, at 17:43 inbound on St Clements Street.

The counting looks good. Over 539 boardings the running total ends the day at exactly zero and never drifts further than three passengers from what the boardings and alightings imply. If that holds across a fleet it is a better sensor than anything in the open data by an enormous margin. Every problem below is about metadata and attribution, not about the counter.

It is worth noting that the operator in this sample is one of the nine already publishing the coarse Occupancy enum to BODS. That makes them an unusually good validation partner: their own three-value enum can be checked against their own numeric count, on the same vehicles, on the same days.

What the file cannot be joined to

Our reporting is built on stop identifiers and timetabled journeys. This file has neither.

What a join needs What the sample has
Stop identity (ATCO / NaPTAN) Absent. A free-text name, plus the position of the bus.
Operator (NOC) Only as a prefix on the vehicle name, OXBC-789. FLEET is "Oxford" -- a depot, not an operator code.
Line identity ROUTE_NAME is 400 -- the number on the front of the bus, not a LineRef or a registered service code.
Direction Absent entirely.
Journey identity TRIP_ID runs 1-11 and resets daily. No relation to a journey code, a block, or a DatedVehicleJourneyRef.
Scheduled time Absent. Observed times only, so nothing can be called early or late.
Vehicle identity that matches AVL OXBC-789 is a fleet number, not the VehicleRef the live feed uses.
Capacity Absent -- so an occupancy of 65 cannot become a load factor.
Timestamps 06/03/2026 is ambiguous between day-first and month-first; times are 12-hour with AM/PM, no date, no timezone. Anything running past midnight sorts wrongly.

Names are not identifiers, and this is the expensive one

Matched against the national stop register, the 34 distinct STOP_NAME values in this file touch 57 distinct physical stops. Eighteen of the 34 names -- more than half -- cover more than one pole:

  • Oxford City Centre, Westgate covers five ATCO codes spread over 335 m.
  • Headington, Brookes University covers three, and the reported positions under that one name span 176 m.
  • Green Road Roundabout and Headington, Green Road Roundabout are the same place under two names. So are High Street, Oxford City Centre, High Street and Oxford High Street (Stop T1).
  • Oxford City Centre, Queens Lane and Oxford City Centre, Queen's Lane West differ by an apostrophe and a suffix.
  • One row is called UNKNOWN.

Above all, the name does not distinguish direction. The two poles either side of a road share a name, and the register gives them different codes precisely because they are different stops. Joining on name silently merges inbound with outbound, which is exactly the distinction a load profile is made of.

The file is a record of doors opening, not of journeys

A row exists only where somebody got on or off. The six inbound trips in this sample produce between 13 and 20 rows each, against a union of 29 distinct stops, and only four stops appear on all six.

That is not a fault -- it is what an event file is. But it means a missing stop is indistinguishable from a stop where nobody boarded, and neither is distinguishable from a stop the bus never served that trip. A load profile built straight off these rows has holes in it that look like zeroes.

Four defects in the data itself

1. Occupancy is a running total that never resets. OCCUPANCY is exactly the cumulative sum of People_In minus People_Out -- across all 168 rows it deviates from that not once. It is not reset at a terminus, at a layover, or overnight. So the day's small residual error accumulates: 16 rows report a negative number of passengers on board, down to -3.

2. The trip boundary is drawn in the wrong place. It falls where the bus moves off, not where the journey ends. The last row of trip 1 is the railway station at 09:17 with 8 boardings and no alightings -- those eight people are boarding trip 2. Worse, trip 3 ends at 11:16 with in 15, out 11: the arriving journey's alightings and the departing journey's boardings are added together in a single row and cannot be separated afterwards. This is also where the negative occupancy comes from.

3. Upper-deck counts are a stair sensor, not a subset. On seven rows the deck count exceeds the door count it should sit inside -- and every one of those rows has People_In and People_Out both zero, with an upper-deck alighting recorded anyway. Read literally, someone left the top deck while nobody left the bus. Read correctly, the sensor is on the staircase, so a passenger who comes downstairs one stop before they alight is counted at the wrong stop. The columns are usable as a deck-split ratio over a journey; they are not usable per stop, and they must never be summed with the door counts.

4. The accessibility columns are empty. WHEELCHAIR_IN, WHEELCHAIR_OUT, PRAM_IN, PRAM_OUT are zero on all 168 rows. That may be a true zero for one vehicle-day, but it is exactly the trap BODS sets: a column existing is not the same as a column being populated, and it is the mistake that would be easiest to make twice. Any contract has to specify populated, with an agreed detection method, not present.

Two smaller ones. Thirty rows have STOP_TIME_FROM equal to STOP_TIME_TO, and 25 of those had passengers move -- a zero-second dwell during which someone boarded. So these timestamps are event marks, not door open/close times, and should not be used as a dwell-time source. And at timing points the window covers the whole layover: 6.5 minutes at Thornhill Park and Ride, 8.6 minutes at Waterstock. Those are scheduled waits, not boarding time.

What we can fix ourselves

Most of it, as it happens -- because we already hold the national stop register and a timetable, and because we already solve the same problem for live vehicle positions.

Snapping to real stops works. Every row carries the bus's own position to eight decimal places. Matched against the 1,060 active bus stops in the register within the area of this route:

rows
An active bus stop within 25 m 161 of 168 (96%)
Within 100 m 165 of 168 (98%)
No stop within 100 m 3 -- two depot rows, which are correctly not stops, and the row named UNKNOWN

But position alone is not enough, and this is the trap. Only 94 of 168 rows (56%) have exactly one candidate stop within 25 m. The rest have two or more, because the opposite pole is usually across a road that is narrower than the fix is accurate. Adding the direction of travel -- derived from the positions either side of each row -- resolves most of the remainder, leaving eight ambiguous groups, all of them multi-stand city-centre locations where the compass bearing in the register does not separate the stands either.

So the rule is the same one that already governs live vehicle matching: match the trip, not the row. A sequence of positions in time order fits one journey pattern far more decisively than any single position fits one stop, and once the pattern is chosen, every stop in it is named. The last few percent come free.

The rest is arithmetic on top of that:

  • Recut the journey boundary at the timetabled journey start rather than at the moment of departure, and split a terminus row's boardings from its alightings by which journey each belongs to.
  • Reset occupancy to zero at each journey start. Negative passenger counts disappear, and the residual error becomes a per-journey quality measure instead of a debt carried all day.
  • Fill the gaps from the timetable. Once the journey is identified, its full stop sequence is known, and a stop with no row becomes an explicit zero rather than a hole -- with the honest exception of stops where the vehicle cannot be shown to have passed.

What to ask the supplier for

Everything above is recoverable, but recovering it costs accuracy we would not have to lose. The following are all things an APC system already knows internally, and asking is much cheaper than inferring:

  1. The ATCO code for each event. This is the single highest-value field in the list; it removes the entire snapping problem.
  2. Operator NOC, line reference and direction as their own columns.
  3. A journey identifier that matches the registered timetable -- a journey code or, failing that, the scheduled departure time of the journey from its origin. Without one of these, punctuality and loading can never be put on the same row.
  4. The vehicle identifier used in the AVL feed, so counts and positions join without a hand-maintained fleet mapping.
  5. Full ISO 8601 timestamps with offset -- 2026-03-06T08:21:17+00:00. This also settles whether 06/03/2026 is March or June.
  6. Vehicle capacity, seated and total, by deck. Without it there is no load factor, only a count.
  7. A journey-scoped occupancy -- reset at each journey start -- or at minimum a flag marking the terminus row, so the boundary can be cut cleanly.
  8. Confirmation that the accessibility columns are populated, with how the detection works and what it does and does not catch.

Items 1, 3, 5 and 6 are the ones that decide whether this data is worth having. The others are convenience.

What it would buy

The value is not the count on its own -- it is that every measure already on this site is currently weighted by buses, and should be weighted by people:

  • Pinch points score a link by slowdown, variability and daily buses. With loadings, "daily buses" becomes daily passengers, and the ranking changes from where buses lose time to where people lose time. That is the number a highway authority can build a business case on.
  • Punctuality currently treats a full peak departure and an empty Sunday one as one observation each.
  • Lost mileage becomes lost passenger journeys, which is what a lost journey actually costs.
  • First and last journey compliance gains a way to say whether the protected journeys are the used ones.

How to test it before committing

One vehicle for one day is enough to judge the shape of a feed, which is what this note has done. It is not remotely enough to judge its accuracy, and the figures above should not be read as if it were -- a single curated sample day that balances to zero could balance to zero for reasons that have nothing to do with the sensor.

A proof of concept should ask for four things:

  • A full week, several vehicles, including at least one double-decker and one single-decker, so the deck sensors can be judged separately.
  • Days we can independently check -- a day with a manual count, or with ticket-machine boardings for the same journeys. Ticketing gives boardings only, and misses concessionary and through passengers differently, so it is a cross-check rather than a truth; but a systematic gap between the two is visible immediately.
  • Vehicles already reporting Occupancy to BODS, so the coarse public enum and the numeric count can be compared directly on the same journeys. If the count says 65 on a vehicle publishing seatsAvailable, one of them is wrong and we would want to know which before either goes on a page.
  • The join tested first, on real days. The measure of success is not "the numbers look plausible" but "what share of rows resolved to a stop and a timetabled journey" -- and per the numbers above, a naive join would quietly succeed at 96% and be wrong about direction on half of them.

The rule this site already applies holds here too: an unlit lane is drawn unlit. A load profile with reconstructed gaps must say which stops were observed and which were inferred, on the page, every time.