Network Pulse

← All notes

Where the network loses time: pinch points, and how they are scored

Status: built and running, on 17 days of observations (2026-08-04 onwards) covering South Yorkshire. Report at /pinch-points; framed on /platform.

Punctuality answers an operator's question — was the bus late? This answers a highway authority's — where is the network losing the time, and what is it costing? Same events, different reader, and the second one can act on roads rather than on schedules.

What a pinch point is here

A link: one stop to the next. Its run time is measured from the observed departure at the first stop to the observed arrival at the second, so the wait at either end is excluded and what is left is time on the road.

Links rather than junctions, and that is a data constraint rather than a preference. There is no junction inventory to hang a score on — NaPTAN has stops, GTFS has stops — so deciding which junction a delay belonged to would be a guess. A link names both its ends, which is better evidence anyway, and in practice a bad link is a junction reading, because a junction is usually what makes the road between two stops slow.

The score

Deliberately the same shape as the Pinch Point Factor Alchera publish for Lancashire, so the two can be read side by side:

Slowdown  ×  (10 + Variability)  ×  Daily buses  =  Pinch Point Factor
Term What it is
Slowdown 1 − quiet run ÷ observed run — the share of the run lost to congestion. A link taking twice as long as it should scores 0.5.
Quiet run The 20th percentile of the link's hourly median run times.
Variability 90th percentile ÷ median − 1 — how much worse an unlucky bus does than a typical one.
Daily buses Scheduled traversals per day, from the timetable.

Slowdown and variability are averaged across the day weighted by how many buses were on the link in each hour, so a link that is dreadful for one hour and deserted for the rest does not outrank one that is bad all day with buses on it throughout.

Three places this departs from theirs, and why

Slowdown is lost time, not a speed ratio. Drawn as a plain ratio of free-flow to congested run time, a worse link scores lower and the product moves the wrong way. It is only monotonic as a fraction of run time lost.

The baseline is measured, not scheduled. Deriving slowdown from the timetable alone was tested here and does not work: the median link's scheduled speed varies by 0.9% across the whole day, because publishers write one running time and reuse it at every hour. The timetable knows how many buses are due — which is where daily buses still comes from — but only the vehicle feed knows how long they took.

Quiet run is the 20th percentile, not the fastest hour. One freak 4am run otherwise becomes the standard every other hour is judged against, and every corridor in the region looks congested.

The (10 + …) damping constant is kept exactly as they draw it. It means variability moves the factor by a few per cent rather than dominating it — defensible (an unreliable link is worse than a uniformly slow one, but not ten times worse), and preserved so the numbers stay comparable rather than because it was arrived at independently.

The number to quote

Time lost per day — excess seconds per bus summed across a day's scheduled buses — not the factor. The factor ranks links; that one is in units somebody can price. North Bridge Road in Doncaster costs 279 bus-minutes a weekday between two stops 568 metres apart.

The live board

The same baseline asked about now: every link a bus has completed in the last 90 minutes, against what that link normally takes at this hour, on this kind of day. Not against a daily average, which would call every corridor congested at 8am and clear at 11pm whether or not anything was wrong. Day types are kept apart for the same reason — Saturday is not a quiet Tuesday, it is a different network with its own peak in a different place.

Why the threshold is not a percentage

A flat ratio was tried first and does not work. At 1.25× it called out 401 of 3,688 links at once: a median over three or four buses is noisy, run times are right-skewed, so small samples read high. A tenth of the network alerting is not a signal anyone can act on.

The baseline already knows each link's own spread, so the threshold comes from that: a link is called out when its typical bus right now is doing worse than an unlucky bus normally does on it — above the 90th percentile of its own hourly baseline — and is at least 30 seconds down. Self-calibrating, so a reliably erratic link is not reported for being itself, and a normally metronomic one is reported as soon as it slips. Same sample: 26 links, not 401.

Hurst Lane is the worked example. It was running 506s against a 196s baseline — 2.6×, and a flat rule would have shouted — but its normal 90th percentile is 606s. It is an erratic link having an ordinary day, and it is correctly quiet.

What this does not cover, and why

South Yorkshire only, today. A link can only be measured where the poller saw the bus at both of its ends, and that needs route shapes in the GTFS export to detect stop crossings from. Shape coverage tracks the result almost exactly:

Area Trips with shapes Adjacent links formed
South Yorkshire 98.0% 99.0%
Tyne & Wear 57.0% 6.0%
Merseyside 33.2% 2.5%
West Yorkshire 79.7% 1.9%
Greater Manchester 3.1% 0.6%

West Yorkshire is the outlier that does not fit the pattern and has not been explained yet — 80% of its trips carry shapes but almost no adjacent links form. Worth a look, because it is the largest network where this ought to work and does not.

Areas absent from the report are not thereby congestion-free. They are unmeasured. The area filter lists only areas with at least 50 scorable links for exactly this reason, and the page says so on its face.

The window is short. 17 days, of which 12 are weekdays and only 2 each are Saturday and Sunday, so weekend baselines rest on very little. Every figure carries its sample count, and an hour with fewer than 10 observed traversals is not published at all — not greyed out, because a figure on a screen gets quoted whatever colour it is. The window can never exceed 35 days: raw observations are pruned at that age, so the baseline outlives the data it was measured from and records on its face the window it was built over.

Bank holidays are not modelled. Daily bus counts come from the timetable's ordinary weekly pattern; calendar_dates exceptions are not applied.

Corridors, and reserving the words "pinch point"

Two things the link-level table could not say, both added after looking at the distribution rather than at any single link.

Most links are not pinch points. The median scored link loses 10% of its run and 4.2 bus-minutes a day. That is a road, not a problem, and applying the phrase to all 4,357 of them made it mean nothing. A link now earns is_pinch_point only by losing at least 25% of its run and costing at least 15 bus-minutes a day — about 3.4% of them. Both conditions matter: a short link can double in time and cost almost nothing, and a very busy one can bleed minutes a day without ever being dramatic.

The loss is corridor-shaped, and per-link ranking hides it. The ten worst individual links carry under 6% of the time this network loses; you would have to fix five hundred to recover half of it. Meanwhile ~2.5 km of the A61 through Woodseats spread itself across eight separate rows and got named nowhere. pinch_corridor groups consecutive links along one street and direction, so that reads as one entry losing 7.7 hours a day instead of eight small ones.

Corridors are built from the links and never replace them. The stop-to-stop rows are untouched, and clicking a corridor filters to them — that is the level any intervention is actually designed at, because a bus lane starts and ends at a place.

Two grouping traps, both hit before they were fixed:

  • Group by street and locality. "Doncaster Road" is a different road in Rotherham, Barnsley and Mexborough; grouping on the name alone produced a single 25 km "corridor" spanning all three.
  • Take direction from the links' own geometry, not NaPTAN's per-stop bearing. A bearing changes as a road curves, which split the A61 into separate S and SW corridors and lost two thirds of it. Each street's dominant axis is measured first, then each link placed on one side of it.

The page states the distribution under the table, so a reader can tell whether row 30 is a serious problem or an ordinary road.

Is it the road, or is it the buses?

The report's biggest blind spot, and the one thing it cannot answer from bus data alone. A bus losing two minutes where the traffic beside it is also losing two minutes is congestion; a bus losing two minutes while the traffic flows is a bus priority problem. Both look identical here, and they call for opposite interventions — a bus lane versus a junction scheme.

So the worst links are sampled against TomTom's Flow Segment Data, which returns currentTravelTime and freeFlowTravelTime for the road fragment nearest a point. Its ratio is the same construct as our slowdown, so the two subtract: what is left is the part of the delay that is specific to buses.

Ratios only, never seconds. TomTom answers for the fragment nearest the link's midpoint, which is not our link — different start, different end, different length. A ratio is dimensionless and survives that mismatch; subtracting its seconds from ours would compare two different pieces of road.

Check the matched road class before believing a reading. TomTom answers for the road fragment nearest a point, and near a junction that can be the side street rather than the road the buses are on. The very first sweep did exactly this: Glossop Road — an A road — matched to FRC4, with currentTravelTime equal to freeFlowTravelTime, which is TomTom's signature for "no live data, here is the reference profile". Read naively that says traffic flows freely while buses lose two thirds of the run — the single most misleading thing this comparison could publish, and it would have been the top finding on the page.

Only FRC0–FRC3 matches are used, on the grounds that bus routes are on classified roads. Identical current/free-flow readings are deliberately not excluded on their own: on a genuinely quiet major road at 3am that is the true answer. Discarded readings are counted and shown, because "we threw away every reading here because they were on a side street" is a different statement from "we have no readings", and the difference matters to someone deciding where to spend money.

The budget is enforced in code, not in the schedule. The free tier is 20,000 requests a month. 20 links × 2 sweeps an hour × 14 hours × ~30 days is ~16,800, and traffic_poll.py counts the month's spend from the samples themselves before each sweep — so a timer that misfires, a retry loop, or a few manual runs cannot quietly turn into a bill. If the clock and the budget disagree, the budget wins.

Unset TOMTOM_API_KEY and the job does nothing and says so; the report shows "not connected" rather than an empty column, and distinguishes that from "connected but this link has too few readings yet". Sampling covers 06:00–20:00 only — the hours either side are the ones where buses and traffic already agree.

Street works nearby

Roadworks arrive from Street Manager (DfT), free under the Open Government Licence. A pinch point that can say "and there are works here until the 19th" names a cause and a date it ends.

Push, not pull. There is no polling API and no sample data: access is over AWS SNS, so the endpoint has to exist and be publicly reachable before registration can succeed. Ours is POST /api/roadworks/sns, registered for the permit and activity topics. Section 58 was deliberately not subscribed — it covers restrictions preventing further works, so it says nothing about why a bus was delayed.

It is a public inbound endpoint, the only one on this site. Nothing is trusted until, in this order: the topic is one of Street Manager's three published production topics; the signing certificate URL is on an AWS SNS host, checked before it is fetched; and the signature verifies against that certificate. The order matters — validating the certificate URL after fetching it would be no validation at all, and skipping it turns the endpoint into a request forgery primitive, since an attacker could sign with their own key and supply their own certificate to verify it against.

Order events by version, not by timestamp. Street Manager stamps each object with a version that increments per change. SNS does not promise delivery order, and two events on one object can share a timestamp — so a redelivered "permit granted" arriving after "work started" would walk the status backwards.

Coordinates are British National Grid, not degrees. Eastings and northings. Treating them as lat/lon puts every work in the country off the coast of Africa; the guard is a magnitude test, since no longitude exceeds 180.

Proximity is not attribution. Works are matched to a link if they fall within 250 m of its midpoint. That is a "same bit of street" test — the link is a couple of hundred metres of road and the work's position is the midpoint of a line the highway authority drew. The page says works near this link, never works causing this delay.

No backfill, ever. The table only learns about works whose events arrive after the subscription is confirmed. A job already underway today is invisible until its next event, which makes "no works on record" much weaker evidence of a clear road than it looks — and the page says so rather than showing an encouraging blank.

Where it runs

  • pinch_rollup.py, nightly at 04:10 via pinch-rollup.timer, after the punctuality rollup because that job prunes the observations this one reads. The whole window is rebuilt every night — the baseline is a median over a rolling window, so a day joining or leaving it moves figures already written, and there is no incremental form that stays correct.
  • Tables pinch_link (scored summary) and pinch_link_hourly (the baseline). Neither is pruned.
  • pinch_web.py serves /pinch-points and /api/pinch/*.
  • traffic_poll.py, half-hourly 06:00-20:00 via traffic-poll.timer, into traffic_samples.
  • roadworks_web.py receives Street Manager's SNS notifications into roadworks. No timer -- it is a push feed. /api/roadworks/status shows what has arrived.