Independent research · SafeFuture Lab, rCITI UNSW · 9 October 2026
Does Compass IoT know how fast
Sydney actually drives?
We compared Compass's open-API operating speeds with what two instrumented SafeFuture cars drove on the same road segments, at the same times of day, over five months. Compass differs from the cars by more than measurement noise can explain. The data do not show that the speed limit or land use drives that difference. Queried hour by hour, Compass does register reported crashes, but only modestly, and it misses the brief slowdowns our cars met.
Press → or Space to step through. All figures are aggregates; no routes, trips or locations are shown. Not affiliated with Compass IoT.
The answer
Six questions, decided by rules set before we looked
A finding counts only if it survives a multiple-testing correction (Holm) within its family. For H1–H3, both cars must also agree in direction.
| Question | What we found | Units | Verdict |
|---|---|---|---|
| H1 · Is Compass more than 5 km/h off the cars? | Off by 7.3 km/h on average (CI 7.0–8.6), Compass 2.9 km/h slower. Beyond the 2.3 km/h noise floor the gap is 5.0 km/h (CI 3.7–6.2). Both cars agree. | 205 | Supported as a discrepancy |
| H2a · Does the posted limit pull Compass speeds? | Limit coefficient 0.20 (CI −0.03 to 0.51), Holm p 0.30. Holden +0.22, Navara −0.29. | 144 | Not supported |
| H2b · Does the gap differ by speed-limit band? | Joint test p 0.67, Holm p 0.91. | 144 | Not supported |
| H3 · Does the gap differ by land use beside the road? | Joint test p 0.45, Holm p 0.91; the cars disagree in direction. | 144 | Not supported · exploratory |
| H4 · Is Compass slower in the hour our cars met a slowdown? | Mean rank 0.50, exactly what no response would give (p 0.54). | 60 | Not supported |
| H5 · Is Compass slower near a reported crash, and more than on nearby roads? | Mean rank 0.36 vs 0.50 if no response; nearby same-day roads 0.48. Median −6%. | 386 | Supported modest effect |
Units are road-segment × time-window combinations with at least 5 passes on at least 3 days. H2/H3 use the 144 units where Compass's speed limit agrees with TfNSW's official zone. H4/H5 come from Phase E (slides 12–14); their units are events.
How it was done
Two cars, one road network, matched segment by segment
Every car pass was snapped to the exact Compass road segment and direction, then compared with Compass's speed for that segment in the same weekday/weekend and hour band.
Speed limits come from TfNSW's official speed zones, not from Compass. Land use comes from NSW planning zoning beside the road. The analysis protocol was frozen and externally critiqued before any comparison was computed.
From GPS points to comparable units
143,275 GPS points became 205 comparable units
Most data drop out because Compass splits roads into very short pieces: 41% are under 30 m, too short for a car logging every 4 seconds to cross with three readings.
H1 · Accuracy
Compass tracks the cars, but sits below them on fast roads
Each dot is one road segment in one time window. Dots on the diagonal mean perfect agreement; the shaded band is ±5 km/h.
- Correlation is high (concordance 0.94): Compass ranks roads by speed well.
- Single segments can be far off: 95% of gaps fall between −19.8 and +14.0 km/h.
- On motorways the cars sit at ~100 km/h and Compass at 92–99 km/h.
140 publishable units shown (segments driven on ≥3 trips, away from home and work areas). Statistics use all 205 units.
H1 · How much is real?
About 5 km/h of the gap is real, not noise
A car's median from a handful of passes is noisy on its own. We bootstrapped every unit's passes to measure that noise: it accounts for 2.3 km/h of the 7.3 km/h gap. What remains, 5.0 km/h, sits right at the 5 km/h tolerance.
H1 · Where the gap is
The faster the road, the lower Compass reads
Close to the cars on slow streets, 5.6–8.1 km/h below them once the cars drove above 50 km/h. Above 70 km/h even the 95% range of gaps sits entirely below zero.
This fits either a Compass statistic that smooths towards typical traffic, or drivers who are faster than average on fast roads. Two cars cannot tell these apart.
Who is driving matters
The Navara drives closer to the limit on 70–80 km/h roads
Car speed minus the posted limit, by limit. On busy 50–60 km/h roads traffic sets the pace for both cars; on 70–80 km/h arterials the Navara runs 7–12 km/h nearer the limit.
- 80 km/h zones: Navara above the limit on 38% of passes, Holden on 10%.
- Same roads: on 109 segments both cars drove, the Navara was 1.9 km/h faster.
- Gap to Compass: −4.0 km/h for the Navara, −2.3 km/h for the Holden.
Post-hoc analysis prompted by the drivers' own description. It does not change any verdict, but it shows the Compass–car gap depends on who is driving.
H2a · Speed-limit effect
No stable pull towards the speed limit
If Compass leaned towards the posted limit, its speed would rise with the limit even at the same car speed (coefficient g above zero). Across specifications, g wanders from −0.18 to +0.30 and flips sign between the two cars.
g = extra km/h of Compass speed per 1 km/h of posted limit, at equal car speed. Lines are 95% intervals. The registered test could detect g ≥ 0.34. The all-hours view (bottom) mixes times of day and is exploratory only.
H3 · Land use (exploratory)
No clear difference by what is beside the road
Mean Compass–car gap by the zoning beside the road. The formal test, controlling for limit, road class and time, gives p = 0.45, and the two cars disagree in direction.
- A trap we fixed: NSW classified roads are zoned SP2 Infrastructure, so a buffer round a main road mostly measured the road itself. We used the zoning beside the road instead.
- Groups were fixed by unit counts before any gap was computed.
Time of day (hypothesis only)
Compass reads lower in the daytime than overnight
Weekday gaps are around −3 to −7 km/h between 10 am and midnight. The overnight band has only 8 units, and only 8 segment pairs were seen both day and night, so this is a question for the next study, not a finding.
Phase E · Short windows
Does Compass notice when something goes wrong on the road?
Compass will answer for a single hour on a single day. We asked it about the hours around 400 crashes reported on TfNSW Live Traffic, 200 breakdowns and 66 slowdowns our cars ran into. Each event is compared with the same hours on the same weekday, up to four weeks either side.
Crash times come from a public archive of the TfNSW Live Traffic feed, saved about every 30 minutes. The test plan was pre-registered (protocol v1.2) and critiqued before any event-hour Compass data was requested.
Phase E · Result
Compass slows near reported crashes, but only a little
- Crash day is the slowest of its comparison days in 34% of cases, against 13% by chance. Median speed change: −6%.
- It is local. Same-class roads 2–5 km away on the same day barely move (0.48), which argues against a purely day-wide cause such as rain.
- But usually faint. Only 13.5% of crashes show as at least 20% slower and the slowest day (5% by chance). Crashes active for more than ~18 minutes: −7.6%; shorter ones: −2.9%.
- The cars' slowdowns don't show up in the hourly average (0.50). They were short, some may be specific to the car, and with about 22 independent locations this test has little power.
Robust to clustering by date, road and 2 km area. Bars are approximate 95% intervals. Compass returned data for 95–98% of event windows.
Phase E · What it looks like
One clear signal, and two more typical cases
Compass speed on the event road, compared with the typical of the other days, for each same-weekday date. The event day is in blue.
The strongest crash in the sample, with a lane closed (Sunday 2 August, 1–4 pm), drops 92%. A typical crash (Tuesday 15 September, 8–11 am) does not stand out from ordinary week-to-week swings. Neither does a car slowdown (Monday 20 July, 5 pm), even though the car was at 18% of its usual speed. No locations are shown.
Read before quoting
What could make this wrong
- Two cars, one household. This is agreement with these drivers, not with the truth. A third speed source on the same segments is the next step.
- Habitual routes. Results describe often-driven roads, mostly Sydney's North Shore, not the network.
- Short segments excluded. Junction-heavy streets are under-represented; the planned segment-joining fallback produced no units.
- Map matching. 77.5% of GPS points matched (median 4.7 m from the road); 70% agree with OpenStreetMap's nearest road, 77% away from junctions.
- Different measures. The cars give distance ÷ time including stops; Compass's averaging method is undocumented.
- Compass speed-limit field: matched TfNSW exactly on 82% of compared segments and within 10 km/h on 93%.
- Day-of-week filter: Compass applies it in UTC, so up to 2% of days fall in the wrong weekday/weekend window.
- 66 corridor clusters for the limit and land-use tests; uncertainty may be understated.
- Process deviations are all logged with timestamps (13 orchestrator entries). Land use was redefined, and a grouping anchor was committed late; H2b and H3 are therefore reported as exploratory.
- Phase E: crash times are listing times, not crash times; locations are reported points; one-hour windows dilute short events; 400 of 1,747 eligible crashes were sampled.
- Reproducibility: data building reproduces exactly from frozen inputs; the analysis also needs two frozen covariate files kept with checksums.
How it was checked
Built by separate agents, attacked by outside critics
No step graded itself. Builders, analyst, replicator and critics worked from committed files only; external critics from two different model families reviewed the plan, the results and the final wording. Phase E was reviewed externally by ChatGPT Pro; both Grok reviews of it failed to return. A separate agent reran it offline from a fresh clone and reproduced every result exactly.
Reproduce it
Everything is versioned, logged and rerunnable
- One command, offline from frozen inputs:
./run_all.sh - Database:
data/analysis_final.duckdb— every stage table, the run log and the Compass request log - Run log: every stage start and end with input and output checksums, row counts, agent and git commit
- Results:
outputs/RESULTS_FINAL.json— each headline number with its source table - Protocol and deviations: frozen protocol tags v1.0, v1.1 and v1.2 (Phase E); deviation log O1–O15; Phase E reruns offline with
./run_phaseE.sh - Independent replication: a separate agent rebuilt the pipeline through the analysis models from a fresh clone with identical package versions. Result: reproduced with differences, all from one cause: its covariate stage fetched extra OpenStreetMap road names live, which changed the corridor clusters (130 → 102). Headline intervals moved by at most 0.16 km/h and no verdict changed; the match rate, TfNSW limits and land use reproduced exactly.
Next steps
- Add a third speed source (loop detectors or another fleet) on the same segments.
- Rerun the analysis on regenerated covariates.
- Join short Compass segments into road sections to recover junction-heavy streets.
- Test the day-versus-night pattern with matched segments.
- Ask Compass about sub-hour windows, and test incidents against loop-detector data.
Raw GPS traces stay on SafeFuture Lab servers. Data sources: SafeFuture telematics server (Traccar), Compass IoT open tile API, TfNSW speed and school zones (CC BY 4.0), NSW Planning Portal zoning, OpenStreetMap (ODbL), TfNSW Live Traffic incidents via the public archive github.com/jxeeno/nsw-livetraffic-historical.