“Insufficient evidence” is a first-class outcome.
Most systems are designed to return an answer. RoadPatch is designed to stop when the submitted evidence cannot support one yet, before an unsupported result travels into procurement or deployment.
Two shipped projects show the discipline between a technical result and a decision: what the evidence supports, what it does not support, and who still holds authority.
A replayable evidence gate for public-sector evaluators deciding whether a computer-vision road-inspection model has enough preserved evidence to justify a larger local replication study. It is deliberately not an approval system: deployment authority remains with a named human evaluator.
Most systems are designed to return an answer. RoadPatch is designed to stop when the submitted evidence cannot support one yet, before an unsupported result travels into procurement or deployment.
The repository makes no current claim of GPU speedup, accuracy improvement, training-time reduction, cost saving, TensorRT result, multi-node result, or customer outcome. Writing down those boundaries is part of responsible technical communication.
Reports and manifests are bound by SHA-256.
Declared taxonomies are checked before comparison.
Declared data overlap is recomputed from shipped JSON.
The Decision Card remains non-authoritative.
The gate validates what the shipped package claims and whether that package is internally supportable. It does not independently establish the underlying ground truth.
Two-person hackathon build with Bobin Joseph, credited in the project README. Bobin brought the concept and the CO₂ domain research; I did the majority of the implementation, including the agentic layer, the simulation rules and the model benchmark. The repository is hosted on his account.
My contribution focused on testing whether a model-generated policy result was stable enough for a committee to defend, rather than merely producing a persuasive number.
Five model setups plus one deterministic baseline produced divergent results: one was roughly −22%, while another was roughly +1.5%.
The work was to establish which result a committee could defend, and where model choice became load-bearing.
The shipped working configuration ran entirely on local NVIDIA DGX Spark hardware with no cloud calls.
CarbonCrest found that a universal flat credit is very largely deadweight: most of the budget rewards people who had already switched.
Both projects are about the same discipline: deciding what a technical result is actually worth before anyone acts on it.
That also means writing down plainly what is not being claimed.