PROOF FIRST / proof before paymentSEND THE BUG ↗

RADAR METHODOLOGY

HOW THE RADAR
COUNTS THINGS.

CURRENT STATUS: TESTED AND COLLECTING. Before we used the Radar, we checked it against 52 real technical cases: 16 found correctly, 1 missed, 0 false alarms. The home page shows a dated snapshot of counts from the last 30 days. Counts that need more data than we have are not shown.

WHAT THE RADAR LISTENS TO

Permitted public technical-problem posts, collected through official or public interfaces: GitHub issues and discussions, and Hacker News. Communities whose terms forbid automated access aren’t collected automatically. A person may add a specific public post by hand. It’s a limited, targeted sample, not the whole internet.

PUBLIC SIGNAL

A post by a person describing a concrete technical problem they are experiencing. Posts written by AI agents or automation are excluded from the public figures and counted separately. A public signal is not a verified bug.

SAID FIXED / STILL BROKEN

Among qualifying human public AI repair-attempt reports in the observed dataset, the count and share where the poster subsequently reported that the original behavior remained broken or another regression appeared.

RECURRING PATTERNS

Each signal is assigned to one pattern from a fixed list of symptoms. A pattern recurs when several independent people report it within the window. A pattern is a group of similar symptoms, not a root cause. Pattern counts are not shown until there is enough data.

REPRODUCED FAILURES AND VERIFIED FIXES

These come only from actual Proof First cases, never from the Radar. A case counts as verified only with BEFORE = FAIL, SAME TEST LOCKED, repair, the same test rerun, and AFTER = PASS, with the evidence saved. Controlled demonstrations on sample apps, such as the video on the home page, are never counted as customer cases. A model saying “fixed” isn’t verification.

HOW IT’S CHECKED

An automated classifier (Jev) assigns bounded labels to each post. It doesn’t repair software. We compared its answers with a reviewed sample of 52 real cases: people decided every disagreement and a random spot check, with an AI giving an independent second opinion. It is a small test, not a guarantee, and the individual counts shown on the home page are made automatically and are not reviewed one by one. Counts change as new reports are collected.

WHAT THE RADAR KEEPS

Every post the Radar collects is public. For each one it keeps the post’s text, link and date. It does not store the writer’s username as such; it keeps a one-way code, used only to count distinct writers. The link and the text can still contain public names, such as a repository owner or an @mention. These are people who posted in public, not Proof First customers. The Radar never emails or privately messages them; any reply stays public, in the same thread. Posts are labelled automatically, using a cloud classifier (Jev, from TypeSafe) that receives the public post text. The labels are not reviewed one by one, and the Radar is a limited sample, not a complete record of anyone or anything.

The data is kept while the research is active. If you wrote a post the Radar holds and want it removed, email christopherarivero@gmail.com with the link to the post and I’ll delete it. Everything else about what Proof First collects is on the privacy page.

KNOWN LIMITATIONS

← BACK TO PROOF FIRST · SEND THE BUG ↗