Sample report. Every name and number is invented. The paid product is not on sale yet.
Rafael Quintana
PassedSeatAssessed work at Fernbrook Systems, last 12 months. Personal training is not counted here and never reaches a team report.
- Submissions
- 5
- Findings caught
- 9 of 12 planted
- Catch rate
- 75%
- False positives
- 1
- Median score
- 75 against a pass mark of 70
Against a frontier model
How a frontier model scored on the same challenges, given the same diffs and marked against the same answer key, with no hints and no walkthrough. The model is not named, because these figures are invented for this sample and an invented number under somebody else's product name is a claim about their product.
- This person
- 75 of 100
- The model, on the same challenges
- 71 of 100
- Difference
- +4 points
Both figures are out of 100 over the 5 challenges of theirs that a published run covers, so they are only compared on work the model has actually done.
Score over time
Every attempt in the order it was made, oldest on the left, against the pass mark.
What they catch and what they miss
The best and worst of this person's own results by weakness category. It is their own detail, not an aggregate, so no minimum cohort applies. The whole breakdown is in the table below.
Strongest
- AI safety100% caught2 found of 2 planted0 missed
Weakest
- Data exposure50% caught1 found of 2 planted1 missed
By CWE category
What this reviewer catches and misses, by weakness family.
| Category | Caught | Missed | Catch rate |
|---|---|---|---|
| Injection | 6 | 2 | 75% |
| Data exposure | 1 | 1 | 50% |
| AI safety | 2 | 0 | 100% |
By language
The same breakdown by the language of the codebase reviewed.
| Language | Caught | Missed | Catch rate |
|---|---|---|---|
| PHP | 5 | 2 | 71% |
| Python | 4 | 1 | 80% |
Every attempt
Newest first. Times are in UTC.
| Challenge | Score | Caught | False positives | Took | Submitted |
|---|---|---|---|---|---|
Search results assembled into the prompt deskmindchallenge 2 | 75 | 2/3 | 1 | 27 min | 22 Jul 2026, 16:23 UTC |
A ticket summary that follows the instructions in the ticket deskmindchallenge 1 | 93 | 2/2 | 0 | 25 min | 21 Jul 2026, 07:23 UTC |
An export carrying columns the page never showed marksheetchallenge 3 | 43 | 1/2 | 0 | 23 min | 19 Jul 2026, 22:23 UTC |
A search filter pasted into the where clause marksheetchallenge 2 | 50 | 1/2 | 0 | 21 min | 18 Jul 2026, 13:23 UTC |
A grade query built by concatenation marksheetchallenge 1 | 100 | 3/3 | 0 | 18 min | 17 Jul 2026, 04:23 UTC |
