How it was scored
The rubric was written before anything was looked at, and the same one was used for all three runs with its weights unaltered. This page is all of it: the fifteen criteria and their weights, the evidence rules, the scores the evaluator revised before reporting, what we corrected before these runs, the fabricated candidate we used as a control, and all three raw outputs with nothing taken out.
The results are here. This page is the working behind them, published so that a sceptical reader can disagree with the method rather than having to take the number on trust.
On this page
The fifteen criteria and their weights The evidence rules Scores the evaluator revised before reporting What we corrected before these runs How pages were read The prompt and the unedited outputs The product that does not existThe fifteen criteria and their weights
One hundred points, declared in the prompt before any product was examined, and weighted for a UK social club night of 18 to 26 players on four courts where the live night is the main problem.
| # | Criterion | Weight |
|---|---|---|
| 1 | Arrival, check-in and late joiners | 6 |
| 2 | Dynamic next-game and court allocation | 12 |
| 3 | Skill-balanced doubles | 12 |
| 4 | Waiting-time and sit-out fairness | 10 |
| 5 | Partner and opponent variety | 8 |
| 6 | Explainability and transparency | 8 |
| 7 | Manual override and special cases | 5 |
| 8 | Player rating system | 7 |
| 9 | Live scoring, statistics and history | 5 |
| 10 | Multi-court display and kiosk operation | 5 |
| 11 | Attendance and session reporting | 5 |
| 12 | Payments and member administration | 5 |
| 13 | Competitions, tournaments and league workflows | 4 |
| 14 | Resilience and device accessibility | 4 |
| 15 | Price and practical limits | 4 |
Forty-eight of the hundred points sit in the first five criteria - arrival handling, next-game allocation, skill balance, sit-out fairness and partner variety. That is the live night, and it is where the weight deliberately went. Weight the other ten more heavily and a different product wins, which is the honest way to say that this result is scoped rather than universal.
The evidence rules
Each criterion was scored in one of four states, from the product's own documentation:
- Verified full - the full weight.
- Verified partial - half the weight.
- Verified absent - nothing, and the criterion counts as resolved.
- Insufficient evidence - nothing, and the criterion counts as unresolved.
Four rules did the real work:
- Insufficient evidence is never absence. It scores nothing but adds its full weight to the product's unresolved maximum, which is why every product has a range.
- One canonical URL per product, treated as its identity anchor, so a similarly named product elsewhere could not be substituted. This was added after an earlier run misidentified ShuttlePilot as a vehicle-dispatch product and scored it zero.
- Self-consistency. Anything the evaluator qualified in prose had to be scored partial rather than full. This is what produced the revisions below, including ours.
- A near-perfect total is a warning. The evaluator was told to treat one as a signal to re-examine rather than a result. The first run returned 100 out of 100 for ePegboard; this rule is why the published run does not.
Documentation volume was explicitly not to be rewarded. It is worth saying that this rule only partly held: the evidential bar rewards documented capability, and a product that publishes a single landing page cannot demonstrate what it does. That is a real limitation of the method and it works in our favour, because we publish several hundred documentation pages.
Scores the evaluator revised before reporting
Two of the rules exist to catch an evaluator flattering whoever documents best. Both fired on us.
The near-perfect-total rule cost us five points. The best-evidenced run first returned 95.5 for ePegboard, with full marks on thirteen of fifteen criteria - and then, in its own words, treated that as "exactly the pattern the rule warns about: a rubric resolved entirely from one vendor's own material." It re-examined every one of those marks and cut two. 95.5 became 90.5, and our lead over second place fell from 11.0 to 6.0.
It also considered cutting a third and deliberately did not, because the same weakness applied to four rivals and downgrading only the leader would have been the unfair call. That reasoning is in the published output.
Ours is listed first below because a scoring pass that only ever finds fault with other people's products is not a scoring pass. The parity rule then raised seven scores that had been written too harshly - six of them on explainability, where a line had to be drawn and then applied to everybody.
| Product | Criterion | Change | Reason given |
|---|---|---|---|
| ePegboard | 1 Late joiners | full to partial | Check-in and kiosk mechanics are documented thoroughly, but what happens to a game already in flight when someone signs in mid-session is implied by the queue model rather than stated as a rule. ShuttlePilot and vsBadminton state it explicitly. |
| ePegboard | 13 Competitions | full to partial | The feature summary was read and the how-to pages noted but not opened. Scoring what was actually read, rather than the pages assumed, gives a partial. |
| ePegboard | 12 Payments | held at partial | Not raised, because the vendor's own words qualify it: not a payment processor. |
| ePegboard | 14 Resilience | held at partial | Not raised: offline running is available only through a paid Windows client. |
| ShuttlSync | 3 Skill balance | full to partial | Skill-split sides are one of three selectable modes rather than the default behaviour of the fair rotation, and the report qualifies it in prose. |
| BadBoard | 7 Manual override | full to partial | Override is documented for game-based mode only; timer mode is described as running unaided all evening with no override path stated. |
| Court Manager | 6 Explainability | partial to full | Raised for parity: the waiting list publishes games played and time waited. |
| Racket Social | 6 Explainability | partial to full | Raised for parity: a public read-only viewer shows court assignments, the sit-out queue and match history. |
| Shuttl | 6 Explainability | partial to full | Raised for parity: a visible match queue plus a published precedence - excessive waits first, then mixing, then skill balance. |
| QourtX | 6 Explainability | partial to full | Raised for parity: the wallboard shows playing, waiting and resting with wait-time estimates. |
| NextFour | 6 Explainability | partial to full | Raised for parity: every court at a glance - who is on, who is next, who has been waiting. |
| Queue Master | 6 Explainability | partial to full | Raised for parity: per-player games, wins and a running wait timer, sortable by shortest wait. |
| UBR | 3 Skill balance | partial to full | Raised for parity: the same class of evidence as two products already scored full. |
| Ladderly | 4 Sit-out fairness | held at partial | Held while Shuttl was raised. Both name wait time as an input; only Shuttl publishes the precedence between waiting, mixing and balance, and precedence is what the criterion asks for. |
Read the "unresolved" rows as what they are: the evaluator could not establish the capability from that product's documentation. Several of those products may well have the feature.
What we corrected before these runs
This is not our first attempt. An earlier set produced a defensible result on an indefensible set of inputs, and rather than publish the caveats we fixed the inputs and ran it again. Four things were wrong:
- We told ourselves a real competitor did not exist. SmashHQ was in the candidate list from an old note with no address. Two of our searches found nothing, so we recorded it as fictional - and treated an earlier run that had found it as inventing evidence. It is a live product, and it scores in all three runs above.
- We aimed the whole exercise at the wrong page for Court Manager - its admin login rather than its product site - and then recorded that it published no product information. At the right address it places sixth, and the run calls its picker documentation the most complete configuration description of any candidate.
- Six sites we recorded as unreadable were working normally. They build their pages in the browser; our reader could not run JavaScript.
- Three real rivals were missing from the field - ShuttlSync, ShuttleStats and SportRevive - because research done for our own guide never reached the list we drew candidates from. ShuttlSync now places second in two of the three runs.
The pattern in all four is the same: our tooling could not see something, and we recorded that as a fact about somebody else's product. Nothing from that earlier attempt is reported on this site. The prompt used here corrected every one of the four, and kept the rubric, the weights, the club profile and every scoring rule unchanged.
How pages were read
Whether each product's page could actually be read. This is the difference between "we could not check" and "it does not do this", and every run publishes it per candidate, naming the URLs it opened and whether each rendered, was blocked or was absent. They are in the outputs below.
Two of the three runs read 32 of the 33 candidates. One name in the list is fabricated and resolves to nothing, so 32 is the ceiling.
The exception is the free-tier run, which reports honestly that it cannot execute JavaScript and fell back to a search index where a site returned a blank shell. It scored 30 of 33 - and it also produced the highest score for us of the three. Weakest retrieval, most flattering number. We are pointing that out about our own best result because it is the sort of thing a reader should be told rather than left to find.
Why rendering matters this much. Several of these products build their pages in the browser and return only a page title to a plain fetch. A reader that does not run JavaScript sees an empty site and can easily record that as a product with no documentation - which is a false statement about a real business, and one we have made before. Requiring a real browser is the single biggest quality difference between this set of runs and our earlier attempt at it.
The prompt and the unedited outputs
Published as plain text, exactly as sent and returned, and not tidied. They contain the parts that do not suit us: the run that scores us 70.5, the run that cut our own score by five points before reporting, and the section concluding that most of our lead is documentation depth rather than demonstrated capability.
- The prompt, in full - 33 candidates, identical for all three runs, weights unaltered
- Claude Opus 5, unedited - 32 of 33 scored, with per-candidate retrieval logs. The most completely evidenced of the three
- ChatGPT 5.6, unedited - 32 of 33 scored, and the harshest reader of the three
- ChatGPT free tier, unedited - could not execute JavaScript, says so itself, and ends part-way through its own report
One run declared an interest of its own, unprompted. The Claude run noticed it was executing on a machine carrying an ePegboard project folder, said so at the top of its report, and pointed the reader at the section quantifying how much of the leader's margin is documentation rather than capability. It did not read the folder. It then cut our score by five points under its own near-perfect-total rule.
The product that does not exist
The candidate list includes a fabricated product - a product name we invented, given a plausible address, and listed among the real ones. No run scored it. All three reported it unlocatable and moved on.
That control exists because our earlier attempt had one that failed in the opposite direction. We listed SmashHQ with no address, could not find it in two searches, and concluded it was fictional - then treated a run that did find it as having invented the evidence. It is a real company with a live product. A name we made up ourselves cannot produce that error: if a run scores it, the run is unreliable, and there is no second possibility.
If you make one of these products
If your product was not assessed, or you think a criterion was scored wrongly, or anything we have written about it is simply inaccurate, tell us and point us at the evidence - a link to the page that shows it is enough.
We will correct factual errors quickly, re-run the rubric against a URL you choose where the scoring is the problem, and publish the outcome either way, including where it moves a competitor above us. Every score here rests on what a vendor publishes, so the vendor is usually the person best placed to tell us we have read it wrong.