Three use cases. One scoring engine.
Every card below is a live score sitting in the database — click any one to see the per-metric breakdown, judge rationale, and ranked suggestions. To score your own URL, ticket, or PR, you'll need an account.
Score any webpage through a customer lens.
Same judge, four real homepages. The rationale cites copy from each page — these are real scores, not screenshots.
Judge homepage (this site)
Pending — this score is being generated.
Stripe
Pending — this score is being generated.
Linear
Pending — this score is being generated.
Vercel
Pending — this score is being generated.
Score any text artifact — even a Linear ticket.
The spec-completeness rubric scores a vague 'Fix slow dashboard' ticket against scope, acceptance criteria, success metric, and edge cases. Same engine works for any text artifact.
ENG-1234 · "Fix slow dashboard"
Pending — this score is being generated.
Score a real GitHub PR — diff, description, and all.
A merged Vercel/SWR PR fetched live via the GitHub API and judged by code-quality. Connect the GitHub App and this happens automatically on every PR.
vercel/swr · PR #4243
Pending — this score is being generated.