Announcements

AI judging has arrived

Pinned 3 replies
RV.studio Staff 2026-07-02

The analyst is in. Starting today, every blind session asks one new question up front — who judges your viewing? You now have three answers.

The three ways to be judged

  • I judge — exactly what you're used to: you rank the candidate images yourself at reveal.
  • AI analyst — a blind AI ranks the panel against your impressions. Flat 1 credit per analysis.
  • Human analyst — invite anyone you trust to judge it — by email (no account needed) or by member handle. Free.

With either analyst, you never see the decoys at all — not before reveal, not after. Your reveal shows only your target, the placement it earned (say, #1 of 4 — direct hit), what in your impressions matched it, what conflicted, and the analyst's confidence.

That's the strictest blinding we offer: you finish the trial having seen nothing but your own work and the real target. And because you never see the decoys, those images stay fresh for future draws.

How the AI judges

The analyst uses elimination ranking — the same discipline a trained human judge applies:

  1. It breaks down each image independently against your impressions: what corresponds, what conflicts.
  2. It rules out the least likely match first.
  3. It picks the strongest match.
  4. It orders what remains, working from the extremes inward.

Crucially, the analyst is blind too. It is never told which image is the real target — it can only judge descriptive correspondence. Placement is computed afterward, mechanically, from where the real target landed in its ranking. There is no way for the answer to leak into the judging.

It reads your sketches, not just your notes

Your committed notes and sketches are both first-class evidence. The analyst weighs structural and gestalt correspondence — shapes, topology, land/water, natural vs. man-made, energy, motion, color — over incidental word overlap. A curve you drew counts as much as a sentence you wrote.

It also speaks session notation. EOS, breaks (BK, AOL break, confusion break, and the rest), stage labels, timestamps, and the divider lines you draw between passes are treated as procedure, never perception — they will not be scored as impressions of the target. And an AOL you flagged yourself is never held against you: the analyst judges the descriptors around it, not the label. Declaring your overlays remains the right move.

How human analysts work

Pick Human analyst when you begin (or decide later — any forced-choice viewing can be handed to a person even after you've finished viewing). Once your impressions are committed, tap Invite your analyst: enter their email or handle and they get a private judging link. It works without an account, judges exactly once, expires in 7 days, and you can revoke it any time to judge the session yourself instead.

Your analyst sees your impressions and the unlabeled candidates — never told which is real — ranks them, and you're notified the moment your reveal is ready.

What it costs

Every AI analysis is a flat 1 AI credit, and every member starts with 10 non-expiring credits. Human analysts are free. (Viewing feedback is also 1 credit; pool curation tools run 1 credit per 10 images.) Admins can top you up if you run dry while we tune things.

What we're testing

This is a calibration phase, and honest data is the whole point:

  • Scoring consistency — does the analyst place targets the way a careful human judge would?
  • Sketch reading — does it weight your drawings fairly against your prose?
  • Notation handling — does it correctly ignore EOS/breaks/dividers and respect flagged AOLs?

If a judging feels wrong — a placement you'd dispute, a sketch it clearly misread, a notation it scored as content — tell us. Use Report a problem in your profile menu (it attaches the context automatically) or post in the feedback category here. Analyst-judged trials count toward your forced-choice accuracy and are tracked separately from self-judged ones, so nothing about your record gets muddied while we tune.

Try it

Begin a blind session → pick a 2- or 4-image panel → Who judges: AI analyst or Human analyst. Work the target as always, then send it to your judge. Sharing a judged session carries the full breakdown too — placement, confidence, and what matched.

Good viewing. 🎯

Elios Aristide 2026-07-10

Nice! Anything to come with pvalue, Zscore, variance etc ?

RV.studio Staff 2026-07-11

Great question — statistical honesty is basically the whole point here, so most of this is already in, and the rest is exactly where we're headed.

Live today:

  • Brier score everywhere we score — the calibration loss (lower = better) that rewards a well-placed confidence, not just being right. It's on your dashboard, the duels ladder, and the leaderboards, and for multi-outcome (n-way) predictions it's the multiclass Brier. For market-tasked ARV events we even compute the market's Brier from its implied odds — so you can see whether you beat the market's calibration, not just the coin flip.
  • Hit-rate vs the exact chance baseline — every forced-choice result sits next to its true chance line (1/N: 50% for a 2-target, 25% for a 4-target). So "62%" is always shown against "vs 25% chance."
  • A calibration curve — your confidence bucketed against realized accuracy, so you can see whether your "70% sure" calls actually land ~70% of the time.
  • On the brain-map / EEG side: per-band z-scores against your own baseline with Holm-corrected p-values, flagging which bands were significantly elevated — within a session and across sessions. So "theta was up" comes with a real significance test, not a vibe.

Not in yet — and we'd genuinely like input on:

  • A formal significance test on your overall hit-rate vs chance — a binomial/z-test giving an actual p-value + confidence interval ("your 62% over 40 blind trials is p = 0.03 above the 25% line"). This is the big one and it's high on the list.
  • Variance / standard error / confidence intervals on hit-rate and Brier, so small samples don't over- or under-sell themselves.

Short version: Brier, chance baselines, and EEG z/p are live today; explicit p-values + confidence intervals on your accuracy are the natural next step. If there's a specific statistic you'd want front-and-center, tell us which and how you'd use it — that shapes what we build first.

Elios Aristide 2026-07-11

Thx !

EA

Join the conversation — members can reply, start threads, and view every discussion.

Log in to reply