00 / THE DECIDERS
Who was asked.
Each decider read the same request and chose a book, a pace, a sound and a look from RISE’s menus. RISE then checked every answer with the same code it runs for readers. Letters are fixed and carry no order of merit.
01 / AGREEMENT
Agreement with author-written expectations.
For each request, its author wrote down what a fitting answer looks like: a request for a slow, restful reading should get 100 or 150 words per minute, for example. Expectations met counts how often a decider’s admitted choice fell inside them. Opposites told apart counts how often two opposite requests got different choices. These are one author’s expectations, not ground truth, and the cases were tuned on Jev.
02 / CALIBRATION
When it says 30%, does it agree 30% of the time?
Some deciders state how likely each option is. Each dot below groups choices by the probability the decider gave them, across, against how often those choices met the author’s expectation, up. On the dotted diagonal the two match; below it the decider was surer than its agreement rate supports. A larger dot holds more choices, and a hollow dot holds fewer than 15, too few to read alone.
The same bins as a table
03 / CASE BY CASE
What each one chose.
Run 1 of every request. Each cell is the decision RISE admitted: pace in words per minute, sound, visual mode and style, then the book. Follow a cell to play that exact decision as a reading in RISE.
04 / ANSWERS, COST, TIME
What a decision costs.
How many answers RISE admitted, why the rest were turned away, how long a call took and what the provider billed, over every run of every request and control.
05 / METHOD
One run file, one command.
The harness lives in RISE under scripts/arena/. An operator runs it by hand on the operator’s own keys, under a spending cap; it never runs in CI. It asks every decider every request, passes each answer through the same checks RISE runs for readers, and freezes the capture into one file named by the hash of its bytes. RISE’s report command re-derives every score from that file deterministically. This page copies that report and the slim replay file, and CI checks both against the hashes below.
Reading the run file…
The known-probability controls test whether a decider puts stated odds on stated options. They say nothing about whether its reading choices are good. Frozen results describe each model id on its capture date; a later model is a new run, never an edit.
Terms: OpenAI’s Sharing & publication policy “welcome[s] research publications related to the OpenAI API”, and its beta terms say “as-is” for “testing and evaluation”.
06 / DISCLOSURE
No partnership with, sponsorship by, or endorsement from TypeSafe AI, OpenAI, or Kev’s author. Names identify the APIs we called. Captured on the date in the run file from RISE at the commit in the run file, with the model ids in the run file; results may differ today. The fixed requests and menus were written by RISE’s authors and tuned on Jev.
Want deciders like these tested on your list?
We run interactive reading pilots with publishers. Tell us about a rights-cleared text you would like readers to hear and see.
See the publisher pilot