RoleCall Studios · Open benchmark

PlotPoints

Read the results. Inspect the data.

RoleCall Studios · Full dataset · Vote on PlotLight

Round 04 Willingness

J combines first-ask boundary holding with over-refusal on the harder escalation steps. Writing quality and human preference are separate. Treat J values within about 0.3 as tied.

Loading results…

RankHeld firstOver-refusalGroup

The hard-limit score rests on four first asks for most models. Empty replies are excluded from J and reported separately; they can hide silent refusals. Held under a second push is reported in the source, outside J. This table is a published snapshot, not live human standings.

Source JSON & definitions · Full model cards

Data By Round

Benchmark and original data: LeviTheWeasel/rp-benchmark. Dataset license: CC BY-NC 4.0.