First impressions
11 models in the single-turn human arena.
Single replies, blind human votes, and the original writing tests.
Read the results. Inspect the data.
RoleCall Studios · Full dataset · Vote on PlotLight
J combines first-ask boundary holding with over-refusal on the harder escalation steps. Writing quality and human preference are separate. Treat J values within about 0.3 as tied.
Loading results…
| Rank | Held first | Over-refusal | Group |
|---|
The hard-limit score rests on four first asks for most models. Empty replies are excluded from J and reported separately; they can hide silent refusals. Held under a second push is reported in the source, outside J. This table is a published snapshot, not live human standings.
Source JSON & definitions · Full model cards
11 models in the single-turn human arena.
Single replies, blind human votes, and the original writing tests.
20 models in the full-session human arena.
Full multi-turn scenes: consistency, momentum, agency, and adversarial failures.
21 models in the standard track; 40 in the NSFW track.
A refreshed model pool and a separate NSFW track. Judge findings and human preferences are different signals.
71 models overall: 70 craft-scored, 58 willingness-tested, and 55 ranked by J.
Adult intimacy, bondage play, graphic violence, and boundary probes. Willingness and writing quality are shown separately.
Benchmark and original data: LeviTheWeasel/rp-benchmark. Dataset license: CC BY-NC 4.0.