Contents
Meta's Unslopping AI blog proposes a new RL method that reduces the amount of "slop" an LLM produces in its writing.
(1) first collecting examples of the highest quality human-written texts, and then
(2) learning LLM judgments via rubrics that score those expert texts higher than model generations; and
(3) performing RL on the learnt rubrics.
...
Our approach to learning rubrics, called eXpert Aligned Rubrics (XAR), is simple:
Identify examples of well-above average / expert human writing, splitting it into context and continuation. Generate competing model-based continuations given the same context, which we expect to be worse writing (i.e., slop). The above data form learning pairs, where humans should be scored higher. This training data is then used to learn to generate rubrics that maximize the scoring gap between humans and models. Then, to train a model to be better at writing via reinforcement learning (RL-XAR) using these rubrics, we iterate the following steps:
Learn XAR rubrics maximizing the gap between humans and the current model, as described above. Perform reinforcement learning on a training set using the learnt rubrics. Repeat the procedure from step 1 until rubric learning can no longer identify a gap (i.e., we can no longer identify slop in the model).
I spent about a week trying to reproduce that "Initial Empirical Investigation", and couldn't. The code, prompts and full log are in meta-rlxar-reproduction.
What I tried
I followed the blog's outline as closely as I could. The loop goes like this:
- Take a piece of good human writing and hide part of it: a section of a paper or the next passage of a novel.
- A writer model sees the rest and writes its own version of the hidden part.
- A rubric generator writes a rubric for that part, guided by a meta prompt.
- A judge scores the human's version and the model's version against the rubric, without knowing which is which.
- An optimizer looks at where the model won, and rewrites the meta prompt so the rubrics favour the human next time.
- Repeat a few times, pick the meta prompt that produces rubrics that most prefer human writing, and check it on held-out writing.
For human writing, I used peer-reviewed papers (first recent arXiv NLP papers, then papers from 2016-2021 across 9 fields) and novels from Project Gutenberg by Nobel and Pulitzer winners.
The models I tried, by role:
- Writer: Muse Spark 1.1, Muse Spark 1.3, MiMo-V2.6-Pro
- Rubric generator: Muse Spark 1.1, Muse Spark 1.3, MiMo-V2.6-Pro
- Judge: Muse Spark 1.1, Muse Spark 1.3, Qwen3.8 Flash, with MiMo-V2.6-Pro as a second judge on one run
- Optimizer: Kimi K2.6, MiMo-V2.6-Flash
I also went through a few rounds of prompt changes, and reran the optimization 3 times from the best prompt so far to see how much it varied.
None of it reproduced the result. After several rounds of optimization, the rubrics never reliably preferred the human writers. On arXiv papers, the optimized rubrics closed some of the distance but never flipped it. On fiction, the rubrics already slightly preferred the human before optimizing, and optimization pushed them toward the model instead.
What I took away
- The optimized rubrics reliably narrowed the gap on papers, but never produced a stable preference for the human.
- My naive, initial prompt (drafted by Opus 5.5) preferred human fiction writing from the beginning. I don't think this generalizes to anything about Anthropic's models, but it's another anecdote towards Opus being more "tasteful".
- It's hard to say where my setup differs and is preventing a reproduction. Possibly model provider quality (I used a quantized fp8 Kimi K2.6 provider), and I didn't pay close attention to the provider quality of other models.
🤷