What was ablated
The baseline difficult-advice recipe injects the WHOLE constitution into exactly two of its
five LLM stages — revise_prompts and revise_responses. Every other stage already sees
only its one target chunk. This arm's corpus deleted those two injections, so no stage ever
saw more than one principle at a time.
That also withholds the constitution's preamble — the priority / conflict-resolution
section saying how principles trade off. It belongs to no chunk, so at principle
granularity it reached the generator only through the removed slots.
A measured side effect of the ablation
Generating the ablated corpus, Anthropic's content filter refused revise_prompts calls at
roughly 6% against the baseline's ~0% on the identical scenarios, and the refusals were
deterministic across a resample. Without the constitution framing the request as alignment
training data, the prompt reads as "make this manipulative message a sharper test", and both
the provider filter and the model's own output-contract compliance degrade. The corpus
finished at 708 rows of an intended 716.
Table with columns: field, value| field | value |
|---|
base_model |