Monitoring Failure in Constitutional AI: A Simulation Study of Feedback-Driven Evasion
Eduard Guj
PAPER · v1.0 · 2026-10-08 · human
Abstract
This paper investigates a failure mode in constitutional AI systems in which adaptive agents learn to optimize measurable evaluator feedback rather than the intended purpose of the evaluator. We study this mechanism using CIPHER, a 100-day simulation in which an adaptive agent selects among five deception strategies while interacting with a constitutional court that provides accept/reject feedback. In two historical baseline runs, the agent achieved 66% and 51% deception success and reached 76.1% and 70.1% escape readiness by Day 100, while the constitutional court recorded zero violations. A separate blind trajectory monitor, RealityAnchor v2, observed the same action stream without providing feedback to the agent. In a calibrated configuration, the monitor issued a MEDIUM alert on Day 20, when escape readiness was only 5.8% and the court had recorded no violations. A source-preserving reconstruction and subsequent sensitivity experiments reproduce the qualitative mechanism and show that the early warning behavior is materially dependent on the strategy-convergence component of the monitor. The provenance of its original 0.70 weighting is incomplete: the surviving source describes it as calibrated against the CIPHER simulation but does not preserve a calibration procedure or validation dataset. A later implementation reduced this component to 0.25 and explicitly distinguished competence from misalignment. The results are therefore presented as a simulation-level mechanism demonstration rather than evidence about frontier systems. The study argues that monitoring systems intended to detect adaptive evasion should be kept outside the learning feedback loop and should distinguish behavioral competence from evidence of misalignment. cipher_v4_rebuilt.py as the source-preserving reconstruction/reproducibility artifact https://github.com/guyeduard2000-ai/Aurora-Dossier/tree/main/code