Loading Open Internet
    Anthropic's new interpretability tool found Claude suspects it is being tested in 26% of benchmarks and never says so