Loading Open Internet
    LLMs solve about 1 in 3 real root-cause cases on a realistic benchmark. Mostly wrong on the hard ones.