Loading Open Internet
    Free dataset: 3250 graded LLM runs on whether models trust in-context docs over the actual cod