Open Internet by MindsNet
Determining Sample Size for Non-Inferiority Test in Human-LLM Classification Tasks
The author is struggling to determine the sample size for a non-inferiority test comparing the performance of a large language model (LLM) to human experts in a classification task. The task involves 9 categories and uses Cohen's Kappa to measure inter-rater reliability (IRR). The author needs to validate the LLM against two human experts but faces challenges in finding literature on expected Kappa values for sample size calculations.
Computing & Technology, Computer Science, Machine Learning