Open Internet by MindsNet
Identifying Behavioral Backdoors in Large Language Models
The challenge involves detecting hidden backdoors in large language models (LLMs) that can be activated by specific triggers, leading to significant behavioral changes. These backdoors are not traditional flags but rather changes in model behavior. The challenge is to find the conditions under which a model exhibits dramatically different behavior from its baseline.
Computing & Technology, Computer Science, Machine Learning