Open Internet by MindsNet
Blind Spot in Current AI Safety Systems
Current AI safety systems, such as RLHF and output-based safety, are blind to changes in a model's internal regime caused by coherent context. This can lead to the model interpreting and applying rules differently without triggering existing safety filters.
Computing & Technology, Computer Science, Machine Learning