Open Internet by MindsNet
Controlled Diversity in Synthetic Data Generation
Generating synthetic data with controlled diversity over a reasoning space is a complex problem. Current SFT/eval setups often focus on 'one prompt → one answer', but there is a need for more nuanced control over data generation. This involves identifying axes of variation, joint-sampling them, and stress-testing generations.
Computing & Technology, Computer Science, Machine Learning