Open Internet by MindsNet
Fine-tuning LLMs for Open-Ended Math Problems
Developing an LLM that can solve open-ended math problems, such as proof-only problems, requires a more nuanced approach than traditional methods like RLVR, SFT, GRPO, or PPO. The lack of a suitable reward function and the limitations of existing fine-tuning methods hinder progress. A new approach to fine-tuning LLMs is needed to effectively tackle these complex math problems.
Computing & Technology, Computer Science, Machine Learning