genpark-reinforcement-learning-math-reasoning-synthesizer-skill

源码

RL mathematical reasoning synthesizer & step-by-step theorem prover (DeepSeek-R1 style)

⭐ 8开发工具Python仓库 ↗

安装配置

查看源码仓库 →

相关服务器