GitHub
Open-AgentRL
观察池 · 暂无正式排名
该项目当前未满足“两个有效维度 + 两种数据源”的主榜门槛。
— 未排名
可观测采用度缺失 —
动量当前有效 · 2026-09-12 6
关注度当前有效 · 2026-09-12 0
信号可信度 依据当前有数据的独立评分维度数量计算。
中2/3 · 1 种数据源
项目介绍
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
An overview of our research on RLAnything.
In this work, we propose RLAnything, a reinforcement learning framework that dynamically optimizes each component through closed-loop optimization, amplifying learning signals and strengthening the overall system:
An overview of our research on agentic RL.
In this work, we systematically investigate three dimensions of agentic RL: data, algorithms, and reasoning modes. Our findings reveal:
We also contribute high-quality SFT and RL datasets, demonstrating that simple recipes enable even 4B models to outperform 32B models on challenging benchmarks including…
各数据源
635 Star
- Star 635
- Fork 59
- 提交 30
- 发布 0