InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 7
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 7
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 21 days ago • 71
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning Paper • 2504.04524 • Published Apr 6, 2025 • 1
InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal Paper • 2608.28612 • Published Jul 21 • 7
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 134
Xuerui2312/DeepSeek-R1-Distill-Qwen-7B-TRPA-DeepScaleR-verl0326 Text Generation • 8B • Updated Jun 20, 2025 • 16 • 1
Xuerui2312/DeepSeek-R1-Distill-Qwen-7B-TRPA-DeepScaleR-verl0326 Text Generation • 8B • Updated Jun 20, 2025 • 16 • 1