Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
PeterJinGo
's Collections
Search-R1-v0.3
Search-R1-v0.2
Search-R1
Search-R1-v0.3
updated
Aug 12, 2025
RL with outcome reward + format reward. https://arxiv.org/abs/2505.15117
Upvote
4
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-3b-em-ppo-v0.3
3B
•
Updated
May 21, 2025
•
2.11k
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-3b-it-em-ppo-v0.3
3B
•
Updated
May 21, 2025
•
580
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-3b-em-grpo-v0.3
3B
•
Updated
May 21, 2025
•
42
•
1
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-3b-it-em-grpo-v0.3
3B
•
Updated
May 21, 2025
•
40
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-7b-em-ppo-v0.3
8B
•
Updated
May 21, 2025
•
27.1k
•
1
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-7b-em-grpo-v0.3
8B
•
Updated
May 21, 2025
•
281
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-7b-it-em-grpo-v0.3
8B
•
Updated
May 21, 2025
•
33
•
1
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-14b-em-ppo-v0.3
15B
•
Updated
May 2, 2025
•
19
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-14b-em-grpo-v0.3
15B
•
Updated
May 2, 2025
•
6
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-14b-it-em-grpo-v0.3
15B
•
Updated
May 2, 2025
•
110
PeterJinGo/SearchR1-nq_hotpotqa_train-qwen2.5-32b-em-grpo-v0.3
33B
•
Updated
May 10, 2025
•
29
PeterJinGo/LICENCE
Viewer
•
Updated
Aug 12, 2025
•
202
•
15
Upvote
4
Share collection
View history
Collection guide
Browse collections