SearchLM NL2BM25: teaching Qwen2.5-3B to generate Tantivy boolean queries via SFT + GRPO. Covers reward hacking (GRPO v1) and the shaped-reward fix (GRPO v2). Supreeth/searchlm-nl2bm25-sft Text Generation • 3B • Updated about 1 month ago • 7 Supreeth/searchlm-nl2bm25-sft-v2 Text Generation • 3B • Updated about 1 month ago • 14 Supreeth/searchlm-nl2bm25-grpo Text Generation • 3B • Updated about 1 month ago • 16 Supreeth/searchlm-nl2bm25-grpo-v2 Text Generation • 3B • Updated about 1 month ago • 13
SearchLM NL2BM25: teaching Qwen2.5-3B to generate Tantivy boolean queries via SFT + GRPO. Covers reward hacking (GRPO v1) and the shaped-reward fix (GRPO v2). Supreeth/searchlm-nl2bm25-sft Text Generation • 3B • Updated about 1 month ago • 7 Supreeth/searchlm-nl2bm25-sft-v2 Text Generation • 3B • Updated about 1 month ago • 14 Supreeth/searchlm-nl2bm25-grpo Text Generation • 3B • Updated about 1 month ago • 16 Supreeth/searchlm-nl2bm25-grpo-v2 Text Generation • 3B • Updated about 1 month ago • 13
Running Repro: XRPO: Pushing the Limits of GRPO with Targeted Exploration and Exploitation 🎯 Collaborate on a research logbook with an AI coding agent
Sleeping RL VeriRL — Verilog RTL Design Environment 🔬 Step through a Verirl environment by sending actions and view results