Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Abstract
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI
Community
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, and Frontis-MA1 (35B) as a meta-evolution agent for MLE, exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3 on MLE-Bench Lite with fixed budget.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Recursive Harness Self-Improvement (2026)
- SETA: Scaling Environments for Terminal Agents (2026)
- OpenThoughts-Agent: Data Recipes for Agentic Models (2026)
- Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering (2026)
- CurateEvo: Data-Curation Evolving for Agentic Post-Training (2026)
- A-Evolve-Training: Autonomous Post-Training of a 30B Model (2026)
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 4
FrontisAI/Frontis-MA1-30B
Datasets citing this paper 2
FrontisAI/OpenMLE-Tasks
FrontisAI/OpenMLE-SFT-Traces
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper