Renxi Wang

4 posts

Renxi Wang

Renxi Wang

@realReasonWang

PhD Student at MBZUAI Natural Language Processing & LLM Agent

Katılım Temmuz 2022
21 Takip Edilen5 Takipçiler
Renxi Wang retweetledi
Qiyuan Xu
Qiyuan Xu@XeroEssential·
Introducing AoA (Agent over Abstract Syntax Trees), our proof agent built on Isabelle/HOL. 🎯 AoA achieves 99.6% on miniF2F, 89.2% on NTP4VC-Pearl, and 97.7% on NTP4VC-realC, while delivering a 2.3–4.7x reduction in API cost, a 2.9–6.9x reduction in token consumption, and 1.4–2.0x faster total execution time compared to Amazon's Isabelle Agent — which is likewise built on Isabelle and equipped with Sledgehammer. These results are powered by two key innovations: ✨ AoA abstracts away from the concrete syntax of proof languages, representing proofs as abstract syntax trees via a JSON schema. This enables the first effective proof agent on a freshly redesigned language (Isabelle/Minilang) that commercial LLMs have had little exposure to — suggesting that the LLM era, far from stifling new languages, can actually accelerate their development. ✨ AoA also abandons the traditional agent interaction paradigm of source-code editing with line-number indexing, adopting a novel tree-editing model that eliminates the line-number drift issues that conventional agents often struggle with. 📄 Paper: arxiv.org/abs/2607.16372 🦾 Source code: github.com/xqyww123/Isa-M… 🙏 Huge thanks to my amazing co-authors @joshuaongg21 , @realReasonWang , @WendaLi8 , @haonanlp , Luke Ong, and @conrad_watt — it has been a true honor working with you all.
Qiyuan Xu tweet mediaQiyuan Xu tweet media
English
1
5
5
176
Renxi Wang
Renxi Wang@realReasonWang·
🚀 Excited to share 🪽AgentFly — our open-source framework for training LLM agents with reinforcement learning that’s built to scale and stays easy to extend. 𝗙𝗲𝗮𝘁𝘂𝗿𝗲𝘀 • 🔁 Multi-turn, multi-chain RL (e.g. GRPO / Reinforce++) with 🎯 token-level masking for stable training • ⚡ Asynchronous tools & reward execution + 🧠 centralized resource manager for high-throughput rollouts • ⚙️ Predefined tools/environments including code interpreter, search, WebShop, ScienceWorld, etc. • 🛠️ Decorator-based APIs to plug in custom tools & reward functions in minutes • 🧩 Decoupled design makes programming and extending new features effortless • 🖼️ Built-in chat templates and 👀 vision-language (multi-modal) agent training 𝗗𝗶𝘃𝗲 𝗶𝗻 • 📄 Report: arxiv.org/pdf/2507.14897 • 📈 W&B Curves: wandb.ai/AgentRL/Open • 💻 Repo: github.com/Agent-One-Lab/… I’d love your feedback—⭐ the repo, try the quickstart, and open issues/PRs. Let’s push agentic RL forward together! ✨ #LLMAgents #ReinforcementLearning #RLAIF #OpenSource #AI #AgentFly #AgentRL #Agent #LLM #LLMAgent #NLP #Agentic #reward #reasoning #code #tool #feedback #mcp #VisionLanguage #MultiModal #rollout #MultiTurn #toolgen #ToolUse #vllm #AgenticRL
English
0
0
0
21
Renxi Wang retweetledi
Haonan Li
Haonan Li@haonanlp·
We are excited to announce Libra-Leaderboard: The first LLM leaderboard dedicated to balancing safety and capability in LLMs. As AI advances, ensuring its safety becomes more critical than ever. By prioritizing safety measurement, we aim to inspire the AI community to make safety a core factor in model evaluation and development, alongside capability.🚀 #LLM #AI #AISafety #ResponsibleAI Website: leaderboard.librai.tech/LeaderBoard Paper: arxiv.org/abs/2412.18551 Thanks to all the contributors: @han_xudong @_ZenanZhai @honglin_mu @_HaoWang @_ZhenxuanZhang @yilin_link @ShomLinEd @realReasonWang @ArtemShelmanov @xiangyuqi_pton @_YuxiaWang @_DonghaiHong @youliang_yuan @_MengChen @HaoqinT @FajriKoto @ttk_kuribayashi @HaoqinT @rishabh15 @BingchenZhao @yawen_duan @_YiLiu @_YaodongYang @dongyp13 @soujanyaporia @stefan_fee @waterluffy @RenHector @Emad_A_Alghamdi @IGurevych @preslav_nakov @monojitchou @eltimster
Haonan Li tweet media
English
1
15
31
3.9K
Renxi Wang
Renxi Wang@realReasonWang·
RT @llm_sec: Against The Achilles' Heel: A Survey on Red Teaming for Generative Models 🌶️ "Our extensive survey, which examines over 120 p…
English
0
3
0
12