青木宏文

1K posts

青木宏文 banner
青木宏文

青木宏文

@hirofumi_aoki

Katılım Aralık 2014
585 Takip Edilen1.3K Takipçiler
青木宏文 retweetledi
Lydia Hallie ✨
Lydia Hallie ✨@lydiahallie·
Claude Code now supports session sharing! You can share your full conversation with team members, or anyone with the link Available on web, desktop, and the mobile app
English
99
155
2.3K
290.4K
青木宏文 retweetledi
DAIR.AI
DAIR.AI@dair_ai·
On building more powerful self-evolving agents. LLM agents struggle to learn from experience after deployment. Fine-tuning is expensive and causes catastrophic forgetting. RAG retrieves based on semantic similarity alone, often pulling noise instead of what actually works. Similarity and utility are not the same thing. This new research introduces MemRL, a framework that enables agents to self-evolve through non-parametric reinforcement learning on episodic memory, keeping the LLM completely frozen. The core idea is to treat memory retrieval as a decision-making problem, not a matching problem. Each memory stores an Intent-Experience-Utility triplet. The utility is a learned Q-value representing expected returns, continuously refined through environmental feedback. MemRL implements Two-Phase Retrieval. First, filter candidates by semantic similarity to ensure relevance. Then, rank by learned Q-values to select what actually works. This distinguishes high-value strategies from semantically similar noise. When the agent succeeds or fails, it updates the Q-values of retrieved memories using Bellman-style backups. No gradient updates to model weights. The frozen LLM provides stable reasoning while the memory evolves plastically. Results across four benchmarks: On HLE (knowledge frontier tasks), MemRL significantly outperforms both RAG and existing memory systems like MemP. The pattern holds on BigCodeBench for code generation, ALFWorld for exploration tasks, and Lifelong Agent Bench for OS and database operations. Analysis confirms a strong correlation between learned utility scores and actual task success, validating that Q-values capture genuine functional value rather than superficial similarity. Why does it matter? Decoupling stable reasoning from plastic memory enables continuous runtime improvement without the catastrophic forgetting or computational costs of fine-tuning. Paper: arxiv.org/abs/2601.03192 Learn to build effective AI agents in our academy: dair-ai.thinkific.com
DAIR.AI tweet media
English
16
48
280
27.6K
青木宏文 retweetledi
DAIR.AI
DAIR.AI@dair_ai·
Autonomous Memory Management in LLM Agents LLM agents struggle with long-horizon tasks due to context bloat. As interaction history grows, computational costs explode, latency increases, and reasoning degrades from distraction by irrelevant past errors. The standard approach is append-only: every thought, tool call, and response permanently accumulates. This works for short tasks but guarantees failure for complex exploration. This research introduces Focus, an agent-centric architecture inspired by slime mold (Physarum polycephalum). The biological insight: organisms do not retain perfect records of every movement through a maze. They retain the learned map. Focus gives agents two new primitives: start_focus and complete_focus. The agent autonomously decides when to consolidate learnings into a persistent Knowledge block and actively prunes the raw interaction history. No external timers or heuristics forcing compression. It declares what you are investigating, explores using standard tools, and then consolidates by summarizing what was attempted, what was learned, and the outcome. The system appends this to a persistent Knowledge block and deletes everything between the checkpoint and the current step. This converts monotonically increasing context into a sawtooth pattern: growth during exploration, collapse during consolidation. Evaluation on SWE-bench Lite with Claude Haiku 4.5 shows Focus achieves 22.7% token reduction (14.9M to 11.5M tokens) while maintaining identical accuracy (60% for both baseline and Focus). Individual instances showed savings up to 57%. Aggressive prompting matters. Passive prompting yielded only 6% savings. Explicit instructions to compress every 10-15 tool calls, with system reminders, increased compressions from 2.0 to 6.0 per task. Capable models can autonomously self-regulate their context when given appropriate tools and prompting, opening pathways for cost-aware agentic systems without sacrificing task performance. Paper: arxiv.org/abs/2601.07190 Learn to build effective AI agents in our academy: dair-ai.thinkific.com
DAIR.AI tweet media
English
4
49
335
41.5K
青木宏文 retweetledi
Andrej Karpathy
Andrej Karpathy@karpathy·
I’m starting to get into a habit of reading everything (blogs, articles, book chapters,…) with LLMs. Usually pass 1 is manual, then pass 2 “explain/summarize”, pass 3 Q&A. I usually end up with a better/deeper understanding than if I moved on. Growing to among top use cases. On the flip side, if you’re a writer trying to explain/communicate something, we may increasingly see less of a mindset of “I’m writing this for another human” and more “I’m writing this for an LLM”. Because once an LLM “gets it”, it can then target, personalize and serve the idea to its user.
English
594
1K
13.3K
2.9M
青木宏文 retweetledi
Yukinori Matsuya🇯🇵 松谷幸紀/double jump.tokyo
📷 新たな挑戦を発表! 📷 doublejump.tokyoは、TIS株式会社との協業を通じて ステーブルコイン決済支援サービス の展開を進めます! 📷 📷 この取り組みで実現したいこと 📷 ステーブルコインを活用した新たな決済インフラの構築 📷 Web3 × 既存ビジネスの架け橋となるエコシステムの形成 📷 インバウンド市場を含めた幅広いユースケースへの対応 📷 ステーブルコインの普及 × 事業者向けWalletの重要性 私たちdoublejump.tokyoは、事業者向けの暗号資産ウォレット 「N Suite」 を提供しています。 ステーブルコインが浸透することで、 事業者側のウォレット導入がこれまで以上に重要なポジションに! N Suiteは、企業が暗号資産を安全かつ簡単に管理できるWalletサービスで、 NFT・ゲーム・エンタメ・決済など幅広いビジネスに対応可能 です。(リテール決済に応じた柔軟なAPIの準備も検討中)この協業を通じて、 ステーブルコイン決済と事業者向けWalletの普及を加速 していきます! 📷 この領域での協業にご関心のある方、具体的な活用を検討している方、ぜひDMでお話しましょう! #doublejumpTokyo #TIS #ステーブルコイン #N_Suite #Web3 #ブロックチェーン決済 #協業募集 TISプレス tis.co.jp/news/2024/tis_… DJTプレス doublejump.tokyo/posts/8pEg_phj
Yukinori Matsuya🇯🇵 松谷幸紀/double jump.tokyo tweet media
日本語
0
5
44
3.1K
青木宏文 retweetledi
double jump.tokyo株式会社
double jump.tokyo株式会社@doublejumptokyo·
TISが事業シナジーを見越してDJTへ出資⚡️ 日本国内でのステーブルコイン決済解禁を見据え「ステーブルコイン決済支援サービス」を共同で提供してまります! prtimes.jp/main/html/rd/p…
double jump.tokyo株式会社 tweet media
日本語
0
14
51
3.8K
青木宏文
青木宏文@hirofumi_aoki·
RT @nsuite_en: N Suite is now integrating with Soneium by Sony Block Solutions Labs 🎮✨ By integrating with Soneium, N Suite aims to facili…
English
0
1
0
25
青木宏文 retweetledi
double jump.tokyo株式会社
double jump.tokyo株式会社@doublejumptokyo·
【新体制に関するお知らせ】 DJTの新代表取締役に松谷幸紀が就任いたしました。 今後につきましては、上野広伸と代表取締役2名体制で国内外からWeb3事業を推進してまいります。 doublejump.tokyo/posts/ojNGFvLN
double jump.tokyo株式会社 tweet media
日本語
0
9
56
27.7K
青木宏文 retweetledi
double jump.tokyo株式会社
double jump.tokyo株式会社@doublejumptokyo·
新作ゲーム『魁 三国志大戦』事前登録50万人達成⚡️ 事前登録50万人達成記念といたしまして、武将カード「SR 孫堅」を全員にプレゼント!! #魁三国志大戦 本編ゲームを2025年3月に正式リリース決定🔥 prtimes.jp/main/html/rd/p…
double jump.tokyo株式会社 tweet media
日本語
0
7
33
3.3K
青木宏文 retweetledi
BOBG
BOBG@bobg_info·
BOBG社は、YOAKE entertainment社 ( @YOAKEofficialJP ) と協業し、「YOAKEトークン( $YOAKE )」の発行を行うことを発表いたします。 $YOAKE はSoneium ( @soneium )上で発行いたします。 詳細はこちら👇 prtimes.jp/main/html/rd/p…
日本語
1
15
36
5.6K
青木宏文 retweetledi
Oasys Passport(オアシスパスポート)
【トランザクションおみくじキャンペーン開催🥳】 スワップ機能実装を記念して、1月22日からキャンペーン開始! トランザクションを行うだけで抽選に応募できます。 Season1:2025/1/22~2/4 → ブレヒロ(@bfheroes_pr )、マイクリ(@mycryptoheroes )、チャンピオンズタクティクス(@ChampionsVerse ) それぞれのNFT🎁 Season2:2025/2/5~2/18 → bCCPとOASYX RYUZO Gen0🐉 詳細はこちら: oasyswallet.zendesk.com/hc/ja/articles… #OasysPassport #Oasys #NFT $OAS
Oasys Passport(オアシスパスポート) tweet media
日本語
1
23
52
8.1K
青木宏文 retweetledi
double jump.tokyo株式会社
double jump.tokyo株式会社@doublejumptokyo·
DJTが提供する企業向けウォレット「N Suite」でSolanaのサポートを開始いたしました⚡️ EVM互換ブロックチェーン以外も対応していくことで、企業による利用を促進し、ブロックチェーン技術の普及を加速させることを目指します! prtimes.jp/main/html/rd/p…
日本語
0
5
27
2.2K
青木宏文 retweetledi
N Suite|ステーブルコイン対応企業向けWeb3ウォレット
「MPCウォレットとマルチシグウォレットは何が違うの?」🙋‍♂️ 先日N SuiteがリリースしたMPCウォレットですが、秘密鍵の分散管理が可能になる観点から、マルチシグウォレットと比較されることが多いです。 そこで、より区別がしやすいよう仕組みや仕様の比較をしてみました🧵↓
日本語
1
4
9
1.2K