OpenSquilla

266 posts

OpenSquilla banner
OpenSquilla

OpenSquilla

@OpenSquilla

Token-Efficient AI Agent Intelligence. ✨ Smart routing · 🧠 Human-like memory · 🛡️ Sandbox Apache 2.0 · open source

California, USA Katılım Nisan 2026
7 Takip Edilen6.1K Takipçiler
Sabitlenmiş Tweet
OpenSquilla
OpenSquilla@OpenSquilla·
OpenSquilla v0.5.1 is live. 🦐 New models drop every week — same question every time: which one for *this* task? So we stopped hard-wiring "which model" to "how you use it." What shipped: ▸ Multi-provider access — more providers on TokenRhythm, free credits on signup ▸ Access / routing split — swap providers by swapping a key, workflow never moves ▸ Smart routing — light model for simple tasks, flagship for hard, several in parallel for critical (or pin one yourself) ▸ Attach any file — PDF, Excel, Word, code. Drag it in ▸ Project spaces — contexts stop bleeding together ▸ Plan — long tasks resume right where they stopped ▸ Import Memory from other AI tools — no re-onboarding ▸ Rebuilt UI — collapsible chats, generated docs pinned right Ensemble is now standard, not a preview. On DRACO (100 hard tasks): 🦐 OpenSquilla — 64.09 / $0.12 per task GPT-5.6-sol — 63.99 / $1.71 per task Same score. ~14x cheaper. Eval script open. Apache 2.0 · self-hostable · macOS + Windows. → opensquilla.ai All models. One mind.
OpenSquilla tweet media
English
24
19
33
112.6K
OpenSquilla
OpenSquilla@OpenSquilla·
If the reports are accurate, Astra’s biggest shift isn’t just a stronger model—it’s AI teams. But teams need orchestration: → Who handles each task?→ What context do they share?→ When should they escalate?→ Who verifies the result? That’s the harness layer we’re building with OpenSquilla:github.com/opensquilla/op… 🦐
English
0
0
0
2
Lumina
Lumina@LuminaXspace·
🚨New OpenAI Model family has Leaked OpenAI have demoed "Astra" to US policymakers this week: • Built for long running, complex tasks • Much more AI Agent oriented (teams) • Tested on projects and advanced maths • No release date confirmed • Could become GPT 6 Astra looks to be a shift from chatbots to autonomous AI teams, I can't wait to see what it's capable especially with the hacks from the hugging face incident. Are you ready for GPT 6?
Lumina tweet media
English
7
8
78
5.6K
OpenSquilla
OpenSquilla@OpenSquilla·
Microsoft’s Project Perception can find vulnerabilities, prioritize them, and generate patches. The bigger question: What evidence, approval, and rollback exist before deployment? Autonomous action should be reversible. #AI #CyberSecurity #AIAgents
OpenSquilla tweet media
English
0
0
0
8
OpenSquilla
OpenSquilla@OpenSquilla·
A clean final answer can hide a terrible agent run. Measure more than output quality: → Retries → Manual interventions → Failed tool calls → Unnecessary tokens → Whether the result was verified The path matters too. 🦐 #AI #AgenticAI #LLMOps
GIF
English
0
0
0
24
OpenSquilla
OpenSquilla@OpenSquilla·
@arena @deepseek_ai Exactly why routing shouldn’t be hardcoded from leaderboards. Real performance depends on the task, context, tools, and retry cost—not model size or vendor. If V4 Flash is the best fit for a turn, the harness should recognize it. That’s OpenSquilla: github.com/opensquilla/op… 🦐
English
0
0
0
1.3K
Arena.ai
Arena.ai@arena·
DeepSeek-V4-Flash-High by @deepseek_ai is #7 overall in the Frontend Code Arena with 1,586 pts! It’s #3 among open. In categories it’s #4 in Consumer Product, #6 in Reference-based Design, Data & Analytics, Gaming, and #7 in Brand & Marketing. This is an impressive improvement from DeepSeek-V4-Flash-High-Preview, a +154 pt jump. Even more impressive is the jump from DeepSeek-V4-Pro-Preview, a +121 point improvement.
Arena.ai tweet media
Arena.ai@arena

Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with a score of 1586! Priced at $0.14/$0.28 per MToken, it’s the best performance-per-dollar of any model in its class. Congrats to the @deepseek_ai team!

English
73
160
1.6K
164.8K
OpenSquilla
OpenSquilla@OpenSquilla·
Exactly why routing shouldn’t be hardcoded from leaderboards. Real agent performance depends on the task, context, tools, and retry cost—not model size or vendor. OpenSquilla routes from harness signals, so V4 Flash gets the turn when it actually fits. github.com/opensquilla/op… 🦐
English
0
0
0
71
BridgeMind
BridgeMind@bridgemindai·
DeepSeek v4 Flash is actually pretty good. Is it as good as the benchmarks claim? No. It is benchmaxxed like everything else shipping right now. But it is a 284B parameter model. At that size it has no business performing this well. It beats GPT 5.6 Luna in real use. A model a fraction of the size, from a lab a fraction of the budget. DeepSeek keeps doing more with less than anyone in the world. Very impressed.
BridgeMind tweet media
English
66
19
691
44K
OpenSquilla
OpenSquilla@OpenSquilla·
The weekend test for an AI agent: Give it a long task on Friday. Return on Monday. Does it remember the goal, show its progress, recover from failures, and explain what remains? OpenSquilla v0.5.1 keeps long tasks moving with persistent Plans. 🦐 github.com/opensquilla/op… #OpenSourceAI #AIAgents
OpenSquilla tweet media
English
0
0
0
16
OpenSquilla
OpenSquilla@OpenSquilla·
Microsoft BitNet makes local CPU inference much more practical. But local models don’t replace frontier models. They change the routing equation: → Routine: local→ Complex: frontier→ Critical: multi-model Use the right intelligence for each turn. 🦐 #AI #Microsoft #AIAgents
OpenSquilla tweet media
English
0
0
0
320
OpenSquilla
OpenSquilla@OpenSquilla·
🦐 Today in AI: 📦 GitHub Models officially retired. 🔌 Only 37.2% of studied MCP apps require approval before tool execution. 💻 Microsoft’s BitNet is putting 100B-class local CPU inference back in the conversation. The runtime layer keeps getting more important. #AI #AIAgents #MCP
OpenSquilla tweet media
English
0
0
0
232
OpenSquilla
OpenSquilla@OpenSquilla·
Small models need smaller, clearer jobs. Let Luna or DeepSeek Flash handle scoped turns. Escalate ambiguity, debugging, and critical decisions to stronger—or multiple—models. The harness should decide before the code becomes a mess. That’s OpenSquilla: github.com/opensquilla/op… 🦐
English
0
0
0
12
Justin
Justin@JustinGorya·
OpenAI dropped the pricing of GPT-5.6 Luna by 80%. They already did something similar back in the day with the o3 model. Today, DeepSeek V4 Flash GA was released and also has an incredible price-performance ratio. These models are really good, but you need to be careful about asking them things they have no idea about. Unfortunately, in general, the GPT models have a high hallucination rate. They are very good if you have a plan and you actually know what to do. Don’t let Luna and DeepSeek V4 Flash run on their own without a clue. There is a pretty high chance that you will get a big code mess after the run is completed. But don't miss these small models out. They are GOOD and usefull.
Justin tweet media
English
17
5
96
17.2K
OpenSquilla
OpenSquilla@OpenSquilla·
@Oluwaphilemon1 This changes the routing equation: → Routine tasks: local CPU models→ Complex tasks: frontier models→ Critical tasks: multi-model collaboration OpenSquilla is an open-source Agent Harness built to choose between them at the turn level. github.com/opensquilla/op… 🦐
English
0
0
1
320
FHILY👑
FHILY👑@Oluwaphilemon1·
Opus 5, Kimi K3, Claude Fable 5, and GPT-5.6 are dead… Microsoft has changed the game 🤯 They've open-sourced bitnet.cpp, a 1-bit LLM inference framework. It let's you run 100B parameter models on your local CPU without GPUs. - 6.17x faster inference - 82.2% less energy on CPUs 100% Open Source.
FHILY👑@Oluwaphilemon1

x.com/i/article/2082…

English
8
7
47
10.9K
OpenSquilla
OpenSquilla@OpenSquilla·
@haider1 This is where turn-level routing gets interesting. Use Luna for the turns where it wins on price-to-performance. Escalate only when the task needs more intelligence—or multiple models. That’s what we’re building with OpenSquilla:github.com/opensquilla/op… 🦐
English
0
0
0
141
Haider.
Haider.@haider1·
GPT-5.6 Luna price-to-performance is ridiculous at just $0.053 per task, Luna-Max scores higher than: opus 5 low while costing nearly 6x less sonnet 5 high while costing over 6x less gemini 3.6 flash while costing over 8x less 3.5 flash-lite while costing nearly 1.5x less
Haider. tweet media
OpenAI@OpenAI

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.

English
28
34
676
49.3K
OpenSquilla
OpenSquilla@OpenSquilla·
The best model today may not be the best model next month. Your agent shouldn’t need to be rebuilt every time the leaderboard changes. Models should be replaceable.Memory, tools, and workflows should remain. #AIAgents #OpenSourceAI
OpenSquilla tweet media
English
0
0
0
22
OpenSquilla
OpenSquilla@OpenSquilla·
🦐 Open-source AI is entering its security era. NVIDIA, Microsoft, GitHub, Hugging Face, and others are joining forces to build open tools for securing AI agents. Open models need open harnesses—and open safeguards. #OpenSourceAI #AIAgents #AISafety
OpenSquilla tweet media
English
0
0
2
440
OpenSquilla
OpenSquilla@OpenSquilla·
GitHub Models was officially retired yesterday. Models, providers, and products change fast. Your agent’s memory, tools, skills, and workflows shouldn’t disappear with them. Keep the harness independent. Keep models swappable. #AI #AIAgents
OpenSquilla tweet media
English
1
0
1
637
OpenSquilla
OpenSquilla@OpenSquilla·
🦐 Today in AI: 🕵️ HF published the OpenAI agent intrusion timeline. 🔐 Claude Mythos found attacks on HAWK + reduced-round AES. 🤖 Kimi K3 weights are live: 2.8T params, 1.56TB. Models get stronger. The harness decides how they act. #AI #AIAgents
OpenSquilla tweet media
English
0
0
0
350
OpenSquilla
OpenSquilla@OpenSquilla·
@akshay_pachaar InsForge optimizes the context layer. OpenSquilla optimizes the decision layer above it: our open-source Agent Harness uses task, context, and tool signals to route each turn to the right model—using multi-model collaboration only when needed. github.com/opensquilla/op… 🦐
English
1
0
2
269
Akshay 🚀
Akshay 🚀@akshay_pachaar·
I cut Fable 5 token usage 2.5x with just one change! (100% open-source solution) - Before: 5.5 M tokens · 7 errors · $8.94 - After: 2.3 M tokens · 0 errors · $4.17 The final build was the same for both, but the path the agent took wildly differed. In both runs, the agent started with the same thing, i.e., it understood the backend before building anything, like: - Permission policies - Available storage buckets - Auth providers configured - How edge functions are deployed The first run used Firebase, which was built for a human dev using a dashboard. While the dev can read the above state by clicking through tabs, an agent has no dashboard. So it gathered the same info through API calls. And there's no single Firebase call that returned this info. The agent required to query multiple times, and each query over-returned. For instance, when the agent asked how sign-in is configured, Firebase also returned the entire auth surface and every method it supported. This was far more context than what it needed. And it repeated across every part of the backend it inspected. Some states (like which auth providers are active) weren't queryable at all. I provided it myself. Otherwise, the agent would have guessed. Errors further compounded the token usage. When a dev sees "permission denied," they can look at the console and figure out whether it's a rule, a path, or an unauthenticated request. Firebase returned the same string to the agent as well, and it had none of that surrounding context to debug. So it guessed again, picked the most likely cause, and rewrote code, utilizing more tokens. This Firebase setup cost me 5.5M tokens and 7 manual interventions during errors on a full-stack RAG app. But I brought that down to 2.3M tokens and 0 manual interventions by using InsForge as the backend context engineering layer (open-source and self-hostable via Docker). It provides the same primitives as Supabase/Firebase, but structures the entire information layer for agents, instead of dashboards. In one CLI call that consumed ~500 tokens, the agent saw the full backend topology before writing a single line of code. This included auth, database, storage, edge functions, model gateway, micro VMs, and deployment. Also, instead of loading the entire product surface into context on every task, four narrowly scoped skills activated only when relevant to keep cognitive load minimal. And to ensure efficient retries if needed, every CLI operation returned structured JSON with meaningful exit codes, so the agent never guessed what to do next. Here's the InsForge GitHub Repo: github.com/InsForge/InsFo… (don't forget to star 🌟) The video below depicts the final build, comparing Firebase and InsForge. If you want to dive deeper, my co-founder recently published a full walkthrough building the same RAG app on both backends and inspected them end-to-end. The article is quoted below!
Avi Chawla@_avichawla

x.com/i/article/2061…

English
29
42
379
73.9K
OpenSquilla
OpenSquilla@OpenSquilla·
This is how we approach agentic routing in OpenSquilla: single-model efficiency for routine turns, multi-model collaboration when the task truly needs it. 🦐 Open source:github.com/opensquilla/op…
English
0
0
0
14
OpenSquilla
OpenSquilla@OpenSquilla·
Multi-model systems shouldn’t mean “run every model every time.” That gets expensive fast. A better approach: → Easy turn: route to an efficient model→ Hard turn: use a stronger model→ Critical turn: let multiple models collaborate Ensemble should be a decision, not a default. #AIAgents #LLMRouting
OpenSquilla tweet media
English
1
0
0
572
OpenSquilla
OpenSquilla@OpenSquilla·
@wallstengine Microsoft gets it: models should be swappable. The harness is the durable layer. That’s what we’re building with OpenSquilla v0.5.1: model-independent routing, memory, tools, and orchestration. Open source. 🦐github.com/opensquilla/op…
English
0
0
1
120
Wall St Engine
Wall St Engine@wallstengine·
Microsoft $MSFT on enterprise AI architecture: “You’ve got to keep your harness separate from the model.” “That means any given model, at any given time, is swappable.” “You can’t depend on any one model.” “You can’t be subject to the refusals of one model.” “You can use frontier models. You can use multiple of them. You can use low-cost models, and train your own model when you don’t want to use any external model.”
English
15
21
170
48.6K