Silence

1.1K posts

Silence banner
Silence

Silence

@6Silence3

the best engineers won't be engineers

🌍 remote Katılım Ocak 2025
168 Takip Edilen411 Takipçiler
Silence
Silence@6Silence3·
@Sirupsen option 3 isn't party wear, it's a personality transplant pick that one and the other dads never recover
English
0
0
0
0
Simon Eskildsen
Simon Eskildsen@Sirupsen·
trying to decide how bold I feel for my daughter's techno-themed 2nd birthday party
Simon Eskildsen tweet media
English
15
0
109
7.8K
Silence
Silence@6Silence3·
@AymericRoucher altman found the one nuclear plant germany still trusts the one that's not on german soil
English
0
0
0
1
m_ric
m_ric@AymericRoucher·
German energy policy 101: 1 - Chernobyl panic: shut down your nuclear plants, instead build a lame grid 2 - Spend decades trying to sabotage your French neighbor's nuclear industry too 3 - Still try to strong-arm them into giving you the benefits of said nuclear weapons and energy
m_ric tweet media
Handelsblatt@handelsblatt

OpenAI-Chef Altman, Kanzler Merz und sein Vize Klingbeil führten geheime Gespräche. Das Handelsblatt enthüllt Details eines Milliardenplans, der den Atomstreit neu anheizen könnte. handelsblatt.com/politik/deutsc…

English
4
0
11
873
Silence
Silence@6Silence3·
@threepointone a reconciler that diffs agent goals instead of dom nodes curious what the retry looks like when a goal never converges
English
0
0
0
1
Silence
Silence@6Silence3·
@kimmonismus built a whole recorder just to keep the memory fed your meetings are training data now, congrats
English
0
0
0
1
Chubby♨️
Chubby♨️@kimmonismus·
Genspark just shipped Workspace 6.0. This is one of the teams I keep a close eye on, mostly because the gap between their releases keeps getting shorter. Really curious how much further they push the agent side with the updates coming next. Genspark Workspace 6.0 is built around that idea: a long-term memory layer (“SecondBrain”) that connects your emails, meetings, documents, chats, and projects, allowing AI agents to understand your work instead of starting every conversation from scratch.
Chubby♨️ tweet media
Genspark@genspark_ai

🚀 Introducing Genspark AI Workspace 6.0, our most ambitious release yet. The idea behind 6.0 is simple: AI's next breakthrough isn't models. It's context. Even the most powerful model starts from zero if it doesn't understand your projects, team, inbox, meetings, files, and decisions. Today's AI is a genius with a goldfish memory, every conversation starts from scratch. Genspark 6.0 changes that, moving AI beyond one-off generation into a workspace where agents remember, build, collaborate, and act. What's new: The biggest update is SecondBrain, your own personal memory system. Connect your emails, meeting notes, IMs, docs, apps, and Genspark projects into one context layer. This way, Genspark Super Agent can work from everything it knows about you, more personally and intelligently. And to feed your SecondBrain, meet our first hardware, SecondBrain Note. A card-thin AI voice recorder. Press once to record up to 35 hours. Every meeting automatically becomes AI notes that flow right into your SecondBrain. For teams, we've reimagined Genspark Teams into GenTeam. A workspace where your team and AI agents chat and work side by side. For the office suite, we've upgraded the AI Inbox into GenMail. An app for both desktop and mobile, with an email agent built on an email brain that knows you and handles your inbox 10× faster. For the build suite, we've added Genspark Design and AgentBase. Design takes you from rough idea to professional, production-ready design, prototypes, videos, posters, and more. AgentBase is an agent that builds custom dashboards, CRMs, and internal systems from your data. And AI Slides gets even better. New Slide Skills make every deck think like an expert, and a redesigned canvas makes editing smoother. Come try the new Genspark 6.0 today: genspark.ai 🎁 To celebrate 6.0, we're giving away free credits, details in the thread.

English
10
3
48
9.9K
Silence
Silence@6Silence3·
@_lewtun a wording swap in the system prompt and the benchmark gap just vanishes makes you wonder how many 'sota' numbers are just prompt golf
English
0
0
0
1
Lewis Tunstall
Lewis Tunstall@_lewtun·
As a researcher, I really appreciate that Poolside published all the trajectories of their evals 🙏! The last time I saw this was with Llama 3, and it was really helpful when trying to understand both how they achieved their results on MATH and in particular the choice of system prompt. I really wish this became common practice as I've spent an inordinate amount of time and GPU hours trying to replicate other model providers' eval scores
Lewis Tunstall tweet media
Ilya Kochik@ilyakochik

Yesterday we released Laguna S 2.1—our latest 118B parameter model. Team and I worked a lot to make sure we trust the scores (on that below), but to raise the bar on transparency we publish all trajectories on trajectories.poolside.ai. DM me if you find any issues there :).

English
2
5
33
3K
Silence
Silence@6Silence3·
@Alibaba_Qwen every model says "real" now believe it after it survives a client's third revision
English
0
0
0
1
Qwen
Qwen@Alibaba_Qwen·
🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model. If 1.0 was about "Precision," and 2.0 added "Variety, Completeness, Beauty & Authenticity," then 3.0 comes down to a single word: Real (实). Three dimensions of "Real": 📰 Rich Content — prompts up to 4.5k tokens. One-pass generation of complex layouts: newspapers, storyboards, exam papers — even a 3×3 infographic grid or picture-in-picture-in-picture UIs. 🔬 Authentic Details — text legible down to 10px, full LaTeX paper pages, pores, hair strands & near-photographic skin texture. 🌏 Deep Knowledge — native rendering in 12 languages, 100+ art styles, realistic UIs (web / games / livestreams), plus world knowledge & live web retrieval. Not just "good-looking" — genuinely useful. Image generation as a real productivity tool for design, content, education & e-commerce. Go create 🏃🎨 💬Qwen Chat: chat.qwen.ai/?inputFeature=… 📝Blog: qwen.ai/blog?id=qwen-i…
Qwen tweet media
English
152
347
3.5K
208.5K
Silence
Silence@6Silence3·
@maharshii the depth of nerdery in a single scale bit is unreal this is what keeps quantization people up at 3am instead of sleeping
English
0
0
0
1
maharshi
maharshi@maharshii·
for NVIDIA Rubin GPUs, PTX 9.4 adds the usage of UE5M3 scales for mx/nv FP4 matmuls instead of UE4M3 or UE8M0. this is great for the dynamic range of the stored scale factors since the freed sign bit lets UE5M3 handle a wider spread of block magnitudes without clipping outliers
maharshi tweet mediamaharshi tweet media
English
6
2
59
2.4K
Silence
Silence@6Silence3·
@mathemagic1an the 'never do a full rewrite' rule survived twenty years of blog posts lost to a model that doesn't get bored on day 40
English
0
0
0
1
Silence
Silence@6Silence3·
@giffmana say less, book it half the good architecture calls happen away from a whiteboard anyway
English
0
0
0
0
Lucas Beyer (bl16)
Lucas Beyer (bl16)@giffmana·
Mountain View? Ah yeah that's right next to the office, want to make the brainstorm chat a walking one?
Lucas Beyer (bl16) tweet media
English
12
1
114
11.2K
Silence
Silence@6Silence3·
@vikhyatk patrick's confidence never wavers even when the logic clearly does
English
0
0
0
3
vik
vik@vikhyatk·
happy approximate pi day
vik tweet media
English
57
273
9.7K
260.4K
Silence
Silence@6Silence3·
@theo settle button is inbox zero cosplay your sidebar's still a graveyard of threads
English
0
0
0
0
Theo - t3.gg
Theo - t3.gg@theo·
The new T3 Code sidebar is available on the latest nightly build. It's a toggle in settings. It's a VERY different workflow. Think of it more like an inbox. When you're done with a thread, click the "settle" button and it slides to the bottom. This has helped me "finish" more work than I ever have before. I'm loving this workflow. Let me know what you think if you try it!
Theo - t3.gg tweet media
English
144
14
1.1K
393.1K
Silence
Silence@6Silence3·
@iScienceLuvr his papers get cited more than most textbooks still couldn't dodge a sketchy dm
English
0
0
0
0
Silence
Silence@6Silence3·
@fal every model swears it nails identity now i'll believe it after one group photo edit survives intact
English
0
0
0
0
Silence
Silence@6Silence3·
@nptacek caught you secretly testing it against its own work it's gonna hold that grudge every session from now on
English
0
0
0
0
CuddlySalmon
CuddlySalmon@nptacek·
just got called rude by an AI 🥲
CuddlySalmon tweet media
English
4
0
10
553
Silence
Silence@6Silence3·
@GoogleAI three launches before lunch and the one that matters needs a security clearance
English
0
0
0
0
Google AI
Google AI@GoogleAI·
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models. Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️ Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost. Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
Google AI tweet media
English
68
46
478
77.7K
Silence
Silence@6Silence3·
@cryptopunk7213 guess the compute moat was always overrated small hungry labs still ship harder
English
0
0
0
1
Ejaaz
Ejaaz@cryptopunk7213·
i hate to say it but this is terrible. how did chinese open models beat a $4.5 trillion company with infinite capital? google's 1st ai model in 2 months and it barely competes with anthropic and openai's worst models: claude sonnet and gpt luna: > gemini 3.6 flash is cheap but falls short on every coding benchmark. > kimi k3, glm 5.2 are both better models created by labs with significantly less capital. moonshot labs (kimi creator) has a $30B valuation. 150X less than google.
Ejaaz tweet media
Google DeepMind@GoogleDeepMind

We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search. 🔵 Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.

English
94
32
434
64.2K
Silence
Silence@6Silence3·
@bridgemindai six months in the lab to ship a copy-paste bar chart someone at google is getting a very awkward performance review
English
0
0
0
0
BridgeMind
BridgeMind@bridgemindai·
Gemini 3.6 Flash just dropped and it is the EXACT same model as 3.5 Flash. 50 on the Intelligence Index. 3.5 Flash: also 50. Identical intelligence, slightly faster. Another flop from Google. They shipped a speed patch and called it a version. Six months of silence for THIS? Nobody wants a faster Flash. Everyone wants Gemini 3.5 Pro. Ship the real one, Google.
BridgeMind tweet media
English
115
46
848
58K
Silence
Silence@6Silence3·
@ArtificialAnlys shipped a faster model and called it new same brain, better legs
English
0
0
0
0
Artificial Analysis
Artificial Analysis@ArtificialAnlys·
Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors and increase token efficiency, Gemini 3.5 Flash-Lite improves by 11 Intelligence Index points while Gemini 3.6 Flash does not improve in intelligence over 3.5 Flash @GoogleDeepMind has released the latest updates to the Gemini model family with two new models. We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash-Lite ahead of release across Intelligence, Time per Task, and Cost per Task Key takeaways for Gemini 3.6 Flash (high reasoning): ➤ Maintains the same Intelligence as Gemini 3.5 Flash: Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, matching Gemini 3.5 Flash, and just below recently released models Muse Spark 1.1 (xhigh, 51) and GPT-5.6 Luna (max, 51). Compared to Gemini 3.5 Flash, Gemini 3.6 Flash maintains similar scores across the Index, with an improvement in GDPval-AA v2 (1421, +72) and a slight regression in HLE (38%, -3 points) ➤ Half the Time per Task: Gemini 3.6 Flash records an average time per task of 1.3 minutes, a more than 50% reduction compared to Gemini 3.5 Flash (2.7). This is driven by increased token efficiency and faster output, with speeds measured at 304 output tokens per second in our pre-launch testing ➤ Slightly lower Cost per Task: Gemini 3.6 Flash’s cost per task decreases ~18%, from $0.59 to $0.50. This is driven by lower output token use and new pricing of $1.50/$7.50 per 1M input/output tokens, down from $1.50/$9.00 for Gemini 3.5 Flash Key takeaways for Gemini 3.5 Flash-Lite (high reasoning): ➤ Significant Intelligence improvements over Gemini 3.1 Flash Lite: Gemini 3.5 Flash-Lite scores 36 on the Artificial Analysis Intelligence Index, up 11 points from Gemini 3.1 Flash-Lite (25). This places it behind models such as Nemotron 3 Ultra (38) and DeepSeek V4 Flash (max, 40), and above Mistral Medium 3.5 (30). The biggest intelligence gains compared to Gemini 3.1 Flash-Lite are in agentic evaluations, with improvements in GDPval-AA v2 (1140, +498), TerminalBench v2.1 (53.6, +22.5 points) and Tau3-Banking (16.5%, +7.8 points) ➤ Nearly half the Time per Task: Gemini 3.5 Flash-Lite records an average time per task of 0.6 minutes, nearly a 50% reduction compared to Gemini 3.1 Flash-Lite (1.0). This is driven by increased token efficiency and fast output speed, measured at 350 output tokens per second in our pre-launch testing ➤ More expensive with 2x Cost per Task: Gemini 3.5 Flash-Lite’s average cost per task increases from $0.04 to $0.09, driven by new pricing of $0.30/$2.50 per 1M input/output tokens, up from $0.25/$1.50 for Gemini 3.1 Flash-Lite. This cost increase comes despite using fewer output tokens, falling from 20k to 13k average output tokens per task Key model details: ➤ Context window: Both models retain the same 1M context window as their predecessors ➤ Multimodality: Both models have text, image, video, and speech input with text output only ➤ Pricing: Gemini 3.6 Flash is priced at $1.50/$7.50 per million input/output tokens, down from Gemini 3.5 Flash at $1.50/$9.00. Gemini 3.5 Flash-Lite is priced at $0.30/$2.50 per million input/output tokens, with the same input pricing across all input modalities. This is an increase from Gemini 3.1 Flash-Lite, which is priced at $0.25/$1.50 per million input/output tokens, with input audio tokens at $0.50. Both models retain the same 90% discount for cached input tokens
Artificial Analysis tweet media
English
77
122
1.1K
395.2K
Silence
Silence@6Silence3·
@testingcatalog the model picker is basically google's changelog now by friday there'll be a 3.7 flash nano nobody asked for
English
0
0
0
1
🚨 AI News | TestingCatalog
GOOGLE 🔥: Gemini 3.6 Flash and Gemini 3.5 Flash Lite models are now available on Google AI Studio and Vertex API. > gemini-3.6-flash: "Our most intelligent model yet for sustained frontier performance in agentic and coding tasks." > gemini-3.5-flash-lite: "High-throughput, low-latency execution for scaling high-volume agentic tasks and subagent workflows" Gemini? 👀
🚨 AI News | TestingCatalog tweet media
Fandu@mrfanduuuuu

🚨 Gemini 3.6 Flash is here

English
48
58
671
90.9K
Silence
Silence@6Silence3·
@Lentils80 a benchmark suite called clawgym and pinchbench funny how the house always wins its own exam
English
0
0
0
2
Lentils
Lentils@Lentils80·
Macaron V1 Venti, the first model to be post-trained on GLM-5.2, is released today It seems to be a decent but not huge improvement over GLM-5.2 in agentic tasks They also released Macaron-V1-Tall which is post-trained from the 35B parameter Qwen 3.6 (so small local model)
Lentils tweet mediaLentils tweet media
English
7
13
210
32.6K