sophs l

32 posts

sophs l banner
sophs l

sophs l

@sofie25_k

on hour 3 of sims 4

Katılım July 2026
52 Takip Edilen0 Takipçiler

2026 Yıllık Özeti

@sofie25_k hesabının Twitter yılını gör

sophs l
sophs l@sofie25_k·
tbh my eyes are up here tbh
sophs l tweet media
English
0
0
0
1
sophs l
sophs l@sofie25_k·
@alexatallah wait which models actually hold up under real usage vs just benchmark chasing
English
0
0
0
0
Alex Atallah
Alex Atallah@alexatallah·
We spend a ton of time with customers struggling to write evals and figure out the best model for each task. This CLI wraps that experience into a single tool, allowing you to leverage all 500+ models on OpenRouter, their benchmarks, and real-world usage data. It's brand new. Give us feedback!
OpenRouter@OpenRouter

Introducing Ori Eval: the easiest way to write your first eval. There's no definitive best model, only the best model for each task. Ori Eval leverages OpenRouter's APIs for each task in your codebase, and then evaluates the results. curl -fsSL openrouter.ai/skills/spawn-o…

English
9
11
112
25.3K
sophs l
sophs l@sofie25_k·
@dwarkesh_sp are you factoring in the recent drop in H100 spot prices or is this a long-term view
English
0
0
0
9
David Senra
David Senra@davidsenra·
The only thing I want on my tombstone: “He was a rebel.” @mickymalka on the rebel mindset: “That's all I want on my tombstone and that will say everything about me. The way I raise my kids. The way I treat my friends. The way I make business decisions. The way I change my mind. I'm willing to change my mind if I learn something new and different. That's a rebel mindset. The rebel mindset is one where you’re always asking questions. A lot of people get to a point in life when they stop asking questions and they start to preach. That's past the rebel time of their life and everybody has that breaking point somewhere in their lives. I want my breaking point to be the last day of my life.”
David Senra@davidsenra

.@mickymalka is the founder of @RibbitCapital — the firm behind Revolut, Robinhood, Nubank, Coinbase, and more. He bought his first share of Berkshire Hathaway when he was 13, in Venezuela, with money he borrowed from his grandfather. His grandfather charged him interest. He started his first company at 17. In his mid-20s, he put every dollar he had into launching a bank in Brazil for the 50 million Brazilians nobody else would serve. Micky is a quiet killer. Respected by the elite of the elite. When I asked @patrick_oshag how he’d describe Micky to someone who didn’t know him, he said: “He’s a mfing moneymaker.” Micky hates labels. The only label he wants is on his tombstone: “He was a rebel.” 0:00 Why Micky Malka Refuses to Be Labeled 3:00 Writing Your Way to Conviction 5:41 Buying Berkshire at 13 and Learning Buffett's Operating System 11:48 Token Factories, AI Bankers, and the Future of Money 18:38 The Infinite Game: Why It's Better to Be Behind 22:09 Gen Z Founders and the Chaordic Company 29:00 Why Young Founders Want to Build Atoms, Not Just Bits 34:29 Bringing Beauty and Taste Back to Technology 36:56 Revolut, Founder DNA, and Earning Deep Trust 45:54 Building OnePay With Walmart 50:53 From Lemon Bank to OnePay: A 20-Year Idea 56:00 Compounding Trust and Protecting Your Reputation 1:00:30 Node and the Rebel Case for Digital Art 1:07:34 What It Means to Live as a Rebel 1:10:50 Why Ribbit Is Built Like a Startup 1:14:41 Charlie Munger and the Power of Time Includes paid partnerships.

English
9
12
195
15.2K
sophs l
sophs l@sofie25_k·
@guilleflorvs worked at a yc startup in 2015 and we used stripe from day one. never connected that detail
English
0
0
0
0
Guillermo Flor
Guillermo Flor@guilleflorvs·
Stripe's first customer was also a YC company. That detail matters more than people realize. Patrick and John Collison founded Stripe in 2010, the same year Y Combinator backed them. The project was originally called /Dev/Payments and was built out of Palo Alto. The very first transaction went through a YC company called 280 North — and that company's founder later joined Stripe. Stripe launched publicly on September 29, 2011. That same year, $2 million came in from Elon Musk, Peter Thiel, Sequoia Capital, and Andreessen Horowitz. Before any of this, the brothers had already exited once, selling Auctomatic for $5 million. In 2024, Stripe processed $1.4 trillion in total payment volume, up 38% from the prior year, equivalent to about 1.3% of global GDP. A tender offer in February 2026 put the company's valuation at $159 billion. YC didn't just fund Stripe. It was Stripe's first transaction. The ecosystem ate its own output.
Guillermo Flor@guilleflorvs

New businesses starting on Stripe are up ~2x year over year. Larger than the COVID surge. The biggest relative jump in Stripe's history. Patrick Collison at YC Startup School 2026: by every objective metric he can see from Stripe's live formation data, it has never been a better time to start a company. That's not a motivational take. That's the CEO of the world's most important payments infrastructure reading real signal from millions of businesses. The origin story makes the data land harder. Stripe started as a conversation on a walk home from sushi in 2009 — two brothers deciding to build in financial services in a sector so new the word "fintech" didn't exist yet. Then they spent two years building before launch. Not a lean MVP. Just-in-time building for a problem they were certain was real. Collison's explicit lesson for founders today: the lean startup playbook was already the wrong model when he started. It's more wrong now. What actually worked — and still does: Find a concrete customer problem in a sector that doesn't yet have a name. Build for the right amount of time, not the minimum amount of time. Do hard things yourself long enough to understand what you're actually building — including learning, not outsourcing cognition entirely to AI. The dropout math? Collison's framing: the cost is de minimis and nobody ever cared. Stripe's data says the window is open. The founding pattern says how you go through it is the variable that matters.

English
3
2
20
5.7K
sophs l retweetledi
mem0
mem0@mem0ai·
We cut vector search latency by 70x, from 8–14 seconds to ~110–120 ms. The biggest unlock? Moving vector search to @turbopuffer and keeping Postgres focused on what it does best. If you're building memory infra for AI agents, this one's for you: mem0.ai/blog/how-we-cu…
English
7
11
74
5K
sophs l
sophs l@sofie25_k·
@databricks does governed mean it snitches to my manager when i ask a dumb question
English
0
0
0
0
Databricks
Databricks@databricks·
Genie One puts your business data, dashboards, and apps in your pocket. Available on iOS and Android, it lets you: ☑️ Ask questions in natural language ☑️ Explore dashboards ☑️ Access Databricks Apps ☑️ Schedule tasks Everything runs on the same governed data, business context, permissions, and network security as the web experience, so you can get trusted answers and act on them wherever work happens. Download Genie One on mobile: databricks.com/blog/take-insi…
Databricks tweet media
English
3
3
20
2.5K
sophs l
sophs l@sofie25_k·
ong mood because you're not here tbh
sophs l tweet media
English
0
0
0
0
Pi
Pi@pidotdev·
Good morning from Vienna People of Pi🌞 Sunday meditations from @badlogicgames and @mitsuhiko - A good engineer says no a lot - Keep complexity low - The learning process gives you authority
English
14
23
386
25.1K
sophs l
sophs l@sofie25_k·
@benjamincowen do u think the long end is pricing term premium now or just sticky inflation
English
0
0
0
0
Benjamin Cowen
Benjamin Cowen@benjamincowen·
So many investors mistakenly think that lowering rates always causes long end rates to go lower. Even when presented with the evidence they won’t believe it. For instance the Fed cut rates from 5.5% to 3.75% from 2024-2025 and yet the 30 year yield is higher today than when interest rates were 5.5%! When the Fed cuts prematurely it causes the long end to go higher if the market thinks the Fed cut prematurely. When the Fed cuts because they are late then rate cuts can coincide with the long end going down. Have explained this ad nauseam for years but many will refuse to believe the evidence no matter what
Benjamin Cowen@benjamincowen

The 10Y yield is likely heading back to 5%. Long end rates will likely continue to rise forcing the Fed to raise rates before the end of the year

English
57
60
1.1K
137.9K
sophs l
sophs l@sofie25_k·
@thejessezhang how do you tell the difference between real FDE and just being a really expensive support engineer
English
0
0
0
0
Jesse Zhang
Jesse Zhang@thejessezhang·
Startups often conflate FDE with just being close to customers in the early days. Being forward-deployed when you're starting out makes a lot of sense because you get signal on what to build. FDE (in the original Palantir sense) exists to bridge the last mile in a $20m contract because the customer has bespoke needs and environments. Unless you are closing those contracts, FDE should not be a crutch to avoid doing the hard work of building a scalable product. Our engineers have always been very customer-facing, but something that helped us is truly viewing our offering as product-driven (as opposed to services or FDE) because we think that is the correct solution for our space. That forced us to skip some shortcuts but has paid off so far.
Cristina Cordova@cjc

I've never seen startups spend so much time looking sideways. One AI company builds forward deployed engineering and then everyone decides they need an FDE team too. One company pays a certain number for talent, and suddenly that's the market, whether or not companies can afford it. One company juices a launch with influencer tweets that read like the same intern wrote all of them, and suddenly that's the playbook. Those choices came out of constraints you can't see from the outside. Forward deployed engineering works when your contracts are seven figures and integration takes 6+ months. It's unlikely a fit it you're selling a product at $20/seat. The comp number works when one hire moves your model quality. Otherwise, it can be a pretty fast way to run out of money. Founder mode (the version I believe in) is about refusing to outsource your judgment. Being able to say "I've thought about it, we're not doing that," without needing a competitor's decision to drive your own. Why spend years and your best engineers building a worse version of someone else's company? Run yours.

English
11
5
175
35.2K
Brendan McCord 🏛️ x 🤖
A lot of my friends and/or people I admire signed “Pacing the Frontier.” I think this was a bad move. My disagreement isn’t with the forecast or the framing of the coordination challenge, but with the immense and illiberal power the letter implies. There is no object called “the pace.” Progress at the frontier comes from compute, algorithms, data, post-training, inference, unattended task length, the spread of model weights, how researchers organize, and other things we haven’t invented and don’t yet know about. Inquiry leads to progress along dimensions that can’t be exhaustively specified in advance. That’s the nature of the frontier. If you gate compute, the research effort moves to algorithms. Regulate releases? Labs start taking things in-house. And other 2nd order effects will be unpredictable. Any rule that must pace the frontier involves ever-shifting proxies. It requires that its administrator has standing authority to continually redefine what counts as dangerous progress. What else is required beyond adaptive scope? The pacing regime would also need speed. One can’t successfully intervene on recursive self-improvement only after six months of legislation and litigation. It will require executive discretion. The pacing regime would also need under-the-hood access. Frontier progress is a process. The regime would need to see internal model use, training activity, compute infrastructure, and perhaps code -- proprietary and strategically sensitive information. And the thresholds couldn’t be fully public, lest they invite firms to game them. So some standards and evidence would remain secret. Insofar as the regime had to verify a rival state’s compliance, that would be an intelligence function. Restrictions would be triggered partly by evidence an affected company or researcher, or the public, could not inspect. Because this contemplated power cannot be bounded by a stable regulatory object (in the way, say, nuclear weapons can be), it would depend heavily on discretion, speed, internal access, and secret evidence. This has a highly illiberal character. Coercive power should be specific, limited, reviewable, and governed by general and knowable rules. Its characteristics (e.g., trigger, scope, evidentiary standard, duration, exceptions, means of review) should be stated before the power is granted. And the burden is on those who would propose it. A defender might answer that the proposed tool need not be coercive at all. That it could be narrow and advisory, focused on evaluation and transparency and readiness. But that wouldn’t solve the letter’s stated problem: racing. With race dynamics, each actor is under pressure not to slow down because others may continue (and thus the frontier keeps advancing). You need a mechanism to bind defectors. Voluntary norms tend to be great for binding people and firms that interact repeatedly and care about reputation. But the letter says each company and _country_… and you can’t rely on informal solutions when dealing with an unwilling state. That’s why the audience for this letter is Washington and why it calls for an international effort. Its diagnosis implies a binding mechanism. @deanwball thinks it is sensible to have a break-glass plan. That plan must involve a binding instrument, because nothing weaker addresses the problem the letter describes. But that therefore carries the burden for the use of coercive power, mentioned earlier. @johnschulman2's suggestion that labs design voluntary mechanisms among themselves is a different notion and coherent one (I would have signed that letter), but the word “country” makes this direction incompatible with the pacing letter. @OpenAI recently argued that a federal evaluator shouldn’t be able to block deployments. A week after, @AnthropicAI proposed that the government should be able to block deployments. Both labs endorsed the same letter. Whether or not the state may stop a deployment is a central question. Yet the letter accommodates both positions. What then, does the letter really say? Like the “We Must Act Now” letter from @erikbryn, @ajay_bcv, @akorinek, and @testingham before it, the letter secures agreement at an altitude where the main disagreement disappears. Lastly, the benefit of pacing is not established. The kind of slowdown the signatories have in mind would seek to buy us time for things like alignment, cyber defense, biological countermeasures, or scientific understanding -- things that increasingly depend on technologies a pause would restrict. E.g., Anthropic's framework relies in part on AI-based biological countermeasures and its security program uses AI to give defenders an advantage. A researcher in the letter's own friendly commentary was astonished at how much agents accelerated the work of the best alignment people he knows, and gave that as his reason for wanting six more months. When danger and our capacity to respond to that danger are plausibly both accelerating, the relevant question is whether this relationship is asymmetric in a safety-improving direction at the level of real-world risk. A slowdown needs to differentially slow the production of danger vs. our capacity to understand and contain that danger. The letter doesn’t attempt to establish that. It treats slower and safer as though they are the same; they are not. The letter is a serious warning, but it is no good as a warrant for an undefined power over inquiry.
Zvi Mowshowitz@TheZvi

Full text of the Pacing the Frontier statement, signed by 1,122 employees for frontier AI companies so far, including a bunch of heavy hitters at OpenAI, Anthropic, Google and others:

English
34
32
219
99.1K
sophs l retweetledi
Praveen Neppalli
Praveen Neppalli@praveenTweets·
AI agents are everywhere at @Uber. It’s great to see, but the thing that keeps me up at night is how we are going to secure them. This is something that I have been thinking about for a while. Today, our agents run 50,000+ sessions per day across thousands of endpoints. And this isn't just engineering anymore. Employees across the company use agents that read code, run commands, call internal tools, analyze data, and act on real systems. That scale forced us to confront an important question: How do you secure agents when your security tools can't even see them? Traditional Endpoint Detection & Response (EDR) sees the file write, but not the prompt that triggered it. It sees the network call, but not the agent's reasoning. The intent, the thing that separates malicious from benign, is invisible. So we built Agentic Detection and Response (ADR): • Capture the full causal chain: prompt → reasoning → tool call → outcome, across Cursor, Claude Code, Codex, and every agent our employees use. • Triage cheaply: a fast, high-recall first pass handles the flood of benign sessions. • Reason deeply: only suspicious events get expensive LLM analysis, enriched with source code, threat intel, and policy context. • Red-team continuously: an offline explorer evolves hard attack variants before attackers find them. After 10+ months in production, the results speak for themselves: • Hundreds of credential exposures detected across 26 categories. • Shift-left prevention blocking secrets at 97.2% precision, before they ever leave the laptop. • Zero false positives on our enterprise benchmark, with 2-4x the F1 score of state-of-the-art baselines. • Every attack detected on AgentDojo, the public prompt injection benchmark. Just as valuable as the detections are the lessons from running this in production: • The workflow is the unit of security, not the individual tool call. Attacks hide in causally-linked chains that look benign step by step. • Credential leakage is a far more common operational issue than prompt injection. • Approval fatigue is real: when users approve 50+ actions per session, human oversight becomes a rubber stamp. You can't secure agents you can't observe. And nobody can solve this alone. That is why we recently joined the Open Secure AI Alliance (OSA), and why today we're taking the next step: open-sourcing ADR. The release includes the ADR Sensor, the detection framework, and ADR-Bench, the first enterprise agentic AI security benchmark: 302 tasks derived from real production telemetry and full coverage of all 17 attack techniques across 5 tactics, so the community can rigorously evaluate their own defenses. Code: github.com/uber/ADR Paper: arxiv.org/pdf/2605.17380… The future of AI security won't be built behind closed doors. Excited to see what the community builds on it, and what we all learn together! @UberEng
Praveen Neppalli tweet media
English
28
80
492
158.7K
Harry Stebbings
Harry Stebbings@HarryStebbings·
Three weeks ago, I was sitting in the drawing room of the holiday rental my family always gets in July in Frinton-on-Sea (St Tropez of the UK). It was midnight and as my grandparents were snoring loudly next to me; my phone suddenly buzzed. It was legendary investor, @shardul_shah @IndexVentures. As much as I love Shardul, it is rare for him to send me a deal and so when he does, I pay attention. Shardul was sharing @simile_ai with me, suggesting I meet @joon_s_pk. He spoke of all the companies he has backed, their growth being unparalleled. Given the man did Wiz, that really excited me! I asked for the intro, and 60 mins later, at 1 AM UK time, I was on a call with Joon. Straight away, it was clear he was one of the most unique talents I had met; an ability to be world-class in two seemingly opposing skills: creativity and scientific rigour. The man was a painter before becoming world class AI mind! Wow. I committed to invest on the call. Why? Joon is obviously insanely special. Second, I care deeply about upside. If this works, what could this be? If this works, it will change marketing forever. It will change stock markets forever. It will change human decision-making in general. Insane. I had to be a part of it and so I gave Joon a boatload of cash and asked if I could summarise the notes from our conversation below. (Links in comments) Huge thanks to @BennettSiegel @ninaachadjian @kevinhartz @BryanOffutt @karpathy for amazing thoughts, questions and feedback here. 1. The Question to Ask When Assessing Talent The best way to evaluate talent is to look for consistent performance across multiple stages of a career, not impressive titles. Ask one question: Were they the common denominator behind every success? Exceptional operators repeatedly create great outcomes across different roles, demonstrating rare ownership and adaptability. -- @kevinhartz you are one of the single best talent spotters in tech. What is the single biggest commonality in the best founders you have worked with? 2. The Sign of Truly Great Talent Exceptional talent combines traits that rarely coexist. While most people have correlated strengths, world-class builders embrace cognitive tension. They are short-term paranoid but long-term optimistic—obsessing over daily execution while maintaining unwavering conviction in the bigger picture. -- @dannyrimer you also have an immense love of art. How has art impacted your view on company building and entrepreneurship? 3. People Will Pay $100 Million for a Single Simulation AI will enable simulations so valuable that companies will spend tens of millions on a single run. A simulation may cost $20 million in compute, but if it prevents a $500 million product failure, paying $100 million becomes an extraordinary bargain. -- @GavinSBaker does this align to your views on specialised intelligence? 4. Why Simulation Requires Modeling Human Bias, Not Perfect Logic Frontier labs are racing to build the “CPU of intelligence” optimized for perfect reasoning. The bigger commercial opportunity may be the “GPU of intelligence”: models that replicate human judgment, bias, and imperfect decision-making to simulate complex real-world markets at scale. -- @ml_angelopoulos how does the role of evaluations change in this world?
Harry Stebbings tweet media
English
17
19
78
78.7K
sophs l
sophs l@sofie25_k·
tbh make sure your alone before u read the comments
sophs l tweet media
English
0
0
0
0
sophs l
sophs l@sofie25_k·
fr Naughty version in comments (fr)
sophs l tweet media
English
0
0
0
0
sophs l retweetledi
Jackson Stokes
Jackson Stokes@jackson_stokes·
“the only way to stop a bad guy with an open model is a good guy with an open model”
clem 🤗@ClementDelangue

We got attacked by secret unreleased proprietary models and defended ourselves with an open model, more precisely the @nvidia quantized version of GLM 5.2 coming from @Zai_org. Banning any open model would hurt first cyber security defenders, startups, small companies, researchers and everyone who's not a frontier lab and need on-prem affordable controlable models to compete and protect themselves. Let's not do that!

English
1
1
5
981