Raj Saha | Building Cloud Colosseum

620 posts

Raj Saha | Building Cloud Colosseum banner
Raj Saha | Building Cloud Colosseum

Raj Saha | Building Cloud Colosseum

@cloudwithraj

Founder of https://t.co/mDsiTtHlVm | Former Principal SA @AWSCloud

NYC Katılım Ekim 2021
176 Takip Edilen783 Takipçiler
Raj Saha | Building Cloud Colosseum
Most learning platforms measure videos watched. We measure infrastructure built. Since launching Cloud Colosseum, 100+ students have collectively completed 3,000+ hands-on cloud and Gen AI challenges. Some interesting statistics: - Total traffic activity on the app: 830K out of which malicious traffic is 117K. Our students are in all 7 continents! (Map attached) - Total hours spent by students in Cloud Colosseum: 7,590 hours and counting (from May 15th launch to current date) - Top 3 popular challenges: Microservice with AuthN/Z & Custom Domain, AI Agent with MCP & Memory on Agentcore, and ALB Vs NLB (most repeated challenge) - Most exciting feature (measured by student chaos on live classroom) - Player Vs Player Cloud Combat. World's first Cloud Combat podium finish decided by 3 mins difference! - Newest Hands-On Challenges Added: AI Eval, Lambda Cold Start Observe & Mitigate And we are just getting started. The next SA Bootcamp cohort will get access to Cloud Colosseum, including all existing challenges, live Cloud Combat, and several new projects! If you are ready to move from watching to building, Cohort 9 launches Aug 1st. Waitlist here: sabootcamp.com Keep learning, and keep rocking 🙌 🚀
Raj Saha | Building Cloud Colosseum tweet mediaRaj Saha | Building Cloud Colosseum tweet media
English
1
0
3
197
Peter Yang
Peter Yang@petergyang·
If you are over 40 and have the room for it this is genuinely one of the best purchases you can make for mental and physical health.
Peter Yang tweet media
English
55
6
259
26.3K
Chris Munns
Chris Munns@chrismunns·
If anyone sensible at AWS still follows me, you'll work to get this tweet taken down. The comments and quote comments on it are telling that people did not find this funny. There's a long history of people having scary situations involving runaway workloads. In many ppl reported being worried about their livelihood. this channel should maintain a strong bar of balanced messaging that is clear and crisp. not silly "haha" stuff. this is not high standards or customer obsession.
Amazon Web Services@awscloud

Typo alert: Some customers saw quadrillion-dollar AWS billing estimates today. Slight miscalculation on our end (very slight 😅). We're fixing it now. No action needed on your end. Sorry for the confusion. Real question: what will you do with those trillions instead?

English
53
48
1.3K
125.2K
Chris Munns
Chris Munns@chrismunns·
@cloudwithraj Good share. This is exactly the kind of believable bill that would freak someone out
English
1
0
2
156
Raj Saha | Building Cloud Colosseum
Almost got an heart attack when my lead developer messaged me we received $7.6 Million AWS bill!! Fortunately, AWS announced it as a glitch, but here are some cost optimization techniques we adopted in my startup: 1. Billing alerts - Using AWS Budgets with different spending thresholds (80%, 90%, 100%). Btw, this glitch didn’t trigger alerts. 2. Organization SCP - Restricting certain actions such as procuring GPUs when not needed. 3. Granular cost view per student - Built a dashboard where we can monitor each child account’s cost. 4. For Cloud Colosseum infrastructure, implemented all the best practices I learnt over the years. Some notable ones are Compute Optimizer, Power Tuning, Selective Spot use, intelligent prompt routing, delete EC2s on non work hours etc. —- Learn practical insights like these in a weekly newsletter (FREE): cloudwithraj.com/newsletter
Raj Saha | Building Cloud Colosseum tweet media
English
2
1
4
1K
Raj Saha | Building Cloud Colosseum
Close friends and family members asked me “Are you stupid to leave Big Tech?”. The hardest part of building a startup was enduring the pain of the buildup period. Every party I went to, I couldn't say I work at Big tech anymore, and couldn't reveal my product either. Now that my product is launched, we have a runaway for 18+ months with 4 full time employees. I go into rooms head held high, but the biggest realization is this: The smartest don’t win in business, it’s the one who can endure the most pain. If you are looking to do any change - physical, mental, career, you will face doubts and challenges. Just keep your heads down, endure, and keep moving. “Do what they think you can’t do” --- cloudcolosseum.io is a real-world Cloud and Gen AI playground, with a leaderboard that grades what you build inside your account, and world first player-vs-player cloud combat - plus the interview answers, LinkedIn posts, and shareable badges recruiters actually look for. Check it out.
Raj Saha | Building Cloud Colosseum tweet media
English
0
0
5
186
Raj Saha | Building Cloud Colosseum
Everyone is talking about LLM evals. But what happens after your AI agent passes every eval and still fails in production? If users immediately start asking questions you never anticipated and were not in the eval dataset, your eval results don’t matter. To fix this, combine eval with observability. But how? I was reading the O'Reilly book Observability Engineering (2nd Edition), chapter 21, section "Using Evaluations for LLM Reliability." The book suggests closing that gap by using production telemetry. Use real inputs, outputs, traces, tool calls, and errors to continuously improve your eval dataset. Observability as eval feedback loop - pretty neat! I also enjoyed chapters on using production observability to improve your code, end-to-end observability for LLMs, and the role of AI Agents for observability. 👉Download the complimentary copy of the Observability Engineering, 2nd Edition book: fandf.co/3R9yEeq Personally, in a world where companies rise and disappear overnight, it’s refreshing to see honeycomb.io still pushing boundaries after a decade. Huge respect for Liz Fong-Jones and Charity Majors - I’ve followed their work for years and learned a lot from both. Thanks to Honeycomb for partnering with me on this post. #genai #observability #honeycomb
Raj Saha | Building Cloud Colosseum tweet media
English
2
0
1
232
Raj Saha | Building Cloud Colosseum
AI isn’t magic. It’s a factory and most people only ever look at one machine. 🏭 Swipe for the full floor plan, in plain english: 🔧 LLM - the machine. Stamps out language, but frozen at training time. 👷 Agent - the manager. Plans, uses tools, checks results, repeats. 📦 RAG - raw materials. Fetches exact facts, feeds them in fresh. 🗄️ Vector DB - the warehouse. Shelved by meaning, not by name. 🔌 MCP - the plugs. One standard socket for every tool. 🛑 Guardrails - the e-stops. Fail closed, not open. 🔍 Evals - quality control. Measurements, not vibes. The weakest station sets the ceiling. A great model with no guardrails still ships a bad product. -- Subscribe to my free newsletter to get real-world system design, Cloud Gen AI interview questions, and career switch steps: lnkd.in/g6_ZUzuq #genai #agent
Raj Saha | Building Cloud Colosseum tweet mediaRaj Saha | Building Cloud Colosseum tweet mediaRaj Saha | Building Cloud Colosseum tweet mediaRaj Saha | Building Cloud Colosseum tweet media
English
0
0
1
92
Raj Saha | Building Cloud Colosseum
How would you self-host an AI agent including the LLM in Kubernetes? This question is getting more popular fast, driven largely by companies worried about sending proprietary data to a model provider. The agent code itself is straightforward. Containerize it, push it to Amazon ECR, and run it as a pod. The model is the hard part. A model has two pieces: the model image, containing the tokenizer and configuration, and the model weights, which for an 80 billion parameter model means 80 billion floating point numbers the prompt runs through. Model weights live in S3 and run on EC2 with Nvidia GPUs, or on Inferentia using compiled models for Neuron cores. vLLM virtualizes access to the model so it can scale under load, the same way a hypervisor virtualizes a bare metal instance into multiple EC2 instances. MCP servers run via FastMCP for tools. Memory runs on an open-source vector database like Milvus, backed by object storage on a persistent volume. Karpenter and horizontal pod autoscaler handle scaling the whole thing. The balance: Full self-hosting buys you security and control over your data. It costs you real operational complexity across every layer, from GPU provisioning to vector database management. Plenty of teams mix and match, self-hosting models while hosting memory through AgentCore, depending on what they actually need to have more control over. --- Subscribe to my free newsletter to get real-world system design, Cloud Gen AI interview questions, and career switch steps: app.cloudwithraj.com/newsletter
Raj Saha | Building Cloud Colosseum tweet media
English
4
0
4
197
Raj Saha | Building Cloud Colosseum
One misconception I keep running into: "Chinese models like DeepSeek and GLM will send our data to China." No. That's not how any of this works. Let me break it down. If you call DeepSeek's API or Z.ai's API directly, sure, your prompts go to their servers. Same as calling OpenAI's API sends your prompts to US servers. But here's what most people miss: these models ship as open weights. That means you can download the actual model and run it on your own GPUs, in your own VPC, in your own data center. Self-hosted. Air-gapped if you want. When you self-host: - Nothing leaves your infrastructure - No prompts, outputs, or data goes anywhere near China (or anywhere else) - You control the model, the logs, the retention policy, all of it The origin of a model's training lab has nothing to do with where your inference happens. That's an infrastructure decision, not a nationality decision. Curious where you all stand on this. Are you self-hosting open-weight models, or is your org still API-only across the board? --- Subscribe to my free newsletter to get real-world system design, Cloud Gen AI interview questions, and career switch steps: app.cloudwithraj.com/newsletter
English
0
0
3
249
Raj Saha | Building Cloud Colosseum
I see this all the time in interviews. People know the names… but can’t explain when and why to use each - Canary Vs Blue Green vs Rolling Deployment! Canary Deployment: You release the new version to a small subset of users first (1-10%). You observe metrics, then gradually ramp up. It has lowest blast radius, ideal for high risk changes, however it's complex to implement. And, this is how you delight the interviewer - mention that some services in AWS such as Amazon API Gateway, has Canary feature built in! Blue-Green Deployment: You maintain two identical environments - BLUE is current prod (v1 in diagram), GREEN is new version (v2). You can switch traffic instantly, with instant rollback option. However it is expensive because you are running both versions. Delight the interviewer by giving an example such as using Amazon Route53 A record switching between two Application Load Balancers Rolling Deployment: You update instances incrementally in small batches while traffic continues flowing. Use when changes are backward compatible. One good example is Kubernetes which supports rolling deployment out of the box. --- Download this and other cloud interview questions and answers (FREE): app.cloudwithraj.com/freeguide #systemdesign #aws
Raj Saha | Building Cloud Colosseum tweet media
English
0
0
3
87
Raj Saha | Building Cloud Colosseum
OpenAI finances got leaked and they are projected to lose approximately $14 billion in 2026. Gen AI is not going to go away but will be changed the following ways: - Not every use case needs Gen AI. A chatbot that handles 50 FAQs does not need a frontier model. It does not even need RAG in most cases. You can feed those documents into a text-based database and query them directly. A data analytics workflow that runs SQL queries with group by and subqueries does not need Gen AI. It needs well-written SQL. Companies are figuring this out and pulling Gen AI out of use cases where it was never the right tool. What stays is the legitimate use cases. Complex reasoning tasks, multi-step agentic workflows. Personalization at scale. These use cases genuinely need GenAI and the investment is justified. - Gen AI cost optimization will become even more important As companies move from experimentation to production, the question shifts from "can we build this with Gen AI" to "can we afford to run this with Gen AI at scale." The engineers and architects who understand how to optimize inference cost, select the right model for each task, implement semantic caching, and design efficient RAG pipelines will be the ones companies hire. Memory and disk maker Sandisk's stock is up 4510% in less than a year (not financial advice!). Hopefully, memory and graphics card prices come down. I need to build a new PC 😅. Question to readers - what other ways you think Gen AI adoption will change? --- Subscribe to my free newsletter to get real-world system design, Cloud Gen AI interview questions, and career switch steps: app.cloudwithraj.com/newsletter
Raj Saha | Building Cloud Colosseum tweet media
English
1
0
3
126
Raj Saha | Building Cloud Colosseum
23: Graduated college. Interviewer rejected me saying I have stammering problem 24: Started working in Mainframes 35: Realized it’s now or never for career switch. Sick and tired of called skinny and low self-esteem 35. Started learning cloud, creating POCs in Verizon 36. Went to gym for the first time in life 36. Went from no LinkedIn, no networking to talking to cloud managers, showing POCs as Hail Mary. It worked 36. Said No to being promoted in mainframe team, joined Cloud team in Verizon 37. Got promoted to Distinguished Cloud Architect 38. Physique transformed, no longer a skinny awkward kid. Launched YouTube channel, first Udemy course 39. Got into AWS 42. Delivered multiple world-scale projects, spoke all over the world 43. Got promoted to L7 Principal at AWS. 100K YouTube Subscribers 45. Left Amazon. Founded cloudcolosseum.io. Cash positive with 18 months runway 45+. Just getting started Some would say I am late to everything. Everyone has their own timeline and map. Stick to yours. If I can do all these in my late thirties and mid forties, so can you. DO WHAT THEY THINK YOU CAN’T DO #career #provethemwrong #aws
Raj Saha | Building Cloud Colosseum tweet media
English
1
0
5
152
Raj Saha | Building Cloud Colosseum
Real-world learnings from building a startup used by real customers: 1. Security is equally or more important than cost optimization App website had 150K malicious traffic in one week! Prompt injection, trying get access keys, spoof logins, DoS and more. Screenshot attached! 2. A great architecture is what you can maintain, not what looks fancy on paper The real work starts AFTER you release the app - day 2 operations, releasing features fast, fixing issues. We chose architecture, and services we understood very well. 3. Gen AI is better at coding than design and deployment It misses the architectural nuances and I had to be real strict and verbose. Some examples - it doesn't suggest secondary indexes, suggests redundant tables/fields, suggests creating API gateway where we have a Load balancer etc. In general, I had to watch it like a hawk. Context bloat and compaction effects are real. 4. AWS quotas are a thing Ran into quota issues - Bedrock model usage limit, Load balancer path based routing rules limit, API gateway timeout, AWS Organizations limits etc. Finally signed up for AWS Business+ Support plan now that we have runaway. And now we know what limits to look out for! 5. Managing people is hard We have 4 fulltime employees and a couple part time contractors now. I want to design and code features, all by myself. I am learning to fight the urge and train my team instead, even if it takes longer. Delegation is hard, but either we scale, or we fail. Lastly, I am biased, but AWS is still the best cloud to build startups on. The security, Gen AI, compute services, and even the free support has been amazing! I'd have failed if I tried to use multi-cloud from get go. --- Subscribe to my free newsletter to get real-world system design, Cloud Gen AI interview questions, and career switch steps: app.cloudwithraj.com/newsletter #genai #startuplife
Raj Saha | Building Cloud Colosseum tweet media
English
0
0
0
117
Raj Saha | Building Cloud Colosseum
I get this question a lot - “Do you miss working at AWS?”. While it’s hard to beat the fast pace and the thrill of building a startup, I miss having the world's best practitioners one message away. Last week, it was a blast to hang out with my old friends from AWS during NY Summit. Sai Samala, Ashok Srirama, and Sai Charan Teja G. showed me the new self managed Gen AI on Kubernetes workshop they built where they are using liteLLM as the proxy layer before vLLM, smart! Heeki Park, and Dhiraj Mahapatro are building a bunch of neat agentic projects. They both share them on their socials, go follow them. We also went to a great Indian restaurant where we debated about how to precisely measure Gen AI cost, while losing count of the calories of chicken tikka, biryani, and tandoor! Until next time!
Raj Saha | Building Cloud Colosseum tweet mediaRaj Saha | Building Cloud Colosseum tweet media
English
1
0
1
148
wiktor romanowicz
wiktor romanowicz@WikRomanowicz·
At Skool, we had entire quarters where we shipped nothing new. No new features. Just improving what already existed. That’s how you build products that actually win markets.
English
1
0
1
17
Raj Saha | Building Cloud Colosseum
I love containers and use them extensively. I do not love running a Dockerfile fifty times a day to test one change. In production, containers are great. During development, they are a heavy lift: - Rebuild, scan, tag, push every time you touch a dependency - Bind mounts and port forwards just to reach your own tools - Cut off from your own editor, shell, and personal setup What if I told you that you could pull in every dependency with one command? No container or virtualenv required. Flox builds your environment declaratively: - Define everything a project needs in one manifest: packages, tools, environment variables, and services. - Runs natively on your machine (macOS or Linux, x86 or ARM) - Every package, library, and tool locked to an exact version, the same for the whole team on any laptop, and in Prod - When dev is done, turn that exact environment into a container for deployment Same pinned environment everywhere, so "works on my machine" stops being a thing. This is why Flox is trusted by Fortune 50 leaders like NVIDIA. Thanks Flox for partnering with me on this post. If setting up your environment eats more time than it should, check how Flox solved this for Resolve AI: fandf.co/4vVfTdn
Raj Saha | Building Cloud Colosseum tweet media
English
0
0
2
85
Raj Saha | Building Cloud Colosseum
If you are around in NJ, and looking for an AI meetup, checkout luma.com/p4cpr42x, happening this Thursday June 25th at 6:30 PM in Moncler, NJ. I love NYC, but going there is such a hassle. I am so happy to finally see a meetup on the better side of the river hahaha. Thank you Ritik Khatwani from AWS for inviting me to this meetup. If you are attending and bump into me, come say Hi. #aws #genai #NJvsNY
Raj Saha | Building Cloud Colosseum tweet media
English
0
0
1
126
Raj Saha | Building Cloud Colosseum
When I was an L7 Principal SA at AWS, I presented at conferences across the world. But there was one big downside… I was never allowed to attend the talks of my fellow colleagues. This week I attended the AWS New York Summit as an AWS customer and a learner. And honestly? It was fun, insightful, and definitely more relaxing! Getting to learn from my former colleagues, from the other side of the room, hits differently. A few things stood out that AWS is doing differently than anyone else right now: - AWS Context, a new service that automatically maps the relationships across your existing data into a knowledge graph and provides agentic search - Amazon Bedrock Managed Knowledge Base, a fully managed RAG (Retrieval-Augmented Generation) service - AWS Transform continuous modernization detects end of life dependencies, deprecated frameworks, and other common sources of technical debt If you want to go deeper on these, check it out for free here: bit.ly/3SvHK5t It was also a blast seeing my former coworkers and catching up with them. (Thank you Amazon Web Services (AWS) for sponsoring this post) #aws #genai
Raj Saha | Building Cloud Colosseum tweet media
English
1
0
4
166