
Mohammed Alshehri
2.8K posts

Mohammed Alshehri
@M0EGPT
Applied ML\RL, Post Training → building and learning prev @ibmwatsonx





a year later - thanks everyone for all of your help along the way! just the beginning! excited to share more learnings and hot takes from the past few months of tinkering soon :)


true story - my first day as a janitor 3 years ago 🥺


Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control. openai.com/index/safety-a…




How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less? I put together a “little” article explaining how these effort levels are implemented at inference time and during training.

Interesting fact people might not know: Moonshot AI (Kimi models) founder Yang Zhilin, who also goes by the nickname 'Kimi', studied under Jie Tang who is a co-founder of Z AI (GLM models) at Tsinghua university (before CMU)










