JV Roig

340 posts

JV Roig banner
JV Roig

JV Roig

@RoigJV

I make stuff. Security, cryptography, cloud, AI, others. More about me in my LinkedIn account.

Philippines Katılım Haziran 2012
145 Takip Edilen41 Takipçiler
JV Roig
JV Roig@RoigJV·
@CIMBBankPH I can't log in to the app. Can I get support? Thanks.
English
0
0
0
4
JV Roig
JV Roig@RoigJV·
Just another day at the @KamiwazaAI and @Signal_65 AI research lab, I randomly discovered that tokenizer differences between open models now can be up to 30%. Not a rounding error! Hidden dimension that affects cost and perf. More in the blog: lnkd.in/gdKrwdUs
English
0
1
2
1.4K
JV Roig
JV Roig@RoigJV·
@WessLeYankee The shorter the context, the better they are at finding information that actually exists. The fabrication rate is different though. But RIKER-like training can actually mitigate it, and seems a promising area of research: docs.kamiwaza.ai/research/blogs…
English
0
0
1
17
wess trabelsi
wess trabelsi@WessLeYankee·
@RoigJV Thx for this work @RoigJV! I wasn't familiar with RIKER and finds this brilliant. Do you have any insights to share about performance on smaller contexts (5 to 10K) since your baseline is 32K?
English
1
0
0
17
JV Roig
JV Roig@RoigJV·
New paper drop: How Much Do LLMs Hallucinate in Document Q&A Scenarios? It's a lot 😂 This is a big study that took 170+ Billion tokens. And if you are thinking "well, just use temp=0.0!", I have really bad news for you... arxiv.org/abs/2603.08274
English
1
0
3
458
JV Roig
JV Roig@RoigJV·
Remember "you cannot improve what you cannot measure" ? Well, now evaluating and improving knowledge retrieval systems for AI deployments just got way easier: docs.kamiwaza.ai/research/paper…
English
0
0
0
47
JV Roig
JV Roig@RoigJV·
@rohanpaul_ai Hi @rohanpaul_ai, author here. Thanks for the coverage! This 1st paper is very constrained due to the manual effort of studying 900 traces across the 3 models. Soon we'll have a follow-up that uses AI orchestration to reliably analyze at scale. We have 200K+ traces left to study!
English
1
0
3
105
Rohan Paul
Rohan Paul@rohanpaul_ai·
The paper explains why language models acting as agents fail in computer-like tasks and which patterns cause those failures. Concludes that LLM agents become reliable only when they are trained to ground actions, verify data, and recover from errors, not just scaled up in size. It tests 3 models on 900 tasks where they must use files, code, and databases through tools. Here an agentic model is one that takes several steps, calls tools, reads results, and then chooses the next action. The study shows that bigger models are not automatically more reliable, and that DeepSeek mainly wins from extra post training on tool use. Across tasks the authors see 4 recurring failure types, premature action, over helpful guessing, distraction from context, and brittle behavior on long runs. Premature action means guessing table or column names instead of asking the database what structure actually exists. Over helpful behavior shows up when a model invents a similar company or region and reports a number even though the requested one is missing. Context pollution happens when tables or CSV files tempt the model to mix in wrong numbers, and brittle execution appears as malformed tool calls, loops, or losing track of steps once traces are large. ---- Paper Link – arxiv. org/abs/2512.07497 Paper Title: "How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations"
Rohan Paul tweet media
English
9
25
139
8.6K
JV Roig
JV Roig@RoigJV·
See how easy it is for backdoored RNGs to go undetected: I ran an experiment with three supposedly-secure RNGs, but one of them has an easy-to-exploit backdoor. Is it possible to find out which is which, just by examining the random output? linkedin.com/pulse/undetect… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Here's another anti-procrastination and general productivity tip: think task consistency, not task intensity. Read the article to find out how you can use task consistency to overcome procrastination and boost your own productivity. linkedin.com/pulse/producti… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Learn the basics of encryption and randomness - and just exactly WHY and HOW randomness is so important to your privacy and security. You don’t need a math degree. Also, you get to pretend to be a Roman general! linkedin.com/pulse/why-does… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
This is Part 1 of a cybersecurity-and-cryptography-related series on randomness, its importance to your security and privacy, and the challenges in this space. linkedin.com/pulse/randomne… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Procrastination isn’t because you’re bad at time management - it’s because you’re **bad at tactics**. Why is that, and what tactics? Read Psych Fridays #17 to find out more! linkedin.com/pulse/procrast… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Random discussions are essential to unlocking your team’s creativity and innovation. It can turn 250GB zip files into half a gigabyte. Read the article to find out more! linkedin.com/pulse/creativi… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Here’s a quick and handy guide for designing continuous improvement mechanisms for your tech team - and all you need to understand is a one-line joke. linkedin.com/pulse/central-… via @LinkedIn
English
0
0
0
0
JV Roig
JV Roig@RoigJV·
Opportunities due to random events, arbitrary rules, and unforeseen side effects happen all the time, and have a profound effect on human existence, success, and failure. That’s life. Not everything is just hard work. linkedin.com/pulse/possible… via @LinkedIn
English
0
0
0
0