JV Roig
340 posts

JV Roig
@RoigJV
I make stuff. Security, cryptography, cloud, AI, others. More about me in my LinkedIn account.
Philippines Katılım Haziran 2012
145 Takip Edilen41 Takipçiler

Just another day at the @KamiwazaAI and @Signal_65 AI research lab, I randomly discovered that tokenizer differences between open models now can be up to 30%. Not a rounding error! Hidden dimension that affects cost and perf.
More in the blog:
lnkd.in/gdKrwdUs
English

@WessLeYankee The shorter the context, the better they are at finding information that actually exists. The fabrication rate is different though. But RIKER-like training can actually mitigate it, and seems a promising area of research: docs.kamiwaza.ai/research/blogs…
English

New paper drop: How Much Do LLMs Hallucinate in Document Q&A Scenarios?
It's a lot 😂
This is a big study that took 170+ Billion tokens.
And if you are thinking "well, just use temp=0.0!", I have really bad news for you...
arxiv.org/abs/2603.08274
English

Remember "you cannot improve what you cannot measure" ?
Well, now evaluating and improving knowledge retrieval systems for AI deployments just got way easier: docs.kamiwaza.ai/research/paper…
English

@rohanpaul_ai Hi @rohanpaul_ai, author here. Thanks for the coverage! This 1st paper is very constrained due to the manual effort of studying 900 traces across the 3 models. Soon we'll have a follow-up that uses AI orchestration to reliably analyze at scale. We have 200K+ traces left to study!
English

The paper explains why language models acting as agents fail in computer-like tasks and which patterns cause those failures.
Concludes that LLM agents become reliable only when they are trained to ground actions, verify data, and recover from errors, not just scaled up in size.
It tests 3 models on 900 tasks where they must use files, code, and databases through tools.
Here an agentic model is one that takes several steps, calls tools, reads results, and then chooses the next action.
The study shows that bigger models are not automatically more reliable, and that DeepSeek mainly wins from extra post training on tool use.
Across tasks the authors see 4 recurring failure types, premature action, over helpful guessing, distraction from context, and brittle behavior on long runs.
Premature action means guessing table or column names instead of asking the database what structure actually exists.
Over helpful behavior shows up when a model invents a similar company or region and reports a number even though the requested one is missing.
Context pollution happens when tables or CSV files tempt the model to mix in wrong numbers, and brittle execution appears as malformed tool calls, loops, or losing track of steps once traces are large.
----
Paper Link – arxiv. org/abs/2512.07497
Paper Title: "How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations"

English

See how easy it is for backdoored RNGs to go undetected: I ran an experiment with three supposedly-secure RNGs, but one of them has an easy-to-exploit backdoor. Is it possible to find out which is which, just by examining the random output?
linkedin.com/pulse/undetect… via @LinkedIn
English

If you are only sleeping 6 hours or less every night, you are chronically sleep-deprived. linkedin.com/pulse/life-cha… via @LinkedIn
English

AWS Lambda benchmark follow-up. New test results just enforce my recommendations from last time: Go include a pre-compiled DynamoDB resource in your layer.
#AWS #Lambda #serverless #DynamoDB #Python
linkedin.com/pulse/lambdabe… via @LinkedIn
English

Here's another anti-procrastination and general productivity tip: think task consistency, not task intensity. Read the article to find out how you can use task consistency to overcome procrastination and boost your own productivity.
linkedin.com/pulse/producti… via @LinkedIn
English

Learn the basics of encryption and randomness - and just exactly WHY and HOW randomness is so important to your privacy and security. You don’t need a math degree. Also, you get to pretend to be a Roman general!
linkedin.com/pulse/why-does… via @LinkedIn
English

Procrastinating? You’re Probably not Lazy or Bored - You’re Afraid (Psych Fridays #19) linkedin.com/pulse/procrast…
English

Faster AWS Lambda Cold Starts: Pre-Compile Your DynamoDB Resource (Tech Tuesdays #18) linkedin.com/pulse/faster-a… via @LinkedIn
English

This is Part 1 of a cybersecurity-and-cryptography-related series on randomness, its importance to your security and privacy, and the challenges in this space.
linkedin.com/pulse/randomne… via @LinkedIn
English

Procrastination isn’t because you’re bad at time management - it’s because you’re **bad at tactics**. Why is that, and what tactics? Read Psych Fridays #17 to find out more!
linkedin.com/pulse/procrast… via @LinkedIn
English

This week: Tech dinosaurs and meteors. What advice would I give to a tech dinosaur to avoid being wiped out by a meteor? Read on!
linkedin.com/pulse/free-adv… via @LinkedIn
English

Random discussions are essential to unlocking your team’s creativity and innovation. It can turn 250GB zip files into half a gigabyte. Read the article to find out more! linkedin.com/pulse/creativi… via @LinkedIn
English

Here’s a quick and handy guide for designing continuous improvement mechanisms for your tech team - and all you need to understand is a one-line joke.
linkedin.com/pulse/central-… via @LinkedIn
English

Opportunities due to random events, arbitrary rules, and unforeseen side effects happen all the time, and have a profound effect on human existence, success, and failure. That’s life. Not everything is just hard work.
linkedin.com/pulse/possible… via @LinkedIn
English
