Why Busy Servers Change AI Answers
A podcast adaptation discussing the Inference Stack Degradation Theory proposal and numerical stability in AI inference.
Listen to episodeEpisode transcript (139 paragraphs)
You know, usually when we talk about like a medical diagnosis, there's this expectation of total precision. Right. You break your arm, the x-ray shows that jagged white line and the doctor just points and says, you know, there it is. But then you step into the world of neurodevelopment or chronic autoimmune conditions and suddenly that x-ray machine is kind of useless. Yeah, the picture gets really blurry really
fast. Exactly. We are first to navigate this diagnostic landscape that is entirely murky, relying on environmental factors and fluctuating symptoms. And I mean, we just accept that. We accept that biology is messy. Right. We're organic, so it's complicated. But we hold on to this deeply comforting illusion that computers are the exact opposite. I think almost everyone listening to this deep dive has this ingrained
assumption about how machines work. If you give a computer the exact same input and you mathematically tell it to remove all randomness, it will give you the exact same output every single time. It's basically the foundational requirement for how we interact with technology, right? Yeah, we expect binary reality. Four is always the answer to two plus two. Right. We expect determinism at like both the hardware and
software levels. If you write a Python script to calculate payroll, you expect the float values to match down to the cent, regardless of whether you run that script at, you know, noon or midnight. We crave that determinism because it implies total control. It implies. It implies reliability. Well, today we are completely dismantling that assumption. We really are. Because in the world of modern large language models,
the AI systems that frankly are now embedded in our daily workflows, that perfect binary consistency is a complete illusion. Our mission in this deep dive is to unpack a really fascinating paper published in December 2025. Yeah, by independent AI researcher Robel Mumment. Right. It's titled Inference Stack Degradation Theory V1.5 or ISDT for short. We're looking at why these massive AI models act differently
depending on like how busy their servers are, how microscopic math errors basically snowball into totally different text outputs, and how we might actually need weather reports to track the reliability of our AI tools. What Mumment has done here is synthesize this growing unease in the machine learning systems community into a formal falsifiable theory. Which is so needed right now. Oh, absolutely. For the last
couple of years, developers have noticed their AI pipelines randomly changing. Randomly failing or acting erratically. And they usually just blame the model itself. Right. Assuming it's hallucinating or just inherently flaky. But Mumment forces us to stop blaming the conceptual brain of the AI and start looking at the physical stress of the data center. OK. So we need to start with the core mystery that sets up this
whole paper. Let's imagine you're sitting down at a computer opening up a developer console for a major AI model and typing in a prompt. Right. A standard setup. In those settings, there is a parameter called temperature. For those who don't know, temperature controls the creativity of the responses. If I set it to say, bound point eight, the model takes risks. It looks at the probability distribution of all the
words in the English language and occasionally picks a slightly less common word to keep the text interesting. You know. Yeah. The underlying mechanism there is the softmax function operating on the model's output logics. Logics being the raw scores, right? Exactly. The model calculates raw scores for every possible next token and the temperature parameter scales those raw scores. So the model can then use the data
to calculate the probability distribution of the model's output logics before they are converted into actual probabilities. So a higher temperature flattens the distribution, which gives lower ranked words a fighting chance of being selected. Right. But, and here's the kicker, if I set that temperature parameter down to exactly zero, I am telling the model to disable all of that creativity. You're forcing what
developers call a greedy decoding strategy. Greedy decoding. Meaning I want the model to look at the probability of the very next word and 100% of the time pick the absolute worst one. The absolute most mathematically likely word. No dice rolling at all. Right. So logically, if I send the exact same prompt at temperature zero a thousand times, I should get the exact same sequence of words a thousand times. That's the
assumption, yeah. That at temperature zero, the model functions like a highly complex but entirely deterministic calculator. The stochastic sampling is mathematically removed from the equation. But the output still changes. Yes. It does. That is the wild mystery Moomin highlights right in the introduction. Identical prompts at temperature zero are yielding different text. And to understand why, we have to look past
the abstract code and examine the actual physical reality of the servers running these models. We have to talk about the inference stack. Right. How a server's physical busyness physically alters its calculations. So break that down for us. Well, these models are served to the end user through an inference engine. A popular production engine is VLLM, for example. And running these multi-billion parameter models,
requires massive clusters of GPUs. Which are not cheap. No, extremely expensive. So to make this economically viable, companies cannot dedicate an entire GPU to process a single prompt in isolation. The memory bandwidth alone would be completely bottlenecked. So they use something called dynamic batching. Exactly. The inference engine takes your request, holds it for a few milliseconds, groups it with incoming
requests from other users around the world, and pushes them all through the GPU matrix multiplications simultaneously. As a single batch. Okay, so the size and composition of that batch change dynamically, right? Like depending on the traffic load at that exact millisecond. Yeah. If the server is quiet, your batch might contain four sequences. If it's peak hours, your batch might contain sequences of varying lengths.
I picture this kind of like ordering a highly specific custom coffee at a cafe. Oh, that's a good way to look at it. Like if you walk in right when they open, you're the only customer. You order a, I don't know. A half-calf, oat milk, vanilla latte. The barista takes your ticket and makes your drink start to finish. They steam the milk, pull the shot, pump the syrup in a very specific optimized sequence. Because they
have the luxury of time. Right. But if you come back at 8:30 a.m. during the morning rush, the barista is looking at your ticket alongside four others. They engage in dynamic batching. They steam oat milk for your latte and someone else's cappuccino at the same time in a bigger pitcher. They execute the batch. They execute the steps in a totally different interwoven order to maximize their espresso machine. Right. To
get everyone served efficiently. Exactly. You get your drink, but the physical sequence of its construction was dictated by the overall traffic in the cafe. It's a great analogy, but we do have to be careful with where that metaphor breaks down. Oh, how so? Well, in a cafe, steaming the milk before pulling the shot doesn't change the chemical composition of the espresso. The end product is physically identical. It's
identical to the one you got when the cafe was empty. Right. But inside a GPU, changing the execution order of the calculations actually changes the final mathematical answer. Wait, I need to make sure I'm grasping the scale of this. Are you saying that just because my prompt is batched with someone else's, the server calculates the numbers in a different order and that makes the mathematical output literally
different? It does, yes. And it traces back to a property called floating point non-associativity. Floating point non-associativity. Right. Moomin anchors his theory here. And he references prior work by He-Ale from 2025, which proved this phenomenon in large language models. To really get this, we have to look at how computers actually store numbers. Okay, let's do it. Because in grade school algebra, we learn the
associative property of addition. If you add A, B, and C, it doesn't matter how you group them. Like A plus B, then add C is the exact same as B plus C, then add A. Sure. In pure theoretical mathematics with infinite precision, associatives are not always the same. So, the probability of the number being the same is the same. But computers operate with finite memory. They use something called the IEE floating point
standard to represent real numbers. Which is basically scientific notation stored in binary, right? Essentially, yes. You have a sign bit, a certain number of bits for the exponent, and a certain number of bits for the mansissa, or the significant digits. So, they have to shop off the number at a certain decimal point because the physical silicon register only has so much space. Precisely. Let's look at a lower
precision format like FB16, which is heavily used in the field of quantum physics. It's used in AI to speed up processing. You only have bits to store a highly complex number. So, when you add two floating point numbers together, the computer has to align their exponents. Okay, I'm following. If one number is vastly larger than the other, the smaller number's binary representation gets shifted to the right to match
the larger number's exponent. And in doing so, the least significant bits of the smaller number literally fall off the edge of the available memory space. They're truncated. Or rounded away. Yes. So, every single time two numbers are added, a microscopic rounding operation occurs just to make the new number fit back into the box. Yes. Now, imagine a matrix multiplication inside a neural network. We are not adding
three numbers. We're adding millions of numbers together in a process called reduction. Millions. Okay. If the inference server uses dynamic batching during a traffic spike, the underlying computational kernels, which are the lowest level programs executing on the GPU, optimize their execution. execution paths to handle that larger batch, they slice and dice the matrices differently to keep all the compute cores fed.
Which means the sequence of addition changes. The sequence changes, which means the point at which those exponents are aligned changes, which means the exact microscopic bits that fall off the edge of the register are different. The rounding errors compound in a completely different pattern. Oh, wow. So the final numerical value that emerges from that massive block of math is physically different than if your prompt
had been processed in a smaller batch. That is mind-blowing. So just because someone in Tokyo asked the AI for a recipe at the exact same millisecond I asked it to write some code, the execution path on the server regroups the addition, the truncations happen at different intervals, and the actual logit score for my next word is physically altered. Yes. He, at all, proved this definitively. Batch size variability
alters kernel execution paths, which alters floating point reduction order, which creates... It's numerical deviations. But wait, if we know this is happening and we know it ruins the determinism we rely on, why don't developers just force the system to calculate the math in the exact same order regardless of the batch size? Like, just tell the GPU to stop optimizing and just execute sequentially. I mean, the
software exists to do that. They're called batch invariant kernels. These algorithms are designed to enforce a strict reduction order so that the numerical results are identical regardless of batch size. Let me guess. There is a massive catch. Oh. A huge one. Restoring identical outputs comes at an exorbitant computational cost. According to the research cited in Moomin's paper, enforcing batch invariant kernels
introduces a performance degradation of roughly 1.6x to 2.1x. It practically halves the speed of the entire data center. Yeah. Or it doubles the infrastructure costs to serve the same number of users. Because you are forcing the GPU to wait for specific calculations to finish in a rigid sequence rather than parallelizing everything dynamically. I see. So when tech companies... Yeah. ...face the choice between perfect
mathematical consistency and literally doubling their server capacity... Yeah. ...they choose speed and efficiency. Every time. They accept the floating point deviations as a necessary cost of operating at a global scale. Okay. So we have a physical reality where microscopic decimal differences are inevitable under load. A number might change from, I don't know, 0.4500001 to 0.45000002. Right. But how does a
deviation at the seventh decimal place deep in a GPU... ...at the GPU matrix turn into a completely different paragraph of text on my screen? That seems like a massive jump. It is. And to understand it, we have to examine the second major concept in Newman's theory, which is autoregressive amplification. Autoregressive amplification. Right. We need to look at the temporal dimension of language generation. You know,
large language models do not generate an entire response simultaneously. They generate text one token at a time. Token by token. Yeah. And every new token is conditioned on the entire sequence of previously generated tokens. It's a continuous feedback loop. It writes word one. Then it feeds word one back into the massive matrix multiplication to predict word two. Then it feeds word one and word two back in to predict
word three. Exactly. Newman's inference stack degradation theory formalizes this exact vulnerability. The theory states that the floating point drift is recursive. The drift at any given time step is a function of the current system load, the kernel properties, and, critically, the current system load. The drift at any given time step is a function of the current system load, the kernel properties, and, critically,
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of
the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any
given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current
system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step
is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The drift at any given time step is a function of the current system load. The
critique agent is the rigorous statistician. The critique agent is the rigorous statistician. It takes the massive datasets from the experiment agent and interrogates them. It takes the massive datasets from the experiment agent and interrogates them. It runs the hard math. It does. It checks if the token divergence is genuinely accumulating over the sequence length, conforming to the autoregressive theory. And it
executes statistical validations like the McNamara test. And it executes statistical validations like the McNamara test. Wait, why specifically the McNamara test? Because we are dealing with paired nominal data. We are comparing the exact same model's success or failure We are comparing the exact same model's success or failure on the exact same prompt under two different conditions: low load, versus high load. The
McNamara test is mathematically designed to determine the McNamara test is mathematically designed to determine if the marginal frequencies of two binary outcomes are equal. It proves whether the drop in accuracy during high load It proves whether the drop in accuracy during high load is statistically significant, or just a random sampling artifact. And the critique agent also runs bootstrap sampling to establish
confidence intervals around the PWI scores, right? to establish confidence intervals around the PWI scores, right? Yes. It is actively trying to disprove the experiment agent's findings. Yes. It is actively trying to disprove the experiment agent's findings. I love that. And if the data survives the critique agent, it moves to the synthesis agent. This final agent aggregates the validated data, generates the time
series plots mapping server load against the PWI, generates the time series plots mapping server load against the PWI, writes the correlation statistics, and produces a structured manuscript for human review. and produces a structured manuscript for human review. The entire scientific pipeline, from environment setup to statistical validation to final reporting, is executed by autonomous systems. It really is a
necessary evolution. The systems we are building are now so complex and opaque The systems we are building are now so complex and opaque that we require intelligent agents just to map their emergent behaviors. that we require intelligent agents just to map their emergent behaviors. And having a team of AI scientists running around the clock experiments And having a team of AI scientists running around the clock
experiments to figure out if another AI is getting stressed out by server traffic? That's wild. So, we have our agentic swarm calibrated. This leads us to the core experimental protocol. This leads us to the core experimental protocol. Moomin lays out four specific experiments for the swarm to execute Moomin lays out four specific experiments for the swarm to execute to prove the inference stack degradation theory.
Let's start at the micro level. Experiment is designed to confirm the baseline physical phenomenon. Experiment is designed to confirm the baseline physical phenomenon. The swarm needs to reproduce the batch-induced token level divergence. And it uses a short token task from the GSM8K dataset. And it uses a short token task from the GSM8K dataset. GSM8K is a benchmark consisting of grade school math word problems,
right? GSM8K is a benchmark consisting of grade school math word problems, right? GSM8K is a benchmark consisting of grade school math word problems, right? Correct. So the swarm queries the model under simulated load, say, Correct. So the swarm queries the model under simulated load, say, forcing the inference engine to only process four concurrent sequences forcing the inference engine to only process four
concurrent sequences and records the output. Then it blasts the server with simulated traffic, pushing it to concurrent sequences and runs the exact same prompts. So the goal is purely to verify that under high load So the goal is purely to verify that under high load floating-point non-associativity alters the token probabilities enough to generate unique string outputs. Yes, it establishes the baseline drift. It
proves that the barista's workflow optimization It proves that the barista's workflow optimization physically changes the math. But a token math problem is very short. Which is why Experiment stretches the baseline to test autoregressive amplification. The swarm extends the generation limits to tokens. The swarm extends the generation limits to tokens. And critically, it targets precision-sensitive tasks,
specifically code generation using the human evil dataset. specifically code generation using the human evil dataset. In a creative writing task, if the AI swaps the word "crimson" for "red" at token 300, if the AI swaps the word "crimson" for "red" at token 300, the poem still functions perfectly, but code is brittle. It's extremely brittle. Imagine the AI is generating a Python script to sort a database. At token
450, it is writing a conditional loop. At token 450, it is writing a conditional loop. Under low load, the math computes perfectly and it generates a "less than" operator. The loop functions. But under high load, the accumulated floating-point drift warps the probability and the token flips to a "less than" operator. and the token flips to a "less than" operator. Oh man, that single altered character introduces an
off-by-one error, the entire script fails to compile, or worse, it silently corrupts the database. or worse, it silently corrupts the database. Exactly. Experiment aims to prove that as sequence lengths grow, the autoregressive drift branches earlier and more dramatically, making long context generation fundamentally unstable under load. Okay, so Experiment proves the steering wheel is misaligned. Experiment proves
that if you drive for miles, you crash. Experiment zooms out to the macro level. This is where the prompt weather index is tracked over a 24-hour cycle. The swarm simulates a production load schedule. Low traffic from midnight to 8:00 AM, Low traffic from midnight to 8:00 AM, a massive peak load, mimicking a global workday from 8:00 AM to 8:00 PM, from 8:00 AM to 8:00 PM, and a taper back to low load. Every minutes,
the swarm queries a fixed set of prompts covering math, coding, and factual retrieval. And what is the null hypothesis here? The null hypothesis assumes that movement is entirely wrong. Naturally. It states there will be no significant correlation between the PWI scores and the concurrent load level. It assumes that any fluctuations in the model's stability are purely random noise, completely decoupled from the
physical state of the server. But if ISDT holds true, the swarm will chart a distinct degradation of the weather during that 8:00 AM to 8:00 PM window. The exact match weights will plummet, the similarity vectors will widen, and the PWI will reliably track the infrastructure stress. Which brings us to Experiment 4, the ultimate measure of operational impact: task-level correctness degradation. Right, because up until
now, we've been tracking stability. We are measuring whether the AI takes a different path or uses different words. But Experiment asks the critical question: Do these microscopic math errors actually cause the AI to fail the task entirely? Does the model become demonstrably capable when the server is busy? The swarm runs prompts per benchmark under both load conditions and measures absolute accuracy. Did it solve
the math problem? Did the code pass its unit tests? Newman hypothesizes a statistically significant accuracy drop of 5% or more during high-load conditions, particularly on those long, precision-sensitive problems. Wow. 5% is a staggering number in this industry. It really is. I mean, companies spend tens of millions of dollars on compute to train new models to achieve a or 3% gain on these exact benchmarks. If the
physical server load is shaving 5% off the top of the model's intelligence dynamically throughout the day, that is a massive hidden tax. It is. But we must acknowledge the immense difficulty of proving this isolation. The critique agent has to be absolutely flawless here, because there are massive systemic confounders. Confounders being other variables that could cause the model to fail under load, unrelated to our
floating-point math theory. Precisely. A busy inference server is a highly chaotic environment. You don't just have dynamic batching altering the matrices. You have extreme memory pressure. You have request queues piling up, leading to latency timeouts. And most importantly, you have KV cache dynamics. Let's unpack the KV cache, because it's a critical piece of transformer architecture. KV stands for key value. In an
attention mechanism, when an LLM generates a token, it computes specific matrices, keys, and values that represent that token's relationship to the rest of the text. Instead of recalculating those massive matrices for every single previous word every time it wants to generate a new word, it saves them in a dedicated memory bank called the KV cache. It's like a student keeping their scratch paper on the desk so they
don't have to re-derive the formula for every question on the test. It's a massive speed optimization. It is essential for autoregressive generation. But the KV cache consumes enormous amounts of GPU RAM. Under high server load, with hundreds of users generating long sequences simultaneously, that memory fills up instantly. Right. The inference engine has to aggressively manage this memory, sometimes paging it out to
slower CPU RAM, or evicting older tokens, or employing techniques like page attention to fragment the memory blocks. So the server is violently shuffling memory around just to keep the system from crashing. Exactly. And this cache pressure can alter timing, trigger memory bandwidth bottlenecks, and force entirely different execution paths. A model might fail a coding task under high load, not because of floating
point drift, but because the KV cache thrashing caused a context retrieval error. Ah, I see. So the swarm has to comprehensively log all of this. It has to track GPU memory utilization, cache eviction rates, and paging latency. It has to mathematically isolate the floating point drift and prove that the 5% drop in accuracy is genuinely tied to the autoregressive amplification of the math, and not just the server
choking on RAM. Moomin is very explicit about the falsification boundaries here. If the swarm fails to observe an accuracy difference beyond the predefined margin after controlling for those confounders, then ISDT is falsified for that specific tested regime. The theory welcomes its own destruction if the data doesn't hold. That's good science. So we have mapped the microscopic physics of the floating point
truncations, traced their amplification through the autoregressive loop, defined a weather index to measure the damage, and architected an AI swarm to conduct the experiments. It's a comprehensive framework. It is. To bring this deep dive together, we have to look at the macro implications. If Moomin's swarm runs these tests, and ISDT is proven 100% true, what happens to the tech industry? Well, it triggers a
fundamental reevaluation of how we benchmark and trust these systems. Currently, the industry treats AI models as static artifacts. A company releases a model and publishes a paper stating it scores 92% on the human evil coding benchmark. And we take that number as a permanent attribute of the model's architecture. But if ISDT is validated, those benchmark scores are essentially meaningless without context. I mean,
if they ran that benchmark on an isolated server rack on a Sunday night, the model looks like a genius. But if a user deploys that exact same model into a production environment with massive concurrent traffic, the autoregressive drift engages, the PWI plummets, and the model suddenly performs like an older, less capable version of itself. Benchmarks will be forced to become load-aware. We will have to decouple the
theoretical capability of the neural network weights from the physical realities of the deployment infrastructure. It really reminds me of the concept of hidden technical debt in machine learning systems. Yes, you are referring to the seminal paper by Scully et al. from Google. That's the one. They argue that the actual machine learning code, the elegant neural network, is only a tiny fraction of a real-world AI
system. The vast majority of the system is the complex, messy plumbing, the serving infrastructure, the data pipelines, the load balancers. ISDT is the ultimate manifestation of hidden technical debt. It proves that the model's intelligence is inextricably bound to the physical stress of the silicon it runs on. The server is the intelligence. And this raises an operational reality that developers are going to have to
confront. Mumman touches on the concept of operational fairness at the end of the paper, doesn't he? He does. Because if inference load reliably degrades performance, we are faced with a systemic disparity in service quality. Think about a global user base. If you are a developer in a time zone that happens to align with the global peak usage hours, and you are using an AI assistant to help you write software, you
are querying the model when the dynamic batch sizes are massive and the floating point drift is severe, and you are paying the exact same API costs or subscription fees as a developer in a different time zone who queries the model during off-peak hours. But you are receiving a fundamentally degraded product. You are receiving code that is statistically more likely to contain microscopic syntax errors, due to
threshold crossings and decision boundaries. You are getting a worse AI purely based on the physical traffic of the data center at that moment, and you have absolutely zero visibility into that degradation. This is where the prompt weather index evolves from a theoretical diagnostic into a necessary observability tool. Transparency becomes essential. Imagine your developer environment having a real-time PWI widget in
the corner of the screen. Before you ask the AI to refactor a massive precision sensitive code base, you glance at the widget. If it shows severe inference degradation, a low PWI, you know the underlying math is drifting. You might decide to hold off on the prompt. Yeah, or you might switch to a more expensive dedicated server instance for that specific task. As Muehman concludes, monitoring tools like the PWI may
complement, but they do not replace, the eventual need for batch invariant kernels. But until the hardware catches up, or until companies are willing to absorb the massive compute costs of enforcing sequential math, exposing the weather to the end-user is the most viable path forward. It is a remarkable journey we've taken through the inference stack today. I mean, we started with the physical realities of server
economics, watching dynamic batching alter the lowest level execution paths of GPU calculations. We dove into the binary mechanics of IE754 floating point standards, watching bits fall off the edge of memory registers. Right, and we tracked those microscopic rounding errors as they fed back into the autoregressive loops of language generation, compounding and altering the fundamental meaning of the text. We examined
the prompt weather index, breaking down exact match rates, high-dimensional semantic embeddings, and task-specific failure vectors. And we envisioned an agentic swarm, a team of specialized AI agents working tirelessly in the background to statistically validate the health of their peers. We confronted the reality that the intelligence of our tools is bound to the physical stress of the infrastructure. I want to
leave you with a final thought to mull over today. Go for it. And I think we're going to have a lot of fun. As humans, we have an overwhelming tendency to anthropomorphize artificial intelligence. When an LLM gives us a brilliant, insightful answer, we think of it as "smart." When it hallucinates or gives us a weird, buggy response, we assume the model is having a bad day, or that the training data was flawed. Right,
we project human traits onto it. But what if its bad days aren't conceptual at all? What if they are just the hidden physical reality of finite computational space? What if the mood of the AI is nothing more than the flow of the world, or the floating point friction of a million other people asking questions at the exact same millisecond? It changes how you see the tool completely. It really does. So, the next time
you are staring at your screen and an AI writes a bizarrely buggy piece of code for you, or messes up a simple logic puzzle, don't just assume the neural network failed. Stop and ask yourself, is the model actually hallucinating, or did you just happen to catch it standing in the rain of a bad
Chercher et filtrer l'archive
Chercher par titre, description ou catégorie. Chaque entrée est un artefact concret du système : un article, un cadre, ou une étude appliquée.
Catégories de l'archive
Chaque catégorie montre une facette différente de la même méthode. Les fils de recherche nourrissent les articles, les articles se consolident en œuvres longues, les études appliquées prouvent la chaîne sur des problèmes réels.