The Voynich Manuscript as a Medieval Database
A podcast adaptation discussing proposed structural interpretations of the Voynich Manuscript, not an established decipherment.
Listen to episodeEpisode transcript (118 paragraphs)
I want you to picture a scene for a moment. It's the late 1940s, right? Right after the war. Exactly. Right after the end of the Second World War. And we're in Arlington Hall in Virginia, which was the headquarters for the U.S. Army's cryptanalysis efforts. The absolute epicenter of code breaking. Yeah. And the men and women sitting at these desks, I mean, they are undisputed geniuses. They were the very people who
just shattered the Japanese purple cipher. They helped crack the Enigma machine. Right. They have mathematically deconstructed the most sophisticated, you know, high stakes military encryption on the planet. And now that the war is won, they're handed a new assignment. It sounds kind of like a victory lap, honestly. Like, you know, give the ultimate code breakers a medieval puzzle just to keep their minds sharp. That
was probably the exact thought process. So they sit down with this book. It's an old book written on vellum dating back to the early 15th century. And they fully expect to just dismantle it in, I don't know. A matter of weeks, maybe months. Right. Because what's a medieval scribe compared to the German military? Exactly. But days turn into weeks and weeks turn into years. And these world class battle hardened
cryptanalysts just find themselves staring completely blankly at faded pages covered in an alien script they can't read. Along with those bizarre watercolor illustrations, right? Oh, yeah. The naked women sliding down green plumbing tubes and disembodied roots of plants. That do not exist anywhere on Earth. I mean, they throw everything they have at it. Frequency analysis, substitution matrices, every trick that
brought down the axis powers. And they get absolutely nowhere. Total brick wall. They fail completely. It really sets the stage, doesn't it? Yeah. Because that book is the Voynich manuscript. Yeah. And that post-war failure is really just one chapter in a long, honestly, incredibly frustrating history of the world's smartest people hitting a brick wall. Yeah. Which is why we are here today. Welcome to a highly
requested deep dive. Today's mission for you, the listener, is to wade into the ultimate diagnostic muddy waters. We are tackling the Voynich manuscript, but, and this is key, we are doing it through a completely new lens. A very modern lens. Right. We are exclusively examining two recent, absolutely groundbreaking draft papers from March by an independent AI researcher named Robel Moomin. And what we're going to
explore today is a framework that suggests everyone, I mean everyone, from the 15th century scholars to the 1940s NSA code breakers, they might have been asking the fundamental wrong question. They were looking for meaning when they should have been looking for structure. Exactly. Okay. Let's unpack this. Because if you love a good aha moment, viewing this centuries old mystery through the lens of modern data
architecture is going to completely change how you view history. The core insight we need to explore today is honestly a paradigm shift. For over a century, the assumption has been that the Voynich manuscript is a direct language or maybe a linguistic cipher. You know, you look at a word, you try to find its equivalent meaning in Latin or English or Arabic. Right, like a one-to-one translation. Yeah. But Moomin
suggests it might be something else entirely. He proposes it's a highly structured recording system, basically a database. And the proof doesn't lie in translating the vocabulary. The proof lies in the rigorous mathematical syntax of where the ink actually hits the vellum. Before we get into the heavy math, though, let's just make sure we're all looking at the same picture. Because for anyone who hasn't fallen down
this particular rabbit hole, the Voynich manuscript is a physical book. It's a 240-page illustrated codex. And it's old. Very old. The vellum, you know, the prepared animal skin it's written on, has been radiocarbon dated to between and 1438. And it currently lives in the vault at Yale University's Benecke Rare Book and Manuscript Library. And visually, I mean, it is a masterpiece of the bazaar. The text is written
left to right in this looping, really elegant script that appears literally nowhere else in the historical record. It's not a known alphabet. Just totally unique to this one object. Exactly. But what makes it so captivating and honestly so confusing are the illustrations. You have this massive botanical section featuring highly detailed drawings of hybrid plants, like a botanist stitched together the leaves of one
species, the stem of another, and the roots of a third. Like medieval Frankenstein plants. Right. You have astrological charts. And then you have the famous balneological section, which is just pages and pages of tiny nude figures bathing in these elaborate interconnected pools and aqueducts. Which leads perfectly into why people have been so desperate to read it. You look at those drawings and you think, okay, there
is secret knowledge here. But let's talk about the history of attempts to extract that knowledge. Because understanding the failures is really the only way to appreciate the fact that there is a secret knowledge here. And I think that's why we need to appreciate why Moeman's new approach is so radical. Yeah. The failures generally fall into two major camps. The first camp treats the manuscript as a natural language
hidden beneath a substitution cipher. You know, you see a shape that looks like a bench and you just assume it stands for the letter T. And a loop is an O, that kind of thing. Exactly. Researchers have spent lifetimes trying to map these shapes onto existing languages. I mean, recently, in 2019, a researcher named Gerard Cheshire made international headlines claiming the manuscript was really a work of science. And
he said it was written in a proto-romance language. Oh, I remember that, making the rounds online. People were saying the mystery was finally solved, like case closed. It was heavily publicized, yeah. But it was almost immediately dismantled by the linguistic and cryptographic communities. Why? What was wrong with it? The methodology was highly circular. It relied on taking a string of Voynich characters and
essentially squinting at them, shifting vowels around, and digging through dozens of different Romance language dictionaries until you found a word that kind of sort of fit the picture on the page. Oh, wow. So just cherry picking, basically. Yeah, it wasn't a systematic decipherment at all. It was just an exercise in confirmation bias. So if it's not a hidden language, what's the second camp? The skeptics, I assume.
Right. The second camp is the hoax theory. The idea that if it can't be read, well, maybe there's nothing to read. The most prominent version of this came from Gordon Rugg in 2004. He argued the entire manuscript was a meaningless medieval scam generated mechanically using something called a "cardan grill." I've heard of that, but how does a cardan grill actually work in practice? Think of a cardan grill as like a
piece of stiff parchment or cardboard with little rectangular windows cut out of it. You place this grill over a master table of random syllables and prefixes. The windows reveal a seemingly random string of characters. Okay, so you just copy what you see in the little windows. Then you write the characters down, then you rotate the grill or slide it down a line and write down the new combination. It allows a single
person to generate hundreds of pages of convincing-looking gibberish very quickly. Rugg argued a 15th century fraudster did this to sell a mysterious "magical book" to a gullible nobleman. I mean, it sounds really plausible on the surface. If you're trying to fake a 240-page book, you'd want an automated way to do it, right? So your brain doesn't just start repeating the same three fake words over and over. But I'm
guessing modern computers had something to say about that. Oh, they absolutely did. Later statistical analyses applied rigorous tests to the text and completely undermined the hoax theory. Because real languages follow specific mathematical laws. Like Zipf's Law. Yes, Zipf's Law. It dictates that in any natural language, the frequency of any word is inversely proportional to its rank in the frequency table. The most
frequent word will occur approximately twice as often as the second most frequent word, or as often as the third, and so on. Like the word "the" in English. It's everywhere. While the word "accordion" is incredibly rare. Exactly. And the Voynich text adheres remarkably well to Zipfian distributions. But beyond that, it exhibits long-range word correlations. In a real book, if I use the word "bishop" in chapter 1,
there's a higher statistical probability I will use the word "bishop" again in chapter 2. Because the topic is continuing. Right, the context carries over. The mechanically generated gibberish using a cardan grill doesn't do that. A grill has no memory. It can't carry a theme across ten pages. The Voynich text, however, has long-range structural memory. So we are stuck in this brutal middle ground. The statistics say
it's way too complex and structured to be a random hoax. But our absolute best linguistic efforts say it doesn't map to any human language we know. I mean, how do you even begin to study a text like that in 2026? Obviously, you can't just feed photos of a 15th century book into a spreadsheet. No, you need a bridge between the physical vellum and digital data. That bridge is transliteration. Over the decades,
researchers developed systems to convert the bizarre Voynich squiggles into standard machine-readable Latin characters. So the computer can actually process them. Right. The gold standard today is called EVA, the extensible Voynich alphabet. Specifically, Moomin's research relies on the Takahashi version of the EVA transliteration. Let's get a quick visual for that. If a scribe drew a shape that looks, you know, a
bit like the number 8, the EVA system might map that to the letter D on a keyboard. A shape resembling a tiny gallows becomes a K. Exactly. We aren't saying the gallows' shape actually sounds like a K. We're just giving the computer a label it can understand so we can run the math. That is a crucial distinction. EVA is not a translation. It is merely a digital transcription of shapes. But once you have those 37,000
words transcribed into a clean data set, you can start running heavy computational work. And this is where Roble Moomin enters the picture with his semantic pose-based cognitive encoding hypothesis, the SPBCEH. It's a dense acronym, for sure. But what he's proposing is a total fundamental shift. Because he looked at the graveyard of failed translations and realized everyone was making the exact same assumption. They
were assuming that a cluster of letters bounded by spaces was a traditional word-like, a noun, a verb, an adjective. Right. And he realized that trying to assign semantic meaning to these clusters was a trap. If you don't have the dictionary, staring at the letters simply won't help you. So he threw out the concept of meaning entirely. Just tossed it out the window. Completely. He isolated the most highly frequent
EVA glyph clusters in the manuscript. To give you a sense of scale, these clusters alone make up nearly 26% of every single stroke of ink in the entire 240-page book. Wait, really? That is a massive footprint. And what those clusters are doing, you've basically unlocked a quarter of the book. And Mooman's question wasn't, "What do these clusters translate to?" His question was, "What functional role do they play
based on their positional asymmetry?" Let's slow down and define that term. Positional asymmetry. What exactly are we looking for when we look for that? We are looking at the literal physical placement of a cluster within a line of text. If a cluster of characters is a standard noun, like the word "tree," you would expect to see it scattered fairly randomly throughout sentences. Sometimes it's at the beginning,
sometimes the middle, sometimes the end. It's symmetrical. It just pops up wherever the grammar needs a tree. Exactly. But if a word only ever appears in a very specific physical spot, say, it only ever appears as the absolute first word on a line and never anywhere else, that is positional asymmetry. I see. In English, a capital letter has positional asymmetry. It always sits at the start of a sentence. A period has
positional asymmetry. It sits at the end. So if Moomin finds clusters that strongly prefer specific positions, he's not finding nouns. He's finding the punctuation. He's finding the invisible scaffolding. He is finding the syntax. He is isolating the structural markers that dictate how the information flows. And by analyzing the positional behavior of these high-frequency clusters across all 37,000 words, Moomin
grouped them into a provisional set of functional categories. In his first paper, he identified seven distinct words. These are the seven roles we're going to spend a lot of time on. He named them. Let me make sure I have this list right. Initiator, closure, link, action, mode, temporal, and reference. And for the sake of the data, he labels them R1 through R7. Yes. The entire foundation of this theory rests on
understanding the behavior of the first two roles. R1, the initiator, and R2, the closure. Moomin highlights specific EVA clusters to demonstrate this. For the R1 initiator, he focuses heavily on the behavior of the first two roles. He focuses heavily on a cluster transcribed as "Coketee." And for the R2 closure, he focuses on a cluster transcribed as "Cheti." Now, intuitively, if I hear the word "initiator," I
picture the starting gun of a race. I assume that "Coketee" is going to be sitting at the extreme left margin of the page, acting as the very first word of a new line of text. Sure, that makes sense. And if "Cheti" is a closure, I expect it to be the last thing written before the scribe moves their pen down to the next physical line. That is the logical assumption. But when Moomin ran the positional data, the results
were entirely counterintuitive. The cluster "Coketee," the supposed initiator, is almost never the first word on a physical line. In fact, it occurs in the medial position, meaning sandwiched somewhere in the middle of a transcribed line, 87.9% of the time. Wow. And the closure. The cluster "Cheti" is medial 92.2% of the time. It is almost never the final word on a physical line. They are both overwhelmingly floating
dead in the middle of the text lines. Hold on. That feels like a massive contradiction. How can a researcher look at a word that sits perfectly in the middle of a sentence, nowhere near an edge, and declare with absolute certainty, "This is an opening bracket and this is a closing bracket"? Because Moomin stopped looking at the physical margins of the vellum and started looking at the sequence of the words
themselves, he analyzed every single bittogram in the entire manuscript. Let's define a bittogram for the listeners, just to be clear. A bittogram is simply a pair of adjacent words. If the sentence is "the quick brown fox," your big rims are "the quick quick brown" and "brown fox." Moomin calculated the transition probabilities between his seven rolls. If I am looking at an R2 closure cluster, what is the
statistical probability that the very next word will be an R1 initiator cluster? And the data he got back wasn't just a mild correlation, was it? It was an anomaly of historic proportions. It was staggering. The transition from an R2 closure directly to an R1 initiator, meaning a close immediately followed by an initiator, is the single most overrepresented transition sequence in the entire manuscript. To put a
number on it, Moomin calculated its z-score against a shuffled baseline, and it came back at +9.75. I think we need to linger on that number, because z-scores can sound a bit abstract. A z-score measures how many standard deviations a beta point is from the average. If you get a z-score of 0, that's perfectly average. A z-score of means something interesting is happening. In fields like psychology or sociology, a
z-score of is often considered strong evidence to publish a paper. A z-score of +9.75. I mean, in particle physics, they use a 5-sigma threshold to declare the discovery of a new subatomic particle. So a 9.75 is practically the universe grabbing you by the shoulders and screaming that this is not random chance. That is a brilliant way to contextualize it. The scribe was executing this close to a net transition with
incredible, relentless deliberation across pages. And what this proves is that this pairing defines a definitive boundary. It is a data packet boundary. And because these close to a net pairings are happening overwhelmingly in the middle of the physical lines of text, it leads to the most mind-bending realization of the whole paper. The grammar of the Voynich manuscript completely ignores the physical layout of the
page. Exactly. The packet structure wraps around the line bricks. To give you a visual analogy for this, imagine you are writing a really long, complex essay on a piece of paper. But the rule you've decided to follow is that you are completely forbidden from putting a period at the end of a physical line when you reach the right margin. Okay, I'm picturing it. You are only allowed to place a period when your actual
thought concludes. That thought might end three words into line two. It might end dead in the middle of line five. You place your period. You capitalize the very next letter to start the new thought. And you keep writing. And if an alien came down years later and tried to decode your essay, but they incorrectly assumed that a physical line of ink was the basic unit of your language, they would be hopelessly lost.
Totally lost. They would look at your periods and capital letters showing up randomly in the middle of lines and think it was gibberish. Because the physical margin of the paper has absolutely nothing to do with the structure of the information. The data packet dictates the structure. The margin is just where you ran out of room to move your hand. That fundamentally reframes what we are looking at when we look at the
Voynich. It means the scribe was essentially encoding discrete, self-contained packets of information. When a packet was finished, they explicitly closed it with an R2 cluster and immediately opened the next packet with an R1 cluster regardless of where their quill happened to be resting on the vellum. Here's where it gets really interesting though. I love this concept, but I have to play the skeptic here. Because
anyone can look at a wall of static and draw a constellation. How do we know Moomin didn't just fall into the Texas sharpshooter fallacy? You mean firing a shotgun at a barn and then painting a bull's eye around the biggest cluster of bullet holes? Yes, exactly. Couldn't a researcher just take any weird text, point at a bunch of high frequency words, label them initiators and closures, and then run a math formula
that makes it look like a brilliant discovery? That is the exact trap that destroys most Voynich theories. Finding a pattern is relatively easy. Proving that the pattern is structurally vital to the text, and not just an artifact of your own categorization, is where the real science happens. Moomin anticipated this exact criticism, which is why he dedicated a significant portion of his research to the falsification
protocol. Let's dig into that, because this is where the paper moves from interesting idea to rigorous science. How do you falsify a grammatical theory on a text you can't even read? You do it through a controlled inversion experiment. Moomin set a hypothesis: If my label's initiator and closure are actually representing real directional syntactic functions, then swapping them should completely break the structural
logic of the text. However, if my labels are just arbitrary tags I applied to random common words, then swapping them won't really matter. The math will still look relatively similar. So he builds a mirror universe. He takes his framework, and he flips it. Every R1 initiator, he has suddenly reclassified as an R2 closure. Every closure becomes an initiator. He then applies this inverted framework blindly to three
specific folios pages of the manuscript that were deliberately kept out of the initial training data so they wouldn't bias the results. He used folio F88R from the pharmaceutical section, folio F103R from the recipe section, and folio F75R from the balneological section. The results on folio F75R are incredible. Let's look closely at line on that specific page. In this line, the cluster Coquitie, the one we
identified earlier as a prime R1 initiator, appears five separate times. Under Moomin's original intended rules, reading that line is like reading a list of bullet points. It's five opening actions triggering five new packets of data. In a section that is seemingly detailing step-by-step bathing rituals or biological processes, opening five new data packets makes logical structural sense. But observe what happens
when you apply the inverted rules to that exact same line. Under the inversion, Coquitie is now a closure. That line of text suddenly represents a system executing five structural closures in a row with absolutely zero opening states preceding them. Wait, five in a row? Like just bam bam bam? Exactly. It would be like a computer programmer writing five closing brackets or an author writing five end parentheses in a
single sentence without ever opening one. It doesn't mean anything. It is structurally incoherent. It violates the basic logic of a sequential state system. The original mapping produces a coherent state transition narrative that flows beautifully with the visual context of the page. The inverted mapping produces a catastrophic syntactic pileup. This falsification protocol proves that Moomin's clusters carry true
directional functional asymmetry. You cannot swap their roles without destroying the underlying architecture of the text. Okay, so the roles are real. And they are highly directional. Once Moomin proves this, he takes the next logical leap. He maps the entire grammatical structure using something called a finite state automaton, or an FSA. Now, I'll be honest, when I first read finite state automaton, my brain
immediately glazed over with computer science jargon. But the concept is actually remarkably intuitive. It really is. Think of an FSA simply as a flowchart. It's a mathematical model of computation that defines a system with a limited number of states and strict rules for how you are allowed to move from one state to another. A turnstile is a classic example. A turnstile, like at a subway station? Yes. It has two
states: locked and unlocked. Pushing a coin into it transitions the state from locked to unlocked. Pushing the arm transitions it back to locked. Moomin designed an FSA flowchart for the Voynich packet grammar containing four distinct states. Walk us through the four states of this Voynich turnstile. State one is ideally. This is the baseline. No packet is currently open, and no data is actively being recorded. The
system is just waiting. State two is open. This state is instantly triggered the moment the text encounters an R1 initiator cluster. The boundary is set. The packet has officially begun. Then we move into state three, which is payload. This is the meat and potatoes. Once the packet is open, you get a sequence of R3 links, R4 actions, R6 temporals. This is the actual substance of the information being recorded by the
scribe. And finally, state four is closed. This is triggered the moment the text hits the ground. This is the moment the text hits an R2 closure cluster. The packet is officially sealed. And from this closed state, the strongest mathematical pull in the entire manuscript, that massive +9.75 z-score transition we discussed, immediately drags the system back to the open state with a new R1, or returns it quietly to the
ideally state. It is a beautifully self-contained loop. Idle, open, payload, closed, repeat. Moomin took this specific four-state flowchart and ran it across the entire text of the manuscript to see how strictly the ink on the page actually obeyed these invisible rules. And the results of this multi-scale validation are what cement the packet theory as a genuine breakthrough. Moomin tested the text at two different
scales. First, he tested it at the paragraph level. He prospectively set a threshold. Considering we are dealing with 600-year-old faded ink, ambiguous transliterations, and normal human scribal errors, he decided that if 60% of the transitions followed his strict FSA flowchart, the model was a success. And he beat his own threshold. At the paragraph level, 61.3% of all role transitions followed the idle, open,
payload, closed grammar perfectly. The structure holds up. But then you look at the line level conformance. If you force the FSA to evaluate the text line by physical line, the model completely bombs. Only 9.1% of the physical lines have zero structural violations. A nearly 91% failure rate at the line level sounds like a disaster for the theory. Until you investigate the specific nature of those thousands of
violations. Right. Moomin dug into those 5,934 specific line level violations. And he found they were overwhelmingly the exact same error. There were lines of text that started in the ideal state because it's a new line. But the very first word wasn't an initiator. It was a piece of payload. Or it was a closure. Which is exactly what must happen if the packet wraps around the margin. If an information packet spans
across a physical line break, then naturally the next physical line is going to start right in the middle of a packet. The computer reads the start of the line, assumes it should be in the IDLE state waiting for an initiator, but instead gets hit in the face with mid-packet payload data and it registers a violation. It's incredible. The failure at the line level is the ultimate proof of the paragraph level theory. It
mathematically confirms that the physical layout of the vellum pages has absolutely nothing to do with the actual grammatical structure of the encoding system. The packet is an independent, mathematical layer floating entirely free from the physical constraints of the book. This raises an important question though. We have a mathematically sound role model. We have an FSA proving the packet structure. The obvious
question to ask here is, is it perfect? Did Movenin crack the exact flawless syntactic categorization on his very first try? And the answer, which I think is a testament to the scientific rigor of these papers, is no, he didn't. This brings us to the core of paper 2, which involves a massive scientific pivot. Because the true test of a theory isn't just proving yourself right, it's actively trying to prove yourself
wrong and seeing what survives the fire. Movenin recognized that while the broader packet structure, the boundaries, the wrapping, was undeniable, his specific categorization of the internal payload roles might be flawed. How could he be certain that R4 action and R5 mode were truly distinct grammatical functions, and not just his own human brain imposing artificial categories on a system he didn't fully understand?
So he designs the anti-projection test. He essentially wants to force the computer to tell him if his roles are the best possible way to organize this text. He takes his original roles, and he systematically generates alternative mappings. He scrambles categories together, he splits them apart, creating models with roles, roles, roles. He tests every logical variation against his original 7-role baseline. He ran
these alternatives. He used rigorous structural scoring metrics. And the results were a humbling moment for the initial theory. The 7-role model failed to hit the top in any metric. It ranked 6th out of on classification accuracy, and it ranked 7th out of on Markov transition structure. Let's pause on that term, Markov transition structure. What exactly is the computer measuring when it scores that? A Markov model is
a way to describe a sequence of possible results. A Markov model is a sequence of possible events where the probability of each event depends entirely on the state attained in the previous event. Think of it like predicting the weather using a very simple rule. If it is raining today, there is a 70% chance it will be cloudy tomorrow, and a 30% chance it will be sunny. So you don't need to know the whole history.
Right, you don't need to know what the weather was last week. You only need the current state to predict the next state. So Muvins is asking the computer, "Which of these different role mappings creates the most predictable, reliable weather forecast?" For the grammar of the Voynich manuscript. Precisely. And the computer returned a clear winner. The highest scoring alternative across every metric was a mapping
called QR-AFT_09. And the specific change that QR-AFT_09 made was fascinating. It took Mumin's 7-role model, and it forcefully merged R4, which Mumin called "action," and R5, which he called "mode," into a single, unified category. It told him that the data simply didn't support distinct categories. The statistical behavior of the words in the action bucket and the mode bucket overlapped too heavily. They weren't
distinct grammatical gears. They were just two slightly different flavors of the exact same thing: the inner payload content. So Mumin merged them into a new, consolidated R4, which he simply named "content." We now have a 6-role model. This is a classic example of model parsimony in science. The simplest explanation that accurately fits the data is usually the correct one. Mumin needed to prove mathematically that
merging these two massive categories wasn't just muddying the waters. To do this, he utilized an entropy decomposition test. Entropy is one of those terms that gets thrown around a lot in pop science, usually referring to chaos or the heat death of the universe. But in information theory, it has a very specific meaning, right? In information theory, entropy is simply a measure of unpredictability, or the amount of
surprise in a sequence of data. If I have a biased coin that lands on heads and not on the bottom of the coin, then I'm going to have to use a very specific method to get it to the bottom of the coin. If I have a biased coin that lands on heads 99% of the time, the entropy is very low, because you almost always know what's going to happen. If I have a fair coin, the entropy is high, because the next flip is
completely unpredictable. So how does Mumin apply this coin flip logic to the grammar of the Voynage? He separated the 6-role text into two distinct layers. The first layer is the structural scaffolding: the initiators, the closures, the links. The second layer is the variant payload: the actual content words. In any highly structured recording system, you would expect the structural scaffolding to have low entropy.
The rules of punctuation should be highly predictable. Because structure needs rules. Yes. But you would expect the content payload to have high entropy, because the subject matter being recorded should be highly variable. I see. It's like looking at the blueprint for a modern office building. The structural entropy is incredibly low. The steel I-beams, the load-bearing walls, the elevator shafts, they always follow
strict engineering codes. They go in predictable places. But the variant entropy is sky high. Because once the floors are built, one company might fill their office with cubicles, another might build a yoga studio, and another might pack it with server racks. The structure is rigid and predictable. The content is infinitely variable. That is a perfect analogy. And when Mumin ran the math on the new 6-role model, it
passed the blueprint test perfectly. The structural entropy of the markers dropped to a highly predictable 2.38 bits. Meanwhile, the variant entropy of the new, merged content role rose to a highly unpredictable 2.77 bits. Wow. So the split is really clean. Very clean. The 7-role model had actually failed this specific predictability test. The 6-role model mathematically aligned with the expectations of a real
encoding system. What's fascinating here is what didn't change. Even after smashing two huge categories of words into one and rerunning all the massive calculations, that core heartbeat of the manuscript didn't flinch. The R2 to R1 packet boundary transition, that definitive closing bracket, followed immediately by an opening bracket, survived the merger completely intact. The z-score barely shifted a fraction,
sitting rock solid at +9.71. That is the hallmark of a true underlying phenomenon. If the packet boundary was just an illusion created by bad math, merging the payload categories would have shattered it. The fact that the boundary signal survived proves that the core grammar of the Voynich is robust. It is not an artifact. It is reality. So we have emerged from the deep mathematical weeds with a rock solid, 6-role
grammatical engine. We know the text is broken into discrete data packets that wrap around physical lines. The next obvious step is to take this invisible mathematical filter we've discovered and lay it over the actual physical pages of the book. We had to look at the pictures. This is where Moment introduces what he calls the domain arc. And this is where the math suddenly breathes life into the visual strangeness
of the manuscript. He analyzed how the different illustrated sections of the book utilize these six grammatical roles. And he found that the grammar changes dramatically depending on what the scribe is drawing. Let's start with the most infamous section. The bound neological or biological section. Again, these are the densely packed pages featuring illustrations of naked women, often with swollen bellies, bathing in
interconnected green pools, holding mysterious objects with elaborate pipework flowing through them. It looks almost like an alchemical diagram of human biology. What does our mathematical filter reveal when we lay it over these pools? The math goes absolutely wild with structural boundaries. In this section, the R1 initiators and R2 closures absolutely dominate the text. Initiators make up 12.3% of all classified
tokens, and closures make up 11.1%. To put that in perspective, that is three to five times higher than the boundary density of any other section in the entire book. Three to five times higher. Yeah. That is a massive deviation. Muehman notes that 23.1% of the paragraphs in this specific section are perfectly sealed, meaning they start with a pristine initiator and end with a definitive closure. He describes it as a
boundary-intensive recording environment. The syntax is highly fragmented. Why would a scribe do that? If you look at the illustrations, it starts to make sense. If this section is a medical ledger or a step-by-step recording of a chemical ledger, you are not writing long-flowing philosophical essays. You are logging rapid-fire state changes. Exactly. Open packet. Subject enters pool A. Close packet. Open packet.
Subject interacts with pipe B. Close packet. Right. The heavy use of boundaries perfectly matches a procedure of discrete sequential data entry. The syntax mirrors the function. Now let's contrast that rapid-fire data entry with the exact opposite end of the grammatical spectrum, the zodiac section. The zodiac pages are visually similar to the data entry. The pages are visually stunning. You have a central
illustration of a recognizable zodiac sign like a bull for Taurus or scales for Libra. And radiating outward from that center are concentric rings of stars and tiny human figures with text written in unbroken continuous circles wrapping around the entire page. When you apply the six-roll mathematical filter to these circular pages, the structural profile completely flips. In the zodiac section, the R1 initiator rule
almost ceases to exist. It drops to a negligible 0.5% of the text. The closures vanish as well. Instead, the R3 link rule violently spikes and dominates the grammar. And the sheer brilliance of this is understanding why the math flips. Why would a scribe suddenly stop using opening and closing brackets in this section? Look at the geometry of the page. The text is written in a literal physical circle. The circle has
no linear beginning and it has no definitive end. You cannot have a definitive opening tag and closing tag when the sentence is an infinite loop. So the underlying packet grammar intelligently adapts to the visual domain. It suppresses the boundary markers and relies almost entirely on R3 links to weave the continuous stream of astrological data together. The scribe is fundamentally entering the syntax of the
language to accommodate the physical geometry of the astrological charts. That is mind-blowing. What about the massive herbal and pharmaceutical sections? These are pages dominated by single giant illustrations of the world. Giant illustrations of strange plants or rows of apothecary jars. In those botanical sections, the boundary markers drop back to a moderate level and the new combined R4 content role becomes the
undisputed king. It dominates the text at 11.3% and 11.8%. Which again makes perfect logical sense. If you are a medieval botanist trying to describe the complex hybrid root system of a massive plant or listing the intricate ingredients contained within a pharmaceutical jar, you don't need rapid-fire bullet-proof text. You don't need rapid-fire bullet-point boundaries like the bathing section. You need descriptive
depth. Yes, you need long payloads of highly variable information. You need content. The correlation between the mathematical grammar and the visual subject matter is so strong that Moehmann put it to a blind machine learning test. He built a leave-one-folio-out classifier. He essentially blinded an AI to the illustrations. He fed the computer only the six-dimensional roll frequency vectors for a given page. So the
computer only sees the math. It just sees this page has 12% initiators and 4% content. It doesn't know if there are bathtubs or stars drawn on the vellum. Precisely. And just by analyzing the mathematical rhythm of those six rolls, the classifier could predict which illustrated section of the book that page belonged to with 64.7% accuracy. That's a huge jump. Yeah, when you consider that a baseline system simply
guessing the most common section every time only achieves 57.8%, this is a highly significant step. This is a highly significant signal. So what does this all mean? Let's take a step back and really digest what this domain arc reveals. The scribes who wrote the Voynich manuscript weren't just writing a stream of consciousness diary. They possessed a highly sophisticated domain-specific encoding architecture. They can
consciously deployed one syntactic structure to analyze a plant. They dynamically altered that structure to record circular astrological cycles. And they used a totally different, segmented structure for the rapid-fire data logging in the plumbing diagrams. Which forces us to confront the ultimate million-dollar question: If the manuscript possesses this deeply structured, dynamically adapting packet grammar, what is
the Voynich manuscript? Let's return to those competing hypotheses from the beginning of our deep dive. What does SBBCEH do to the hoax theory? I feel like this framework systematically dismantles the card and grill theory. A piece of cardboard with holes cut in it doesn't know that it's sitting over a drawing of a star. It's sitting over a drawing of a circle, and therefore needs to mathematically suppress its
generation of R1 initiator clusters. A mechanical hoax generator cannot dynamically alter its syntactic output based on visual context. That is the primary defense against the hoax theory. But Mummin introduces a secondary, morphological line of evidence that I think drives the final nail into the coffin of the Ho's hypothesis. Morphology refers to the internal structure of the words themselves, right? How the
letters are built together. Mummin looked at the internal spelling of the specific clusters assigned to these grammatical roles. And he found stunning consistencies. Almost all of the active R1 initiator clusters share the exact same EVA prefix. This specific prefix is overwhelmingly exclusive to the initiator role. It practically never appears in closures, links, or content words. It functions almost like an
invisible structural tag. If you see "quack," you know a packet is opening. And the R2 closures are just as rigid in their morphology. They almost entirely reduce down to combinations of just three consonant stems, "sh" and "elch," paired with four specific suffixes, "80, 80, and 80." So a closure is essentially built from a menu. Pick a stem, pick a suffix, seal the packet. This internal morphology proves that we
are looking at a sublexical structural distinction. This is a grammatical morpheme system, highly consistent with the mechanics of natural languages. A 15th century fraudster using a mechanical grid generator would not perfectly, accidentally pair specific words to mathematical paragraph initiators and specific suffixes to closures with 90% positional accuracy across pages of text. It requires a level of intentional
structural complexity that precludes a simple hoax. Exactly. There is also this fascinating, hyper-specific detail in the paper about two tiny clusters: "kriel" and "kriel." I want to touch on this because it feels like peeking at the source code of the manuscript. It is a remarkable finding. It is a remarkable finding. It is a remarkable finding. Boomin noted that the cluster eel is classified under the R6 reference
rule. But there is a very similar unclassified word: "kriel." When he mapped their frequencies across the different sections, he found that they have a section profile Pearson correlation of 0.940. For context, a Pearson correlation measures how closely two variables move together. A score of 1.0 means they move in perfect, lockstep identical motion. A score of 0.940 means these two words are practically joined
together. A score of 1.0 means these two words are practically joined at the hip. Where one goes, the other follows. Right. Primarily spiking together in the balneological bathing section. It suggests a deeply intertwined syntactic relationship. Boomin hypothesizes that "kohl" might be an inner packet function word, structurally equivalent to the reference word "ole," but designed to operate exclusively deep inside
the payload, rather than near the boundaries. It reminds me of writing code for a website in HTML. You have your structural tags that define the scaffolding, the head and body tags, and then you have highly correlated inner tags that only operate within that specific scaffolding. The "ole" and "kohl" relationship shows a level of nested syntactic depth that, again, points violently away from a hoax and strongly
toward a real intentional encoding system. So if it's not a hoax, what are we left with? Is it a language? Yeah. Has Moomin finally cracked the 600-year-old cipher? This is where we have to be incredibly intellectually honest about the interpretive scope of these papers. Because the internet loves a sensationalist, the internet loves a sensational headline. But Moomin is exceptionally clear. He does not claim to have
translated a single word of the Voynich manuscript. He has not provided a dictionary. He cannot tell you what the bathing women are doing, what the hybrid plants are used for, or what the astrological charts predict. He has not mapped the semantic meaning. What he has done is mapped the skeleton. He's established what the grammar does, not what the grammar means. Which leaves us with two incredibly tantalizing
possibilities for what this book actually is. Possibility one: It is a natural language cipher, but one that possesses a highly visible, incredibly rigid grammatical skeleton of tags and brackets that explicitly wrap around the phonetic vocabulary. And possibility two, which is perhaps even more fascinating, it isn't a linguistic cipher at all. It is a structured non-linguistic recording system, a medieval database,
an early spreadsheet encoded on animal skin. In this scenario, the clusters aren't words meant to be spoken aloud. They are abstract metadata tags indicating the start of a medical record, the nature of a biological variable, and the end of an entry. We have discovered the underlying syntax of an alien system without knowing its vocabulary. We know exactly where the capital letters and the periods are, we know how
the paragraphs wrap, but we remain completely blind to the story being told. And that leads to the ultimate legacy of this research, what Mummin calls the negative constraint. This framework acts as a permanent mathematical framework for all future Voynich research. I love the negative constraint, because every two or three years, a new amateur sleuth or academic makes global headlines claiming they finally
translated the Voynich. They say it's abbreviated Latin or ancient Turkish or a lost dialect of proto-Romance. And usually it takes experts months to untangle their logic and debunk it. But SPBCH changes the game. We now possess a mathematical truth. Any future decipherment attempt must account for the "coq" prefix exclusively initiating data packets. It must account for the three-stem closure morphology. And above
all, it must account for the massive close-to-init transition asymmetry that completely ignores physical line breaks. If your translation theory claims that the cluster "coqui" means the green leaf, but coqueting mathematically functions as a nonlinear structural packet initiator 87% of the time, your translation is demonstrably wrong. The math doesn't care about your linguistic theory. If a proposed translation
cannot comfortably fit inside this six-role structural grammar, it can be immediately and confidently discarded. It is a profound step forward for historical cryptography. We still haven't read the book. But for the first time in six centuries, we finally understand how the book was built. Let's synthesize the journey we've been on today. We started in the murky waters of a 600-year-old mystery, looking at an
impenetrable block of weird text that humiliated the greatest codebreakers of the 20th century. And by fundamentally shifting our perspective, by abandoning the desperate search for semantic definitions and instead rigorously analyzing positional transitions and positional asymmetry, Robomuman has unearthed a hidden six-role grammatical engine. A discrete, packet-based data structure that behaves more like modern
network communication protocols than a medieval diary. We saw how these invisible data packets wrap seamlessly around physical lines of ink, and how the underlying mathematical rules and the underlying mathematical domain profiles adapt perfectly to the visual reality of the illustrations, from the rapid-fire boundary logging required in the plumbing sections to the continuous unbracketed linking necessary for the
circular zodiac rings. The broader lesson here for you, the listener, extends far beyond a single medieval artifact. It's a lesson in how to approach the incomprehensible. Whether you are facing an ancient manuscript, a massive unorganized data set at work, or an overwhelmingly complex problem in your personal life, the instinct is usually to look harder at the details, to try and force a translation when you don't
possess the dictionary. Exactly. But the trick isn't to stare harder at the confusing parts. The trick is to zoom out. Look at the spaces between the details. Structure often speaks so much louder than content. How things connect, how they transition, and the boundaries they respect are usually far more revealing than what those things actually are. It is about finding the underlying architecture beneath the noise.
Let me leave you with a final, lingering thought to mull over. We've spent the last hour exploring a system where a 15th century scribe was manually encoding information in discrete, tagged packets, complete with highly distinct opening and closing metadata. This was happening centuries before computer scientists conceptualized data packets for modern Internet protocols. If that level of sophisticated, modern
cognitive architecture was operating in the 1400s, what other impossible systems are sitting in ancient, dusty archives right now? What other unsolvable historical mysteries are just waiting for someone to stop trying to read them and start decoding their structure? We'll see you next time on the Deep Dive.
Archiv durchsuchen und filtern
Suche nach Titel, Beschreibung oder Kategorie. Jeder Eintrag ist ein konkretes Artefakt des Systems, ein Papier, ein Framework oder eine angewandte Studie.
Archivkategorien
Jede Kategorie zeigt eine andere Seite derselben Methode. Forschungsstränge nähren Papiere, Papiere verdichten sich zu Langtextwerken, angewandte Studien belegen die Pipeline an realen Problemen.