ModMess and SuperFizzBuzz as evolutionary search benchmarks
Draft of 2026.10.02
May include: genetic algorithms ↘ benchmarks ↗ &c.
I have big piles of old notebooks and electronic notes from meetings and seminars and all kinds of just sitting around thinking up stupid shit. This is about a relatively simple problem, which might have been a benchmark idea, or might be something somebody else told me about, or might have been from a seminar?
Later: I am 55% convinced this came from a chat with Terry Jones about his tunable RS Landscapes. Explaining below.
It has no identifying information in that regard, except being from between 2001 and 2004, when other pages in the same notebook were dated. Tell me again about AGI and superintelligence.
Anyway, it looks from context like it was intended for comparing various genetic algorithms, simulated annealing, and other metaheuristics for population-based parametric optimization. There are two variants I can glean here: let’s call them ModMess and ModMess2.
The ModMesses
In both ModMess and Modmess2, a solution (aka “genome”) is an array of G integers, each in the range (1...R) inclusive. In my notes, I stuck to G = 16 and R = 100, but you do you.
A specific landscape for search is defined by picking a specific (uniform random) target genome, and the goal of search is to identify the target genome’s values of the genes in all G positions.
Slightly confusingly, and input vector (or training case) is also an array of G integers in (1...R). For any given input vector, a genome (including the target genome) is used to produce an output vector using the algorithm shown below. To evaluate an arbitrary individual at a specified input vector, we produce the output vector for that individual and return the summed absolute error compared to the output vector produced from the same input applied to the target genome.
How that works, in words:
- We take the input vector of numbers, and construct a new vector of the same length, where each element is the running sum of the preceding values in the input.
- We take these sums and collect a new vector of the remainders obtained modulo the genome’s value at each position.
- If we’re doing
ModMess2, we then take these remainders and collect a vector of them modulo the gene value at the position just to the right (wrapping back to the head as needed). - We compare this vector of integers to the target vector (summed absolute error) to determine the error score.
I asked Opus 5.5 to write me a little legible python version of this for explanatory purposes, and when it calmed down enough to hear me it produced this, which isn’t too bad, right?
""" How a genome is evaluated in ModMess and ModMess2. genome, target and each input are lists of G integers in 1...R (G = 16, R = 100 in our experiments). The target scores 0 on every input. """ from itertools import accumulate def modmess_error(genome, target, inputs): sums = list(accumulate(inputs)) # running sums of the input mine = [s % g for s, g in zip(sums, genome)] # each sum mod the genome's gene there goal = [s % t for s, t in zip(sums, target)] # ... and mod the target's gene there return sum(abs(a - b) for a, b in zip(mine, goal)) # summed absolute difference def modmess2_error(genome, target, inputs): sums = list(accumulate(inputs)) # running sums of the input next_genes = genome[1:] + genome[:1] # gene p + 1, the last wrapping to the first next_targets = target[1:] + target[:1] mine = [(s % g) % h for s, g, h in zip(sums, genome, next_genes)] # mod the gene, then mod the next gene goal = [(s % t) % u for s, t, u in zip(sums, target, next_targets)] # (0 mod x is 0) return sum(abs(a - b) for a, b in zip(mine, goal)) # summed absolute difference def case_errors(error, genome, target, training_inputs): # one error per training case (what lexicase uses); their sum is what # tournament and fitness-proportionate selection use return [error(genome, target, inputs) for inputs in training_inputs] # Worked example, G = 5: # target [19, 67, 30, 48, 75], genome [14, 48, 84, 89, 60], input [33, 95, 11, 78, 37] # running sums [33, 128, 139, 217, 254] # ModMess target remainders [14, 61, 19, 25, 29] # genome remainders [5, 32, 55, 39, 14] error 9 + 29 + 36 + 14 + 15 = 103 # ModMess2 target remainders [14, 1, 19, 25, 10] # genome remainders [5, 32, 55, 39, 0] error 9 + 31 + 36 + 14 + 10 = 100 if __name__ == "__main__": target, genome, inputs = [19, 67, 30, 48, 75], [14, 48, 84, 89, 60], [33, 95, 11, 78, 37] print(modmess_error(genome, target, inputs), modmess2_error(genome, target, inputs)) # 103 100
And… that’s about it.
ModMess is relatively easy to search over, since each gene can be independently optimized. The structure of the landscape of each is still a bit up-and-downy, but it’s not complicated for any reasonable search method.
As far as I’m aware, there’s only one solution in ModMess.
ModMess2 however has a little frisson of nonlinearity thrown in. You take a remainder mod another value from a different gene. So the two genes affecting the outcome, that’s nonlinear epistasis in terms of gene effects. And the fact that if the number you mod by second is bigger than the one you used first, well nothing will happen. But if it’s lower, then something about how different they are will have a big effect.
And (having asked Opus 5.5 to check for me) it’s also true that ModMess2 can have multiple solutions that are not exactly the same as the target genome. As it explained:
… So if a target gene is divisible by the next target gene and is at least as large as the previous one, then any other multiple of the next gene that is also at least the previous one works too. For example, with targets …30, 40, 20… at three adjacent positions, the middle gene can be 40, 60, 80 or 100….
Why?
I was paging through some old idea books looking for GA problems where lexicase selection would be obviously better than fitness-proportionate or tournament selection. And, frankly, a quick hand-wave by Claude to check things seems to support that idea.
One try led me to do a bunch of runs on a lot of landscapes, and yup: ModMess2 is better this way.
One of the interesting things, to me, is the way ModMess2 makes things wobble around when you try to independently optimize one gene at a time (because of that second neighboring gene’s effects).
SuperFizzBuzz
And working through that jostled my memory, and brought back to mind (25 years or so later) where it came from and where it was going in my notes.
The origin story is “sitting and talking with Terry Jones”, probably in the early 2000s, though the topic was some stuff from his thesis back in the mid-1990s.
I don’t know if he ever wrote anything up about “Random Schema Landscapes”, but they’re his, as I recall. Something that didn’t make the cut, maybe? Inspired by working with John Holland so much, and as counterpoint to Stu Kauffman’s NK landscapes, I remember him pointing out how RS landscapes have neutral networks, where NK models are just bumpy rugged bullshit all over.
The premise for RS Landscapes, as I recall them years later?
- the search space is bitstrings of length \(N\)
- we specify an RS Landscape by building a key-value store of \(N\)-bit schemata (as keys) and random doubles (as values); in each schema, a fixed (tunable) number of random locations are specified bit values of
0or1, and the rest are wildcards# - we score a bitstring by adding up the doubles associated with every schema it matches
There’s a parameter in there somewhere, which might have been R for some reason, that’s how many bits are specified. In my head is a confused feeling that it’s S schemata and R bits each, but the name is also from “Random Schemata”.
Anyway, that’s the origin story for ModMess as well, and where ModMess2 comes in, and how we get all the way up to SuperFizzBuzz.
It’s not complicated!
ModMess2 is just ModMess which uses one target. and then re-uses that target rotated one step to the left. That is, we take the inputs modulo the genes at position \(i\), and then take that result modulo the genes at position \(i+1\). As Claude explained above, adjacent genes that are multiples will do some weird stuff.
SuperFizzBuzz has the same general setup: G genes that are integer values from a limited range (that’s important).
But the “target” is not identically matching a genome. It’s defined in terms of multiple targets, each a collection of G integers. Increasing the number of targets makes the problem harder.
The process is:
- We take the input vector of numbers, and construct a new vector of the same length, where each element is the product of an element and its left and right neighbors. (wrapping at the ends)
- For each defined target vector, we take these sums and collect a new vector of the remainders obtained modulo the genome’s value at each position (in parallel, not serially).
- We add those remainder vectors all together
- The result is the sum of those remainders, over all vectors and index
So instead of trying to match a target, now we are aggregating adjacent genes by multiplication, and hoping that those products will be multiples of several factors.
A little example of SuperFizzBuzz
Define the landscape over 10-integer vectors, using these targets:
[80, 99, 67, 84, 83, 69, 11, 43, 98, 82][10, 65, 64, 47, 31, 100, 40, 40, 25, 36]
Let’s score a random solution [65, 98, 72, 68, 68, 54, 32, 14, 71, 42].
The products in each position with its left and right neighbors are [267540, 458640, 479808, 332928, 249696, 117504, 24192, 31808, 41748, 193830].
The index-wise remainders are:
t1 -> [20, 72, 21, 36, 32, 66, 3, 31, 0, 64]t2 -> [0, 0, 0, 27, 22, 4, 32, 8, 23, 6]
The total of those values is 467. I think?
Does this stay interesting? I don’t know!
But it explains where ModMess2 came from, in my notes. One was a sketch of the other.
Or, alternately, it doesn’t matter. Undoubtedly I’ve poked around with RS landscapes a few times, but have nothing “real” to show for that. Just some play toys.