<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://eboshii-dev.web.app/feed.xml" rel="self" type="application/atom+xml" /><link href="https://eboshii-dev.web.app/" rel="alternate" type="text/html" /><updated>2026-09-26T08:02:37+00:00</updated><id>https://eboshii-dev.web.app/feed.xml</id><title type="html">eboshii.dev</title><subtitle>Thoughts on artificial life, finance, and travel.</subtitle><author><name>eboshii</name></author><entry><title type="html">Mooring a 4D boat</title><link href="https://eboshii-dev.web.app/blog/mooring-a-4d-boat/" rel="alternate" type="text/html" title="Mooring a 4D boat" /><published>2026-09-24T09:00:00+00:00</published><updated>2026-09-24T09:00:00+00:00</updated><id>https://eboshii-dev.web.app/blog/mooring-a-4d-boat</id><content type="html" xml:base="https://eboshii-dev.web.app/blog/mooring-a-4d-boat/"><![CDATA[<link rel="stylesheet" href="/assets/css/mooring-4d.css" />

<p>Ahoy there, sailor!</p>

<p>You’ve been away on a long voyage across the four-dimensional sea. Its heaving surface isn’t a sheet of water but a whole volume of it, slamming into your hull from directions no 3D sailor has a name for. The horizon all around you is a sphere rather than a circle. As well as fore and aft, port and starboard, and up and down, there’s a fourth axis, and your boat can steer along it. Charles Hinton, who spent the 1880s trying to teach people to picture four dimensions, named its two directions <em>ana</em> and <em>kata</em>, and they’ve served 4D sailors ever since.</p>

<p>Now the harbour wall is coming up off your ana bow, and it’s time to moor.</p>

<h2 id="1-the-bowline-slips-out-ana">1. The bowline slips out ana</h2>

<p>You throw a line around the bollard, tie a bowline and step ashore. Behind you, the knot slides open: one strand steps ana, slips past another, and the bowline falls apart.</p>

<p>A knot holds in three dimensions because rope can’t pass through rope. Draw a knot flat and at every crossing one strand goes over the other. To swap which strand is on top, the upper strand would have to drop past the lower one right where they cross, and there the two would meet. Solid rope can’t do that.</p>

<p>In four dimensions there’s a way around. Give every point of the sea four coordinates, \((x, y, z, w)\). The first three are the usual ones, and \(w\) measures how far a point is ana (\(w &gt; 0\)) or kata (\(w &lt; 0\)). Mathematicians call this space \(\mathbb{R}^4\).</p>

<p>Lift a short piece of the upper strand ana. It’s now at a different \(w\) from the lower strand, so the two can share the same \((x, y, z)\) without touching. Lower it past the lower strand, bring it back kata, and the crossing has changed without the strands ever meeting.</p>

<p>In the figures, colour is \(w\): blue is kata, cream is \(w = 0\) and orange is ana. Two pieces of rope touch only if they’re in the same place <em>and</em> the same colour.</p>

<figure class="m4-fig m4-panel">
<div id="m4-crossing"></div>
<figcaption>Changing one crossing of a trefoil, the simplest knot. After the change, the rope relaxes in ordinary 3D space and turns out to be a plain loop. The relaxation is a small simulation in which every part of the rope repels every other part, so it never passes through itself.</figcaption>
</figure>

<p>That’s enough to untie any knot. Pick a starting point on a knot diagram, away from any crossing, and walk once around the diagram. At every crossing you reach for the first time, make your strand the upper one. When you reach a crossing for the second time you’re on the lower strand, so leave it as it is.</p>

<p>The new diagram is always the unknot. To see why, give the rope heights that match it. Start at height 1 and let the height fall steadily as you walk, down to 0 as you arrive back at the start, then climb straight up to height 1 there. At every crossing the strand you reached first is higher, just as the diagram says. For each height \(h\) between 0 and 1, the rope now has exactly two points at that height, one on the way down and one on the climb. Join them with a horizontal segment. Segments at different heights lie in different horizontal planes, so none of them meet, and together they fill in a disc whose edge is the rope. Shrinking the rope across that disc turns it into a small circle, so it was the unknot all along.</p>

<p>So every knot turns into the unknot after some crossing changes, and in \(\mathbb{R}^4\) every crossing change is free.</p>

<p>That covers rope lying in ordinary space. A rope anywhere in \(\mathbb{R}^4\) can be moved there first, by shrinking every point’s \(w\) to 0. Two points of the rope would only collide on the way if they had the same \(x\), \(y\) and \(z\), and a slight nudge to the rope beforehand removes any such pairs. So in \(\mathbb{R}^4\), every knot comes undone.</p>

<p>The general idea is <em>codimension</em>, the dimension of the space minus the dimension of the object in it. A rope is 1D. In ordinary space it has codimension 2, which is enough room to go around another strand but not enough to get past it. In \(\mathbb{R}^4\) it has codimension 3, and the spare dimension lets every crossing undo itself.</p>

<h2 id="2-something-to-tie-to">2. Something to tie to</h2>

<p>No knots, then. But a rope doesn’t need a knot to hold. Pass it around a bollard and fuse the ends into a closed loop, and there’s nothing left to untie. You try it on the nearest bollard, and the loop slips off ana. The harbourmaster has watched plenty of visiting sailors do this, and points you down the quay to a different kind of bollard.</p>

<p>To see why the first one failed, start with an ordinary post in three dimensions, and for now assume it goes up forever. A real post has a top you could lift the loop over, and section 4 comes back to that.</p>

<p>Whether the post holds the loop comes down to a count. Stretch a surface across the loop, such as a soap film, and give it a front and a back. Follow the post upward, adding 1 each time it passes through the film from back to front and subtracting 1 each time it passes from front to back. The total is the <em>linking number</em> of the loop and the post. A loop dropped over the post has linking number \(\pm 1\). A loop lying on the quay beside the post has linking number 0.</p>

<p>Two facts make this count useful. First, it doesn’t depend on which film you choose. Two films with the same edge together form a closed surface, and a line running off to infinity at both ends leaves a closed surface as often as it enters. Second, it can’t change while the rope moves without touching the post. Carry the film along with the rope. Crossings in the film’s interior appear and disappear only in pairs of opposite sign, as the post slides over a fold in the film, so they leave the total unchanged. A single crossing can only escape across the film’s edge, and the edge is the rope itself. So no motion takes the loop from around the post to the quay beside it, and the post holds the rope.</p>

<p>The count needs the film and the post to cross at isolated points. In \(\mathbb{R}^n\), an \(a\)-dimensional object and a \(b\)-dimensional one in general position cross at isolated points when \(a + b = n\), and miss each other entirely when \(a + b &lt; n\). In 3D the film is 2D and the post is 1D, and \(2 + 1 = 3\).</p>

<p>In 4D the film is still 2D and a line is still 1D, but now \(2 + 1 &lt; 4\). Nudge the post ana and it misses the film altogether. The rope can then shrink across the film to a tiny loop without touching the post, and float away. To hold the rope, the bollard has to cross the film at isolated points, so it has to be 2D: \(2 + 2 = 4\). In \(\mathbb{R}^4\), a rope loop can only be held by a two-dimensional bollard.</p>

<p>The same count works in any dimension. A <em>\(p\)-sphere</em> is the \(p\)-dimensional version of a circle. A loop of rope is a 1-sphere, and a closed sheet with no holes, like a balloon, is a 2-sphere. A \(p\)-sphere bounds a \((p + 1)\)-dimensional film, so it can be <em>linked</em> with a \(q\)-dimensional object in \(\mathbb{R}^n\) when \((p + 1) + q = n\), or</p>

\[p + q = n - 1.\]

<p>A loop of string on a table links a point (\(1 + 0 = 2 - 1\)), a rope in ordinary space links a line (\(1 + 1 = 3 - 1\)), and a rope in \(\mathbb{R}^4\) links a plane (\(1 + 2 = 4 - 1\)).</p>

<p>Real bollards are solid, though. What counts is a bollard’s <em>core</em>, what’s left if you shrink it without it ever touching the rope. An ordinary post shrinks to the line up its middle. The rope never touches the post as it shrinks, so the linking number doesn’t change, and the line holds the rope just as the post did. So \(q\) is the dimension of the core.</p>

<figure class="m4-fig m4-wide m4-panel">
<div id="m4-shrink"></div>
<figcaption>Each bollard shrinks to its core without the rope coming off. The fade at the top of a post means it keeps going up. The 4D bollard is drawn as three 3D slices, at w = −1, 0 and 1. Its core is a line in every slice, and the lines together make a plane.</figcaption>
</figure>

<p>The bollard you tried was the obvious 4D version of a post. A 4D sea has three horizontal directions, \(x\), \(y\) and \(w\), and that bollard was round in all three and ran up in \(z\). Its exact shape doesn’t matter, though. Any bollard that’s only so wide in each of \(x\), \(y\) and \(w\) shrinks to a line, and a line can’t hold a rope in \(\mathbb{R}^4\). The harbourmaster’s bollard is different. It’s round only in \(x\) and \(y\), and it runs up in \(z\) and keeps going ana and kata in \(w\), forever in both, so its core is the \(zw\)-plane.</p>

<p>The easiest way to see the difference is to look at the harbour one slice at a time. The slice at a fixed \(w\) is an ordinary 3D space, and moving the slice ana and kata shows how things change along the fourth axis.</p>

<figure class="m4-fig m4-wide m4-panel">
<div id="m4-slip"></div>
<figcaption>Drag the slice through w, or let the rope try to escape.</figcaption>
</figure>

<p>The round bollard gets thinner as you move ana, the same way the slices of a ball shrink towards its edge, and then it stops. Beyond its edge the slices are just water. So the rope steps ana past the bollard’s edge, slides sideways, and comes back kata beside it. The long bollard is in every slice, so wherever the rope goes, the bollard is still inside it.</p>

<p>At the long bollard, the harbourmaster brings a coil of rope from the ropewalk, runs it around the bollard and around a cleat of the same long shape on your deck, and splices the ends together where they lie. The splice is necessary. The count that stops a closed loop coming off also stops one going on, so the only way to get a closed loop around the bollard is to close it in place. Your boat is moored without a single knot.</p>

<h2 id="3-missing-knots-take-a-tarpaulin">3. Missing knots? Take a tarpaulin</h2>

<p>The boat is safe, but every knot you know is now useless. You can’t lash a crate, hitch a fender or tie off a sail. The harbourmaster sees you turning a length of rope over in your hands and passes you a tarpaulin. “If it’s knots you want, give up on rope.”</p>

<p>Go back to the codimension count from section 1. A rope has too much room in \(\mathbb{R}^4\). A sheet has less. A tarpaulin is 2D, so in \(\mathbb{R}^4\) its codimension is 2, the same as rope in ordinary space. Closed sheets in \(\mathbb{R}^4\) can be knotted. From here on a closed sheet means a 2-sphere, as in section 2.</p>

<p>The move that undid the bowline doesn’t work on a sheet. Lifting one patch of sheet past another needs a spare dimension, and a sheet in \(\mathbb{R}^4\) has none. That doesn’t prove any sheet is knotted, since some other motion might still undo it. A proof needs an <em>invariant</em>, a quantity that stays the same however the sheet moves without passing through itself, and that differs between the sheet in question and a plain sphere.</p>

<p>The first knotted surface was built by Emil Artin in 1925, by <em>spinning</em>. It takes three steps to see how.</p>

<p><strong>Spinning in 3D.</strong> Hold a semicircle with both ends on a vertical axis and spin it about the axis. Each point of the arc travels around a circle, bigger the further it is from the axis, and the two ends stay where they are. The arc sweeps out a sphere.</p>

<p><strong>Turning in 4D.</strong> In 3D, a rotation turns about a line: points on the axis stay put and every other point moves in a circle. In 4D, a rotation turns about a whole plane. A rotation that mixes \(x\) and \(w\) leaves every point with \(x = w = 0\) where it is, which is the \(yz\)-plane, and moves every other point in a circle. A point that starts in ordinary space at distance \(d\) from the plane swings out ana, is at \(x = 0\) and \(w = d\) after a quarter turn, and arrives back in ordinary space on the far side of the plane after half a turn. The second half of the turn brings it back through kata.</p>

<p><strong>Spinning a knotted arc.</strong> Take a knotted arc in ordinary space with both ends on the plane \(x = 0\) and the rest of it on the side \(x &gt; 0\), and turn it through \(w\) about that plane. Just like the semicircle, each point sweeps out a circle and the ends stay put, so the arc sweeps out a sphere. The sphere contains a copy of the knotted arc at every angle of the turn.</p>

<p>Artin’s invariant was the <em>fundamental group</em> of the space around the sphere, the group of loops in that space up to continuous deformation. He showed that the spun sphere’s group is the same as the group of the space around the original knot in \(\mathbb{R}^3\). For the trefoil that group isn’t commutative. Around an unknotted sphere it’s the integers, so the spun trefoil can’t be unknotted. Rolfsen’s textbook <em>Knots and Links</em> (1976) covers spinning along with other knotted surfaces.</p>

<figure class="m4-fig m4-wide m4-panel">
<div id="m4-spin"></div>
<figcaption>Left: a semicircle spun about a line. Right: a knotted arc turned through w about the green plane. Seen from ordinary space, the turn looks like a squash. Faint copies show where the arc has been.</figcaption>
</figure>

<p>You can also look at the knotted sphere the way you looked at the bollards, one slice at a time.</p>

<figure class="m4-fig m4-wide m4-panel">
<div id="m4-slices"></div>
<figcaption>Slices of the spun trefoil at different values of w.</figcaption>
</figure>

<p>At \(w = 0\) the slice is the knotted arc joined end to end with its own mirror image. For the trefoil that’s a right-handed trefoil joined to a left-handed one. Knot theorists call it the <em>square knot</em>, a name borrowed from the reef knot (the square knot in America), whose drawing it resembles. Move ana or kata and the parts of the arc nearest the plane drop out of the slice. The knot breaks into separate loops, and they shrink and vanish at the sphere’s edge. The slices at \(w\) and \(-w\) have the same shape, because the turn is symmetric.</p>

<p>A tarpaulin also rescues the round bollard the rope slipped off. For a sheet, \(p = 2\), so the rule gives \(2 + q = 3\) and \(q = 1\): a closed sheet is held by a bollard whose core is a line.</p>

<p>It’s easiest to see at a single height. In the 3D slice at a fixed height \(z\), the round bollard is a solid ball. A rope loop around a ball just slides off it, but a sheet can wrap the ball completely. Shrink the ball to its core and you have the \(\mathbb{R}^3\) case of the rule, \(2 + 0 = 3 - 1\), a closed sheet held by a point. In ordinary space that’s a balloon with a speck of dust inside. Nobody would call it mooring, but it meets the definition used all along, since the two can’t be separated without one passing through the other. Here the point is one slice of the bollard’s core, which carries on up and down, so the sheet can’t slip off in \(z\) either.</p>

<p>So gather the tarpaulin around the bollard like a sack and fuse the neck shut, the way the harbourmaster spliced the rope. Don’t knot the neck. Gathered up, it’s a thin strand of sheet, and a thin strand tied like rope slips ana just as the bowline did.</p>

<h2 id="4-casting-off">4. Casting off</h2>

<p>Here is everything in one table. A \(p\)-sphere in \(\mathbb{R}^n\) is held by a core of dimension \(n - 1 - p\). In the table, knots hold only where the codimension is 2.</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>A rope loop is held by a core that is</th>
      <th>A closed sheet is held by a core that is</th>
      <th>Knots in rope loops</th>
      <th>Knots in closed sheets</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>\(\mathbb{R}^3\)</td>
      <td>a line</td>
      <td>a point</td>
      <td>hold</td>
      <td>don’t</td>
    </tr>
    <tr>
      <td>\(\mathbb{R}^4\)</td>
      <td>a plane</td>
      <td>a line</td>
      <td>don’t</td>
      <td>hold</td>
    </tr>
  </tbody>
</table>

<p>One entry hasn’t come up yet. A sheet in \(\mathbb{R}^3\) has codimension 1, and J. W. Alexander proved in 1924 that every smooth 2-sphere in \(\mathbb{R}^3\) bounds a solid ball, so it can be shrunk to a point and can’t be knotted. The table needs the restriction to spheres. A torus can be knotted in \(\mathbb{R}^3\), as the skin of a thickened trefoil.</p>

<p>Codimension 2 isn’t the only place knots can hold, though. In 1962 André Haefliger found smooth 3-spheres in \(\mathbb{R}^6\) that are knotted, in codimension 3.</p>

<p>Two assumptions did a lot of work. First, “held” meant held by topology alone. Real rope also depends on friction and stiffness, and a 4D rope would be a tube thick in three directions, whose mechanics nothing here describes. Second, every bollard went on forever. That’s the same fact as the splice. The count that stops a loop coming off also stops it going on, so the rope had to be spliced around the bollard and the sack closed around it. A real bollard has a top, and a flared top, the sideways pull of the boat and friction keep the rope on. A real 4D bollard would need the same, flaring out ana and kata as well as up, since the long bollard can’t go on forever in \(w\) either.</p>

<p>Knotted surfaces in four dimensions are still an active area of research.</p>

<p>Fair winds, sailor, and keep a tarpaulin aboard. Next time you’re in, ask the harbourmaster about chains.</p>

<h2 id="references">References</h2>

<ul>
  <li>Alexander, J. W. (1924). On the subdivision of 3-space by a polyhedron. <em>Proceedings of the National Academy of Sciences</em>, 10(1), 6–8. <a href="https://doi.org/10.1073/pnas.10.1.6">doi:10.1073/pnas.10.1.6</a></li>
  <li>Artin, E. (1925). Zur Isotopie zweidimensionaler Flächen im \(\mathbb{R}_4\). <em>Abhandlungen aus dem Mathematischen Seminar der Universität Hamburg</em>, 4, 174–177. <a href="https://doi.org/10.1007/BF02950724">doi:10.1007/BF02950724</a></li>
  <li>Haefliger, A. (1962). Knotted \((4k - 1)\)-spheres in \(6k\)-space. <em>Annals of Mathematics</em>, 75(3), 452–466. <a href="https://doi.org/10.2307/1970208">doi:10.2307/1970208</a></li>
  <li>Hinton, C. H. (1888). <em>A New Era of Thought</em>. Swan Sonnenschein.</li>
  <li>Rolfsen, D. (1976). <em>Knots and Links</em>. Publish or Perish.</li>
</ul>

<script src="/assets/js/mooring-4d.js"></script>]]></content><author><name>eboshii</name></author><summary type="html"><![CDATA[Knots don't hold in four dimensions. Here is what does.]]></summary></entry><entry><title type="html">A gradient test for ES-HyperNEAT’s quadtree</title><link href="https://eboshii-dev.web.app/blog/differentiable-es-hyperneat/" rel="alternate" type="text/html" title="A gradient test for ES-HyperNEAT’s quadtree" /><published>2026-09-23T10:00:00+00:00</published><updated>2026-09-23T10:00:00+00:00</updated><id>https://eboshii-dev.web.app/blog/differentiable-es-hyperneat</id><content type="html" xml:base="https://eboshii-dev.web.app/blog/differentiable-es-hyperneat/"><![CDATA[<link rel="stylesheet" href="/assets/css/es-hyperneat.css" />

<p><a href="https://doi.org/10.1162/artl_a_00071">ES-HyperNEAT</a> (Risi &amp; Stanley, 2012) decides where to put hidden neurons by evaluating its CPPN at the \(2^n\) sub-cells of every cell it looks at. That’s why its search slows down exponentially as the substrate gains dimensions. But to first order those \(2^n\) samples only measure one thing, the squared gradient of the weight field, and a gradient costs one forward and one backward pass whatever the dimension.</p>

<p>So I swapped the test. The gradient version made the same split decision as sampling 98.7% of the time, and found equally good networks in about a sixth of the search time on a 3D substrate. A smarter sampler closes some of that gap, and the tree itself still grows exponentially.</p>

<p class="eshn-status"><em>Epistemic status:</em> the equivalence is derived and checked numerically. The training comparison is 8 paired runs on one toy task, which is enough to rule out big differences in performance but not small ones.</p>

<p><a href="/blog/higher-dimensional-substrates/">Part 1</a> introduced ES-HyperNEAT and the steering task used here. The short version: HyperNEAT computes each weight from the coordinates of the two neurons it connects, \(w = f(\mathbf{p}, \mathbf{q})\), using a small network called the CPPN, and ES-HyperNEAT puts hidden neurons wherever \(f\) varies, finding those places with a quadtree.</p>

<p>Part 1 found that matching the substrate’s dimension to the problem helps, and that each extra dimension multiplies the cost of that quadtree search.</p>

<h2 id="1-to-first-order-the-2ⁿ-samples-measure-one-gradient">1. To first order, the 2ⁿ samples measure one gradient</h2>

<p>The quadtree splits a cell when the CPPN’s weights at its \(2^n\) sub-cell centres vary by more than a threshold \(\tau\). Variance across a small cell is a measure of how fast the weight field changes there, which is what a gradient tells you.</p>

<p>Take a cell centred at \(c\) with half-width \(r\). Its sub-cell centres are \(c + \tfrac{r}{2}\sigma\), where \(\sigma\) runs over every pattern of \(\pm 1\) signs. Near \(c\) the field is roughly linear:</p>

\[w\!\left(c + \tfrac{r}{2}\sigma\right) \approx w(c) + \tfrac{r}{2}\, \sigma \cdot \nabla w(c)\]

<p>Over the \(2^n\) sign patterns each \(\sigma_i\) is \(+1\) half the time and \(-1\) half the time, independently of the others, so \(\mathbb{E}[\sigma_i] = 0\), \(\mathbb{E}[\sigma_i^2] = 1\) and \(\mathbb{E}[\sigma_i \sigma_j] = 0\) for \(i \neq j\). The constant \(w(c)\) drops out of the variance, which leaves</p>

\[\operatorname{Var}_{\text{sub-cells}}(w) \;\approx\; \left(\tfrac{r}{2}\right)^{2} \sum_{i=1}^{n} \left(\frac{\partial w}{\partial x_i}\right)^{2} \;=\; \left(\tfrac{r}{2}\right)^{2} \lVert \nabla w(c) \rVert^{2}\]

<p>So the tree, the threshold and everything else can stay as they are. Only one line changes:</p>

\[\text{split if } \operatorname{Var}(2^n \text{ samples}) &gt; \tau \quad\longrightarrow\quad \text{split if } \left(\tfrac{r}{2}\right)^{2}\lVert \nabla w(c)\rVert^{2} &gt; \tau\]

<p>Numerically it holds up well. On random CPPNs the gradient’s estimate is within 5–15% of the sampled variance on the biggest cells, and indistinguishable from it on small ones.</p>

<!-- Diagram for the ES-HyperNEAT posts. The caption is passed in from the post. -->
<figure class="eshn-fig">
<svg class="eshn-diagram" viewBox="0 0 640 190" role="img" aria-label="Left: a cell sampled at four sub-cell centres. Right: the same cell with one point at the centre and a gradient arrow">
  <rect x="40" y="20" width="140" height="140" fill="none" stroke="rgba(235,225,210,0.45)" />
  <path d="M110 20 V160 M40 90 H180" stroke="rgba(235,225,210,0.18)" />
  <g fill="#3987e5" stroke="#0e0a18" stroke-width="2"><circle cx="75" cy="55" r="7" /><circle cx="145" cy="55" r="7" /><circle cx="75" cy="125" r="7" /><circle cx="145" cy="125" r="7" /></g>
  <text x="200" y="70">Sampling</text>
  <text x="200" y="96" class="m">2ⁿ CPPN passes</text>
  <rect x="360" y="20" width="140" height="140" fill="none" stroke="rgba(235,225,210,0.45)" />
  <circle cx="430" cy="90" r="7" fill="#d95926" stroke="#0e0a18" stroke-width="2" />
  <path d="M430 90 L478 56" stroke="#d95926" stroke-width="2.5" marker-end="url(#eshn-arr2)" />
  <defs><marker id="eshn-arr2" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="6" markerHeight="6" orient="auto"><path d="M0 0 L10 5 L0 10z" fill="#d95926" /></marker></defs>
  <text x="520" y="70">Gradient</text>
  <text x="520" y="96" class="m">1 forward + 1 backward pass</text>
</svg>
<figcaption>Same question, asked two ways. The backward pass is backpropagation to the CPPN's inputs.</figcaption>
</figure>

<p>The gradient is cheap because of reverse-mode automatic differentiation (backpropagation). It gets all \(n\) partial derivatives in one backward pass, for a small constant multiple of the forward cost, however big \(n\) is. The chain rule works back from the single output and reuses each intermediate value from the forward pass once, so a test costs one forward and one backward pass instead of \(2^n\) forward passes.</p>

<p>The tree itself doesn’t change, though. A cell the gradient marks as varying still splits into \(2^n\) new cells, so each test gets roughly \(2^n/2\) times cheaper while all \((2^n)^m\) cells are still there.</p>

<h2 id="2-the-two-tests-almost-always-agree">2. The two tests almost always agree</h2>

<p>The approximation breaks when the field is far from linear across a cell. A saddle centred exactly on a cell has zero gradient at the one point the test looks at, so the gradient test thinks the cell is flat and never splits it; the <em>Blind spot</em> button below builds one. Sampling has its own blind spots. A bump centred on a cell gives four identical samples.</p>

<div class="eshn-fig eshn-wide eshn-panel" id="eshn-quadtree"></div>

<p class="eshn-caption"><em>Compare</em> outlines the cells where the two tests disagree.</p>

<p>How often does this matter in practice? We rebuilt trees for random CPPNs in 2 to 7 dimensions and scored every cell with both tests. They agreed on 98.7% of 2,580 split decisions. In 32 of the 34 disagreements the gradient split a cell that sampling left alone, and only 2 went the other way.</p>

<p>That lean matches section 1: on big cells the gradient’s estimate runs a little high, which tips borderline cells over the threshold. An extra split costs a bit of compute. A missed split loses detail, and those were rare.</p>

<h2 id="3-the-saving-grows-with-dimension-but-a-better-sampler-narrows-it">3. The saving grows with dimension, but a better sampler narrows it</h2>

<p>Against sampling as published, the gradient test needs 2× fewer CPPN passes in 2D and 61× fewer in 7D, and the search time follows the same curve.</p>

<div class="eshn-fig eshn-wide eshn-panel">
<div id="eshn-chart-scaling"></div>
<p class="eshn-caption">One depth-2 tree, median over 10 random CPPNs with 4 inputs and 4 outputs. A backward pass counts as one pass; the corner-cache sampler is counted, not run.</p>
</div>

<table>
  <thead>
    <tr>
      <th style="text-align: right">n</th>
      <th style="text-align: right">sampling</th>
      <th style="text-align: right">corner cache</th>
      <th style="text-align: right">gradient</th>
      <th style="text-align: right">time: sampling</th>
      <th style="text-align: right">time: gradient</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: right">2</td>
      <td style="text-align: right">416</td>
      <td style="text-align: right">144</td>
      <td style="text-align: right">208</td>
      <td style="text-align: right">1.1 ms</td>
      <td style="text-align: right">0.75 ms</td>
    </tr>
    <tr>
      <td style="text-align: right">3</td>
      <td style="text-align: right">3,136</td>
      <td style="text-align: right">784</td>
      <td style="text-align: right">784</td>
      <td style="text-align: right">3.9 ms</td>
      <td style="text-align: right">1.2 ms</td>
    </tr>
    <tr>
      <td style="text-align: right">4</td>
      <td style="text-align: right">24,704</td>
      <td style="text-align: right">4,096</td>
      <td style="text-align: right">3,344</td>
      <td style="text-align: right">27 ms</td>
      <td style="text-align: right">2.9 ms</td>
    </tr>
    <tr>
      <td style="text-align: right">5</td>
      <td style="text-align: right">24,832</td>
      <td style="text-align: right">4,540</td>
      <td style="text-align: right">2,576</td>
      <td style="text-align: right">27 ms</td>
      <td style="text-align: right">2.4 ms</td>
    </tr>
    <tr>
      <td style="text-align: right">6</td>
      <td style="text-align: right">786,944</td>
      <td style="text-align: right">67,112</td>
      <td style="text-align: right">27,152</td>
      <td style="text-align: right">873 ms</td>
      <td style="text-align: right">24 ms</td>
    </tr>
    <tr>
      <td style="text-align: right">7</td>
      <td style="text-align: right">6,358,016</td>
      <td style="text-align: right">275,476</td>
      <td style="text-align: right">103,440</td>
      <td style="text-align: right">8.8 s</td>
      <td style="text-align: right">0.10 s</td>
    </tr>
  </tbody>
</table>

<p>It isn’t quite a fair fight, though. The published algorithm samples sub-cell centres, and none of those samples ever get reused. A variant that samples cell corners instead, and caches them, can reuse each corner in every cell that touches it. On the same trees that makes sampling 3× cheaper in 2D and 16× cheaper in 7D.</p>

<p>Against that version the gradient test is actually more expensive in 2D, level in 3D, and only 2.7× cheaper in 7D. Both still climb steeply, because both still build a \(2^n\)-way tree.</p>

<h2 id="4-it-finds-equally-good-networks-with-less-search">4. It finds equally good networks with less search</h2>

<p>Last, back to part 1’s steering task, on both the 3D substrate and the 2D map projection. The two versions share every line of code except the split test: same tree, thresholds, network, task, evolutionary algorithm and settings. They also share random seeds, so each gradient run starts from the same CPPN and gets the same random perturbations as its sampling twin. The two only drift apart once the tests first disagree about a split.</p>

<div class="eshn-fig eshn-wide eshn-panel">
<div class="eshn-row">
<div class="eshn-col"><div class="eshn-cap">3D substrate</div><div id="eshn-chart-method-3d"></div></div>
<div class="eshn-col"><div class="eshn-cap">2D substrate (map projection)</div><div id="eshn-chart-method-2d"></div></div>
</div>
<div id="eshn-strip-method" style="margin-top:1.2rem"></div>
<p class="eshn-caption">Top: mean ±1 standard error over 8 paired runs. Bottom: each seed; a dot inside its ring means both tests did equally well.</p>
</div>

<table>
  <thead>
    <tr>
      <th style="text-align: left"> </th>
      <th style="text-align: right">3D sampling</th>
      <th style="text-align: right">3D gradient</th>
      <th style="text-align: right">2D sampling</th>
      <th style="text-align: right">2D gradient</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">mean final score (8 runs)</td>
      <td style="text-align: right">0.737</td>
      <td style="text-align: right">0.728</td>
      <td style="text-align: right">0.551</td>
      <td style="text-align: right">0.570</td>
    </tr>
    <tr>
      <td style="text-align: left">runs reaching 0.68</td>
      <td style="text-align: right">8 / 8</td>
      <td style="text-align: right">7 / 8</td>
      <td style="text-align: right">4 / 8</td>
      <td style="text-align: right">4 / 8</td>
    </tr>
    <tr>
      <td style="text-align: left">neuron-placement time per genome</td>
      <td style="text-align: right">79 ms</td>
      <td style="text-align: right">14 ms</td>
      <td style="text-align: right">11.5 ms</td>
      <td style="text-align: right">5.0 ms</td>
    </tr>
    <tr>
      <td style="text-align: left">one full training run</td>
      <td style="text-align: right">5.9 min</td>
      <td style="text-align: right">1.7 min</td>
      <td style="text-align: right">1.2 min</td>
      <td style="text-align: right">0.7 min</td>
    </tr>
  </tbody>
</table>

<p class="eshn-caption">Timings on an AMD Ryzen 7 7840HS, 15 runs in parallel. A genome is one candidate CPPN.</p>

<p>Performance came out the same. In 3D the mean final scores differ by less than 0.01, slightly in sampling’s favour (95% confidence interval for the difference: −0.027 to +0.002), and each version won four of the eight pairs. Most of that gap is one gradient run that was still improving when training stopped.</p>

<p>In 2D the gradient version is slightly ahead on average (95% confidence interval: −0.042 to +0.089), but it only won three of the eight pairs, with one tie. Its lead comes from two big wins. The 2D runs that succeeded were the same four seeds under both tests, so a run’s outcome is set by its starting CPPN and the substrate, not by the split test.</p>

<div class="eshn-fig eshn-wide eshn-panel eshn-swimmers" data-runs="3d/sample/*,3d/grad/*,2d-azim/sample/*,2d-azim/grad/*" data-labels="3D, sampling test|3D, gradient test|2D map projection, sampling test|2D map projection, gradient test"></div>

<p class="eshn-caption">All 8 runs of each version, one per chase. Sampling and gradient panels always show the same seed.</p>

<p>Where they differ is time. Placing neurons took about a sixth as long in 3D and under half as long in 2D, so a full 3D training run finished 3.5× sooner. Going from 2D to 3D multiplied sampling’s placement cost by 7 and the gradient’s by 3, because sampling pays for both the bigger tree and the doubled cost of each test, while the gradient only pays for the bigger tree.</p>

<h2 id="5-caveats">5. Caveats</h2>

<p>This is a simplified ES-HyperNEAT written from scratch, with a fixed-shape CPPN trained by an evolution strategy instead of NEAT. Both versions share all of it, so the comparison is fair, but the absolute numbers would come out differently in the reference implementation.</p>

<p>It’s one toy task with 8 paired runs per substrate, which rules out big differences but not small ones. Part 1 describes the fixes that were needed before anything learned at all. Code and raw results are in <a href="/experiments/es-hyperneat/">/experiments/es-hyperneat/</a>, and the two versions differ only in the function <code class="language-plaintext highlighter-rouge">complexity()</code> in <code class="language-plaintext highlighter-rouge">eshn.py</code>.</p>

<h2 id="6-what-a-5-dimensional-creature-looks-like">6. What a 5-dimensional creature looks like</h2>

<p>Swapping the quadtree’s \(2^n\) samples for one gradient leaves the split decisions, and the networks that come out of them, almost unchanged, and it takes the exponential cost out of each test. The tree still branches \(2^n\) ways, so higher dimensions aren’t free. But 5, 6 and 7-dimensional substrates are now cheap enough to try, which raises the question of what a 5-dimensional task actually looks like.</p>

<p>In a life simulation, more of them than you might expect. A cell that senses which way the food is lives in 3D, but one that also cares which way it’s facing has three more axes, for six. A member of a swarm reacting to its neighbours’ positions and velocities lives in the same six. A creature that grows adds time to its 3D body, and a cellular automaton whose cells carry several chemical signals gets an axis for each one on top of its grid.</p>

<p>Outside artificial life it’s the same story. A drone’s controls depend on both its position and its orientation. A robot arm’s state lives in the six or seven dimensions of its joint angles. Medical scans are 3D volumes that change over time, and so is the weather. Whenever a task’s inputs and outputs come with a natural geometry of more than three dimensions, a substrate that matches it might make the right network simple, the way the 3D substrate did for the swimming cell.</p>

<p>Finding out no longer gets exponentially more expensive with every dimension, and I’d love to see what turns up.</p>

<script src="/assets/js/es-hyperneat-post.js"></script>]]></content><author><name>eboshii</name></author><summary type="html"><![CDATA[Part 2 of 2. Replacing the quadtree's 2ⁿ samples with one gradient finds equally good networks for a fraction of the search.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://eboshii-dev.web.app/assets/images/teasers/differentiable-es-hyperneat.jpg" /><media:content medium="image" url="https://eboshii-dev.web.app/assets/images/teasers/differentiable-es-hyperneat.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">3D creatures need 3D brains</title><link href="https://eboshii-dev.web.app/blog/higher-dimensional-substrates/" rel="alternate" type="text/html" title="3D creatures need 3D brains" /><published>2026-09-23T09:00:00+00:00</published><updated>2026-09-23T09:00:00+00:00</updated><id>https://eboshii-dev.web.app/blog/higher-dimensional-substrates</id><content type="html" xml:base="https://eboshii-dev.web.app/blog/higher-dimensional-substrates/"><![CDATA[<link rel="stylesheet" href="/assets/css/es-hyperneat.css" />

<p>Artificial life simulations are full of small creatures that need brains: cells chasing food, swimmers, walkers, whole ecosystems of them. Those brains are usually evolved, and one of the nicest ways to do it came out of the field itself: ES-HyperNEAT (<a href="https://doi.org/10.1162/artl_a_00071">Risi &amp; Stanley, 2012</a>).</p>

<p>ES-HyperNEAT doesn’t evolve a network’s weights one by one. It evolves a small function that draws them from geometry. Each neuron gets coordinates, and the weight between two neurons is that function evaluated at their two positions, so moving a neuron changes all of its weights, and with them what it computes. That lets a creature’s brain mirror the layout of its body, and the same function decides where the hidden neurons go.</p>

<p>The same family of encodings has since evolved cellular automata that grow and copy patterns (<a href="https://doi.org/10.1109/TCDS.2017.2737082">Nichele et al., 2017</a>), and whole ecosystems of neural cellular automata (<a href="https://arxiv.org/abs/2406.09654">Barbieux &amp; Canaan, 2024</a>).</p>

<p>ES-HyperNEAT’s brains are usually flat, though. Its search is a quadtree, built for a 2D sheet, while plenty of simulated creatures live in 3D. What I wanted to know is whether that matters. The test is a single cell chasing a drifting bit of food in 3D.</p>

<p>With the neurons laid out in 3D, the one-line rule “connect neurons that are close” steers about as well as a hand-built controller. Flatten the same neurons onto a plane and the rule falls apart in every layout I tried. In one of them the up and down thrusters end up on the same spot and cancel each other out, so the cell can never change height.</p>

<p>Evolving from scratch tells a similar story. ES-HyperNEAT found a good controller in every 3D run, and in about half the runs on the best 2D layout.</p>

<div class="eshn-fig eshn-wide eshn-panel eshn-swimmers" data-runs="3d/sample/*,2d-azim/sample/*,2d-flat/sample/*" data-labels="3D (all 8 runs)|2D map projection (all 8 runs)|2D, drop z (all 4 runs)"></div>

<p class="eshn-caption">Every evolved run, one per chase. 3D always catches the food, the 2D map about half the time, and drop z can't change height.</p>

<p>The problem is cost. The search that places hidden neurons gets exponentially more expensive with every dimension you add, and <a href="/blog/differentiable-es-hyperneat/">Part 2</a> is about making it cheaper.</p>

<p class="eshn-status"><em>Epistemic status:</em> one toy task, 8 seeds per setup, and a simplified ES-HyperNEAT written from scratch rather than the reference code. The hand-set rule result is big and clear. The evolution result is suggestive, with one open confound: the 3D networks grew about six times as many hidden neurons.</p>

<h2 id="1-hyperneat-computes-weights-from-positions">1. HyperNEAT computes weights from positions</h2>

<p><a href="https://en.wikipedia.org/wiki/HyperNEAT">HyperNEAT</a> puts every neuron at a point in a geometric space, called the <em>substrate</em>, and computes each weight from the positions of the two neurons it connects:</p>

\[w = f(\mathbf{p}, \mathbf{q})\]

<p>The function \(f\) is itself a small network, the <em>CPPN</em> (compositional pattern-producing network). Evolution never touches the weights directly. It evolves \(f\), originally with NEAT.</p>

<!-- Diagram for the ES-HyperNEAT posts. The caption is passed in from the post. -->
<figure class="eshn-fig">
<svg class="eshn-diagram" viewBox="0 0 640 210" role="img" aria-label="Two neurons on a square substrate; their coordinates go into the CPPN, which returns the weight of the connection between them">
  <rect x="20" y="15" width="180" height="180" rx="6" fill="none" stroke="rgba(235,225,210,0.32)" />
  <text x="110" y="208" text-anchor="middle" class="m">substrate</text>
  <circle cx="65" cy="150" r="7" fill="#3987e5" stroke="#0e0a18" stroke-width="2" />
  <text x="48" y="178" class="m">(x₀, y₀)</text>
  <circle cx="155" cy="60" r="7" fill="#d95926" stroke="#0e0a18" stroke-width="2" />
  <text x="128" y="40" class="m">(x₁, y₁)</text>
  <path d="M71 144 L147 68" stroke="#F5F0E8" stroke-width="2" marker-end="url(#eshn-arr)" />
  <text x="118" y="118">w = ?</text>
  <defs><marker id="eshn-arr" viewBox="0 0 10 10" refX="8" refY="5" markerWidth="7" markerHeight="7" orient="auto"><path d="M0 0 L10 5 L0 10z" fill="#F5F0E8" /></marker></defs>
  <path d="M225 105 L300 105" stroke="rgba(235,225,210,0.5)" stroke-width="1.5" stroke-dasharray="4 4" marker-end="url(#eshn-arr)" />
  <text x="262" y="95" text-anchor="middle" class="m">ask the CPPN</text>
  <text x="318" y="62" class="m">x₀</text><text x="318" y="92" class="m">y₀</text><text x="318" y="122" class="m">x₁</text><text x="318" y="152" class="m">y₁</text>
  <path d="M338 58 L380 80 M338 88 L380 92 M338 118 L380 118 M338 148 L380 130" stroke="rgba(235,225,210,0.4)" />
  <rect x="380" y="55" width="150" height="100" rx="10" fill="rgba(28,18,48,0.75)" stroke="rgba(235,225,210,0.32)" />
  <text x="455" y="78" text-anchor="middle">CPPN</text>
  <path d="M396 118 q8 -22 16 0 t16 0" fill="none" stroke="#3987e5" stroke-width="2" />
  <path d="M440 128 q10 -40 20 0" fill="none" stroke="#199e70" stroke-width="2" />
  <path d="M472 130 q8 0 14 -12 t14 -12" fill="none" stroke="#c98500" stroke-width="2" />
  <text x="455" y="148" text-anchor="middle" class="m">sin · gaussian · tanh</text>
  <path d="M530 105 L585 105" stroke="#F5F0E8" stroke-width="2" marker-end="url(#eshn-arr)" />
  <text x="598" y="110">w</text>
</svg>
<figcaption>The CPPN turns two neurons' coordinates into the weight between them.</figcaption>
</figure>

<p>Because \(f\) is smooth, neurons that sit near each other get similar weights. And since \(f\) only ever sees coordinates, a rule like “connect neurons that are close” needs nothing more than the distance \(\lVert \mathbf{p} - \mathbf{q} \rVert\). (HyperNEAT is usually sold on symmetry and repetition. This post leans on something plainer, locality.)</p>

<p>So where do the coordinates come from? The designer places the input and output neurons to match the physical layout. A sensor that looks left goes on the left of the substrate, and a thruster that pushes up goes near the top. That’s the only place the physical world gets into the network.</p>

<p>Hidden neurons don’t correspond to anything physical. What matters is what they’re near. A hidden neuron’s weights are the CPPN evaluated between its position and each input and output, so under “connect what’s close”, one sitting between the left sensor and the left thruster mostly listens to the first and drives the second. Its position is its job.</p>

<p><a href="https://doi.org/10.1162/artl_a_00071">ES-HyperNEAT</a> (Risi &amp; Stanley, 2012) goes a step further and lets \(f\) place the hidden neurons too. Only the inputs and outputs are fixed. Pin one end of a connection to an input neuron, say at \((0, -1)\), and \(f(0, -1, x, y)\) becomes a scalar field over the substrate: the weight a hidden neuron at \((x, y)\) would get from that input.</p>

<p>Hidden neurons go where this field varies. Where it’s flat, neighbouring neurons would get identical weights and compute the same thing, so extra ones would be wasted.</p>

<h2 id="2-a-quadtree-finds-where-the-field-varies">2. A quadtree finds where the field varies</h2>

<p>To find those regions, ES-HyperNEAT keeps splitting the square into quarters:</p>

<ol>
  <li>Evaluate the CPPN at the centres of a cell’s four quarters and take the variance of the four weights, starting with the whole square.</li>
  <li>If the variance is above a threshold \(\tau\), split the cell and repeat on each quarter, down to some maximum depth.</li>
  <li>Every leaf cell whose variance is still above a lower threshold gets a hidden neuron at its centre.</li>
</ol>

<p>The explorer below runs this on the field of a random CPPN, \(w(x, y) = f(0, -1, x, y)\). Colour is the weight, lines are the leaf cells and dots are hidden neurons. Drag \(\tau\) down and the tree digs into the busy regions.</p>

<div class="eshn-fig eshn-wide eshn-panel" id="eshn-quadtree" data-mode="sample"></div>

<h2 id="3-in-3d-the-right-controller-is-a-one-line-rule">3. In 3D, the right controller is a one-line rule</h2>

<p>Since the CPPN only sees coordinates, a substrate works well when its coordinates capture the geometry the task cares about. The weights you need are then a simple function of position, and simple functions are what evolution tends to find first, because it builds up a CPPN one mutation at a time.</p>

<p>The test task is deliberately 3D. A point agent (the cell above) chases a target (the food) along a random 3D curve. It has 14 sensors pointing in fixed directions \(\mathbf{u}_k\), the 6 faces and 8 corners of a cube, and sensor \(k\) reads \(s_k = \max(0, \mathbf{u}_k \cdot \hat{\mathbf{r}})\), where \(\hat{\mathbf{r}}\) points at the target. It steers with 6 thrusters, along \(\pm x\), \(\pm y\) and \(\pm z\), through a network with one hidden layer.</p>

<p>A hand-built controller that just thrusts towards the target scores 0.730. Zero thrust scores 0.225.</p>

<p>That controller is purely geometric. If a sensor fires, the food is roughly in that sensor’s direction, and the thrusters pointing that way will push the cell towards it. So thruster \(j\) should respond to sensor \(k\) in proportion to how well their directions line up, \(\mathbf{u}_k \cdot \mathbf{a}_j\). Put every sensor and thruster on the substrate at its real direction (sensors at radius 1, thrusters at radius 0.5) and “lined up” turns into “close together”, because at fixed radii</p>

\[\lVert \mathbf{p} - \mathbf{q} \rVert^2 = \lVert \mathbf{p} \rVert^2 + \lVert \mathbf{q} \rVert^2 - 2\, \mathbf{p} \cdot \mathbf{q}\]

<p>shrinks as \(\mathbf{p} \cdot \mathbf{q}\) grows. So the locality rule</p>

\[w = b - k \lVert \mathbf{p} - \mathbf{q} \rVert\]

<p>is the controller, routed through a grid of hidden neurons. Short connections are positive and long ones negative. A sensor excites the hidden neurons near it, and they excite the thrusters near them, which point the same way as the sensor.</p>

<p>A flat substrate can’t do this properly. The <a href="https://en.wikipedia.org/wiki/Borsuk%E2%80%93Ulam_theorem">Borsuk–Ulam theorem</a> says any continuous flattening of a sphere sends some pair of opposite directions to the same point. With only 14 sensors you can dodge the collision by tilting the projection, but the map still squashes some directions together and pulls others apart. The ring layout avoids collisions by giving up continuity altogether, so neighbours on the ring point in unrelated directions.</p>

<!-- Diagram for the ES-HyperNEAT posts. The caption is passed in from the post. -->
<figure class="eshn-fig">
<svg class="eshn-diagram" viewBox="0 0 640 170" role="img" aria-label="A sphere with a sensor at the top and one at the bottom; after dropping the z coordinate both land at the centre of the flat map">
  <circle cx="110" cy="85" r="62" fill="none" stroke="rgba(235,225,210,0.32)" />
  <ellipse cx="110" cy="85" rx="62" ry="18" fill="none" stroke="rgba(235,225,210,0.18)" />
  <circle cx="110" cy="23" r="7" fill="#3987e5" stroke="#0e0a18" stroke-width="2" /><text x="122" y="20">+z  "target above"</text>
  <circle cx="110" cy="147" r="7" fill="#3987e5" stroke="#0e0a18" stroke-width="2" /><text x="122" y="156">−z  "target below"</text>
  <path d="M250 85 L330 85" stroke="rgba(235,225,210,0.5)" stroke-width="1.5" stroke-dasharray="4 4" marker-end="url(#eshn-arr)" />
  <text x="290" y="75" text-anchor="middle" class="m">flatten (drop z)</text>
  <rect x="360" y="25" width="120" height="120" rx="4" fill="none" stroke="rgba(235,225,210,0.32)" />
  <circle cx="420" cy="85" r="7" fill="#3987e5" stroke="#0e0a18" stroke-width="2" />
  <circle cx="424" cy="81" r="7" fill="#3987e5" stroke="#0e0a18" stroke-width="2" />
  <text x="494" y="89">same point</text>
</svg>
<figcaption>Projecting along z puts +z and −z on the same spot.</figcaption>
</figure>

<p>Try it below. Pick a substrate and play with \(k\) and \(b\), or hit the search button. The hidden neurons sit on a fixed grid here; in section 4 evolution gets to place them.</p>

<div class="eshn-fig eshn-wide eshn-panel" id="eshn-flatten"></div>

<p class="eshn-caption">Live scores on 16 random chases. The table has the best settings from a grid search.</p>

<table>
  <thead>
    <tr>
      <th style="text-align: left">Substrate</th>
      <th style="text-align: right">Best score with \(w = b - k\lVert \mathbf{p} - \mathbf{q} \rVert\)</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td style="text-align: left">3D (true directions)</td>
      <td style="text-align: right"><strong>0.745</strong></td>
    </tr>
    <tr>
      <td style="text-align: left">2D, drop z</td>
      <td style="text-align: right">0.317</td>
    </tr>
    <tr>
      <td style="text-align: left">2D, map projection (azimuthal equidistant)</td>
      <td style="text-align: right">0.306</td>
    </tr>
    <tr>
      <td style="text-align: left">2D, sensors round a ring</td>
      <td style="text-align: right">0.231</td>
    </tr>
    <tr>
      <td style="text-align: left"><em>reference: zero thrust / hand-built controller</em></td>
      <td style="text-align: right"><em>0.225 / 0.730</em></td>
    </tr>
  </tbody>
</table>

<p>In 3D the locality rule does as well as the hand-built controller. In every 2D layout it lands much closer to zero thrust. A cleverer CPPN could encode exceptions to locality, but evolution would have to stumble on each one.</p>

<h2 id="4-evolution-mostly-agrees-with-one-telling-failure">4. Evolution mostly agrees, with one telling failure</h2>

<p>The hand-set rule used a fixed grid, though. The real test is letting ES-HyperNEAT do the whole job: evolve the CPPN from a random start and place the hidden neurons with its own quadtree.</p>

<details class="eshn-details">
  <summary>How the experiment works: dynamics, learning algorithm, runs and code</summary>

  <ul>
    <li><strong>Task.</strong> The agent integrates its net thrust \(\mathbf{F}\) with damping, \(\mathbf{v} \leftarrow 0.8\,\mathbf{v} + 0.3\,\mathbf{F}\) and \(\mathbf{x} \leftarrow \mathbf{x} + 0.25\,\mathbf{v}\), so it overshoots if it thrusts too hard. A chase lasts 60 steps and scores \(1/(1 + \bar{d})\), where \(\bar{d}\) is the mean distance to the target.</li>
    <li><strong>Learning.</strong> Each run starts from a randomly initialised CPPN and trains it with an <em>evolution strategy</em>. Every generation it evaluates 32 random perturbations of the CPPN’s parameters (16 antithetic \(\pm\) pairs), estimates the gradient of the score from their rank-weighted results, and takes an Adam step. The score is treated as a black box: no gradient passes through the simulation. A run lasts 120 generations.</li>
    <li><strong>Scoring.</strong> Every 5 generations the current CPPN is tested on 16 held-out chases that are never used for training. That test score is what the charts show.</li>
    <li><strong>Runs.</strong> A <em>seed</em> fixes a run’s random initialisation and random choices. There are 8 seeds for the 3D substrate and the 2D map projection, and 4 for each of the two cruder 2D layouts. Everything except the substrate is identical.</li>
    <li><strong>Code.</strong> About 800 lines of NumPy, in <a href="/experiments/es-hyperneat/">/experiments/es-hyperneat/</a>. <code class="language-plaintext highlighter-rouge">python run_all.py</code> reruns everything, and <code class="language-plaintext highlighter-rouge">results/</code> holds each run’s raw output, including its final CPPN.</li>
  </ul>

</details>

<div class="eshn-fig eshn-wide eshn-panel">
<div id="eshn-chart-substrate"></div>
<div id="eshn-strip-substrate" style="margin-top:1.2rem"></div>
<p class="eshn-caption">Top: mean test score, ±1 standard error. Bottom: every run's final score.</p>
</div>

<p>Every 3D run ended up level with the hand-built controller.</p>

<p>The 2D map projection split down the middle. Four of its eight runs got there too, just later, and the other four stalled somewhere between 0.27 and 0.54. The ring layout never got above 0.35.</p>

<p>The “drop z” layout is my favourite failure. All four runs finished at exactly 0.319. Dropping \(z\) sends the \(\pm z\) directions to the same point, which is exactly where a sensor pair and a thruster pair live. The two vertical thrusters always get identical weights, fire together and cancel, so the cell can’t change height whatever the CPPN does.</p>

<p>There’s one confound I can’t rule out. The 3D networks grew about 480 hidden neurons to 2D’s 75. That isn’t because the 2D tree ran out of room, but it does mean this experiment can’t separate “more dimensions” from “more neurons”.</p>

<h2 id="5-each-dimension-multiplies-the-search">5. Each dimension multiplies the search</h2>

<p>In \(n\) dimensions the quadtree becomes a \(2^n\)-tree. A cell has 4 sub-cells in 2D, 8 in 3D and 64 in 6D, and both parts of the cost scale with that:</p>

<ul>
  <li>every test evaluates the CPPN at \(2^n\) sub-cell centres;</li>
  <li>every split makes \(2^n\) new cells, each needing its own test.</li>
</ul>

<p>If the field varies everywhere down to depth \(m\), that’s roughly \((2^n)^m\) cells at \(2^n\) evaluations each, for every input neuron.</p>

<div class="eshn-fig eshn-panel" id="eshn-cost" data-mode="sample"></div>

<p>Here’s what that looks like on random CPPNs in 2 to 7 dimensions, with trees of depth 2:</p>

<div class="eshn-fig eshn-wide eshn-panel">
<div id="eshn-chart-curse"></div>
<p class="eshn-caption">Median over 10 random CPPNs, log scale. Tree sizes vary, which is why 5D is no dearer than 4D.</p>
</div>

<details class="eshn-details">
  <summary>Show the numbers</summary>

  <table>
    <thead>
      <tr>
        <th style="text-align: right">n</th>
        <th style="text-align: right">cells tested</th>
        <th style="text-align: right">CPPN evaluations</th>
        <th style="text-align: right">time per search</th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td style="text-align: right">2</td>
        <td style="text-align: right">13</td>
        <td style="text-align: right">416</td>
        <td style="text-align: right">1.1 ms</td>
      </tr>
      <tr>
        <td style="text-align: right">3</td>
        <td style="text-align: right">49</td>
        <td style="text-align: right">3,136</td>
        <td style="text-align: right">3.9 ms</td>
      </tr>
      <tr>
        <td style="text-align: right">4</td>
        <td style="text-align: right">193</td>
        <td style="text-align: right">24,704</td>
        <td style="text-align: right">27 ms</td>
      </tr>
      <tr>
        <td style="text-align: right">5</td>
        <td style="text-align: right">97</td>
        <td style="text-align: right">24,832</td>
        <td style="text-align: right">27 ms</td>
      </tr>
      <tr>
        <td style="text-align: right">6</td>
        <td style="text-align: right">1,537</td>
        <td style="text-align: right">786,944</td>
        <td style="text-align: right">0.87 s</td>
      </tr>
      <tr>
        <td style="text-align: right">7</td>
        <td style="text-align: right">6,209</td>
        <td style="text-align: right">6,358,016</td>
        <td style="text-align: right">8.8 s</td>
      </tr>
    </tbody>
  </table>

  <p>Timings from an AMD Ryzen 7 7840HS.</p>

</details>

<p>Going from 2D to 7D makes the work about 15,000 times bigger, and that’s one search on one network. Evolution runs a search for every candidate CPPN in every generation. Even in the steering experiment, moving from 2D to 3D stretched a 1.2-minute training run to 5.9 minutes.</p>

<h2 id="6-caveats">6. Caveats</h2>

<ul>
  <li><strong>The prediction was only half right.</strong> The hand-set rule made 2D look hopeless. Evolution then solved the 2D map projection half the time, so a mismatched substrate makes good solutions harder and less reliable to find. It doesn’t rule them out.</li>
  <li><strong>Nothing learned at first.</strong> The first version never beat zero thrust. Random CPPNs saturate every neuron, and opposing thrusters cancel. Fan-in scaling and a zero-mean output shift, applied to every run before any comparison, fixed it. A hidden-neuron cap that was quietly bunching neurons into one corner came out as well.</li>
</ul>

<details class="eshn-details">
  <summary>Simplifications and scope</summary>

  <ul>
    <li><strong>Simplifications.</strong> This ES-HyperNEAT was written from scratch, not taken from the reference implementation. The CPPN has a fixed shape and is trained with an evolution strategy instead of NEAT. There’s one tree per network instead of one per input neuron, a single hidden layer, and a simplified version of the published pruning step. All of that is the same across substrates, but the absolute numbers would come out differently in the reference code.</li>
    <li><strong>Narrow evidence.</strong> It’s one toy task, picked because its geometry makes the point, with 8 runs per substrate, so only fairly large differences show up. It also only compares 3D with 2D.</li>
  </ul>

</details>

<h2 id="7-the-dilemma">7. The dilemma</h2>

<p>Matching the problem’s geometry argues for as many dimensions as the problem has, and most physical problems, from robot arms to drones, are 3D. Every dimension you add multiplies the cost of placing neurons.</p>

<p>There may be a way round it. The quadtree’s variance test is really asking how fast the weight field changes across a cell, which is a question about its gradient, and the CPPN is built from differentiable functions. <a href="/blog/differentiable-es-hyperneat/">Part 2: A gradient test for ES-HyperNEAT’s quadtree</a> swaps the \(2^n\) samples for one gradient and checks whether it finds equally good networks for less work.</p>

<h2 id="references">References</h2>

<ul>
  <li>Barbieux, A. &amp; Canaan, R. (2024). Coralai: Intrinsic evolution of embodied neural cellular automata ecosystems. <em>Proceedings of the Artificial Life Conference (ALIFE 2024)</em>. <a href="https://arxiv.org/abs/2406.09654">arXiv:2406.09654</a></li>
  <li>Nichele, S., Ose, M. B., Risi, S. &amp; Tufte, G. (2017). CA-NEAT: Evolved compositional pattern producing networks for cellular automata morphogenesis and replication. <em>IEEE Transactions on Cognitive and Developmental Systems</em>. <a href="https://doi.org/10.1109/TCDS.2017.2737082">doi:10.1109/TCDS.2017.2737082</a></li>
  <li>Risi, S. &amp; Stanley, K. O. (2012). An enhanced hypercube-based encoding for evolving the placement, density, and connectivity of neurons. <em>Artificial Life</em>, 18(4), 331–363. <a href="https://doi.org/10.1162/artl_a_00071">doi:10.1162/artl_a_00071</a></li>
</ul>

<script src="/assets/js/es-hyperneat-post.js"></script>]]></content><author><name>eboshii</name></author><summary type="html"><![CDATA[Part 1 of 2. ES-HyperNEAT evolved a better brain for a 3D cell when the network itself was 3D.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://eboshii-dev.web.app/assets/images/teasers/higher-dimensional-substrates.png" /><media:content medium="image" url="https://eboshii-dev.web.app/assets/images/teasers/higher-dimensional-substrates.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>