
danburonline@danburonline1d ago
Classical Mind Uploading: A Potential Candidate for the Great Filter (PUBLIC DRAFT)
Disclaimer: This is a long-overdue (and now slightly refined, after a few discussions with @techno0ptimist) draft of an essay I started writing at @LBF_org last year. Btw, fair warning, parts of it are pretty speculative. What I'm sketching is a possible failure mode, i.e., a civilisation uploads before anyone notices the slow drift and then loses the ability to recognise (let alone correct) what's happening to it; whether that could actually end in collapse – and on what timescale – is something that still has to be modelled properly (something I'd like to do in the final version of this essay at some point). Either way, let's go:
You probably all know Fermi's question (where is everybody?), and what @robinhanson made out of it (https://t.co/jtAOvCL43b): If dead matter so rarely ends up as a civilisation that expands and lasts in the universe, then one or more steps along the way must make the whole path extremely unlikely => the Great Filter. If the hard part is behind us (maybe life itself was the fluke, or the jump to complex cells, who knows), fine, then we just got lucky. If it's ahead of us, most civilisations like ours stop somewhere down the line, and we most likely will too.
But not everything counts as a late filter, as the filter only cares about what we'd actually see from down here. So whatever it is has to stop the lasting expansion/conspicuous engineering we'd otherwise expect to see => if a civilisation keeps expanding and building stuff, anything that goes wrong on its inside simply doesn't exist as far as our telescopes are concerned. And it also has to be close to universal => @anderssandberg, Drexler and @tobyordoxford make the point that something which wipes out 99% of civilisations barely helps, as the survivors would still fill the sky (https://t.co/Lodfg9W0S5; btw, the same paper also suggests we may just be alone, in which case there's nothing to explain in the first place, so I'll keep all of this conditional).
Most of the usual suspects for filters, such as nuclear war or engineered pandemics, face a universality problem, because they would have to prevent lasting expansion in nearly every civilisation. My suggestion comes from wanting not to die and not to be killed (basically all the benefits of substrate-independent minds), which I'd guess is about as universal a motive as mortal minds get.
There are two proposed ways out of natural (biological) death*. The one everyone pictures is what I, if you've been following my tweets, call classical mind uploading: scan the brain, then run its dynamics as a computational model (the whole brain emulation avenue, which Sandberg and Bostrom laid out in detail: https://t.co/AnmkPPEVXs). The other is hybrid mind uploading, which is what I actually work on at @synconeticsorg, where biological neural substrate gets replaced gradually inside a living system, so the original is never switched off and restarted somewhere else (we sketched the approach in Death is an Engineering Challenge: https://t.co/a5cK5ublqh, or go and read some of my tweets from a few days ago).
(* I do make a distinction between solving death and stopping dying. I think the most pragmatic way to stop dying is functional brain isolation (which is what we're working on at @eightsixscience), and the most pragmatic way to solve death is hybrid mind uploading)
Now, classical mind uploading is a possible filter candidate, since I think many civilisations of mortal beings with computers will run into the idea that scan-and-copy might be the computationally obvious exit. Add death pressure on top of that (10s of millions of people die every year!) and it could appeal to a lot more people than just a few curious early adopters => and if there were a hidden flaw shared across implementations, it could affect many of them (and yes, that still wouldn't give us universality "for free").
What would that look like from here if another civilisation does classical mind uploading? I think this is where the Barrow scale is really useful. Most people think of advanced civilisations in Kardashev terms, i.e., how much energy they harness (planet, star, galaxy), whereas John Barrow flipped it around and ranked them by how small they can go instead – objects at their own scale, then genes, molecules, atoms, nuclei, elementary particles and, at the very end, spacetime itself (Ćirković has a good discussion of both scales here: https://t.co/eGHxnlEAx0). And if you think about it for a bit, a classically uploaded civilisation could have strong reasons to move down the Barrow scale => everyone lives inside computation, and smaller, denser or colder systems can offer advantages, within engineering limits. Now, push that outwards and you get Robert Bradbury's Matrioshka brain (https://t.co/CXCyDq6q9d): shells nested around a star, each one computing on the waste heat of the one inside, with the outer shells radiating at progressively lower temperatures. Tbh, that's roughly the image I started with when I first asked myself all this, a graveyard of p-zombies in Matrioshka brain situations or something (I've since dropped the p-zombie part => more on that below), i.e., uploads which after a few centuries of drift aren't the people who built them anymore, and maybe, if substrate views are right, not anyone at all.
Push it the other way (inwards) and you end up with @johnmsmart's transcension hypothesis from 2012: more and more compute squeezed into less and less space, until a civilisation ends up somewhere black-hole-like and more or less disappears from view (decent summary here: https://t.co/rDErIn5Ceq). Or they just ... wait. That's basically Sandberg, Armstrong and Ćirković's "aestivation hypothesis" from 2017, and they even put numbers on it => if all you want is as much computation as possible out of a given energy budget, a colder universe could be well worth the wait. Their estimate went as high as ≈10³⁰ times more computation (https://t.co/wC6zXHKxLh). But it depends on their assumptions about the resources available for computation, which Bennett, Hanson and Riedel challenged in 2019 (https://t.co/9b5Nawe6ls). Also, the aestivators in the original paper don't just nod off on the spot btw, they can expand and secure resources first and only then settle down to wait.
With the outward version astronomers at least have something to go looking for, i.e., waste heat, and so far that hasn't turned up much: Griffith et al. went through roughly 100k galaxies in the WISE data in 2015 and found none consistent with more than 85% of their starlight being reprocessed into mid-infrared waste heat (https://t.co/TMxZvT0baJ), and Project Hephaistos searched about 5 million sources in our own galaxy and came up with seven partial Dyson-sphere candidates in 2024, all apparently M-dwarfs (https://t.co/Faag7HJfJc). Candidates only tho => no alien engineering discovered so far, and for two of them a 2026 JWST follow-up already traced the excess infrared emission to background galaxies (https://t.co/7HbbqphWOa).
And you can't really get around the heat either (not with irreversible computing anyway) => Landauer's 1961 bound: erasing one bit dissipates at least kT ln 2, with T = the temperature of wherever the heat ends up, so ≈3 × 10⁻²¹ J per bit at room temperature (https://t.co/6n1ClOdXvh). Tiny! But not zero, so anything computing that way has to leak SOME heat (on paper, reversible computing à la Charles Bennett could dodge most of that). Whether we'd ever pick that up from down here is another story: temperature / distance / what our instruments can actually see etc. So an empty sky fits pretty much everything: nobody ever there / everyone waiting (aestivation) or doing stuff we wouldn't even recognise / everyone really stopped. (If enough civilisations just decided never to expand, that could be a filter of its own btw, just not the one I mean here => mine is a civilisation stopping without anyone ever deciding to stop).
And whenever I bring this up, sooner or later someone brings up p-zombies (in the loose sense btw, the strict ones are physical duplicates: https://t.co/nLHRgWQy9b), i.e., an upload that behaves exactly like you with nobody home. My go-to answer is still @de_dicto's old but gold 1995 BBS paper, "On a Confusion about a Function of Consciousness" (PDF version: https://t.co/59yxFNGYMc), which pulls apart two things we tend to lump together: P(henomenal) consciousness, the experience itself, the "what it's like" (what cold sea water actually feels like, say), and A(ccess) consciousness, i.e., information that's available for reasoning, for report and for steering behaviour. Blindsight's in there too; e.g., Weiskrantz's patient called "DB" from 1974, who could point at stuff in his blind field while insisting he saw nothing at all. My main takeaway from the paper is more of a warning tbh: whenever someone explains what consciousness does (me included!!), check if they've actually explained A and just called it P.
For the filter argument I mostly care about causes. A is defined via information being available for reasoning and control (available ≠ used btw!), which leaves P as the open question => does experience itself do any causal work, e.g., does what it's like to weigh up a decision help cause the decision? If it does (i.e., P has a causal impact on reality) an upload has to account for that contribution somehow, if it doesn't => our choices aren't caused by how they feel to us, not even our reports about how they feel (big implications for free will, another route I'm not going into here yet). (= Jaegwon Kim's exclusion problem btw, which he pushed for decades, e.g., in Mind in a Physical World (1998) => if the physical causes already do the job, what's left for the mental ones to do?! Overview here: https://t.co/L4zitMOymi) And @AlexLerchner argues for a distinction between simulating something and instantiating it in his abstraction fallacy paper from earlier this year: https://t.co/bdURawh1DY.
So say P does make a difference and an upload lacks it (as some substrate-dependent theories would predict). Then whatever P contributed has either been reproduced some other way (=> we're back at the fidelity problem below), or it hasn't, and there's a hole in the upload's dynamics, i.e., one more way for it to diverge, maybe even one of the sources of the slow drift I'm speculating about here, while the upload still seems totally fine at first (whether that could actually lead to delayed failure is again something for a model to figure out). And if P doesn't do any causal work (epiphenomenalism) => losing it can't explain on its own why a civilisation goes quiet. Either way, if it's going to matter for the filter, it has to show up in behaviour sooner or later. Don't get me wrong, an upload civilisation with nobody home would be horrible!! But it could still keep building Matrioshka brains (or whatever) just fine => the case I'm after is a civilisation that can't keep itself going anymore. And for the record, I'm not assuming substrate dependence anywhere in this, and I don't need p-zombies to be possible, or impossible for that matter (Dennett style!), since the failure mechanism I'm proposing is imperfect emulation, not perfect behaviour without experience.
Anyway, for classical mind uploading to fail in a way that would matter for the sky I think three premises have to hold:
First, fidelity is finite, i.e., a practical computational model will simplify something, and every classical upload that can actually be built is an approximation of the brain it was scanned from; my dear friend Kentaroh Takagaki and his collaborator Frank Ohl's comparison of cortex with the heart shows quite nicely how big that gap is (https://t.co/ipJbcKeqcC). In their deliberately simplified comparison you can get away with treating the heart as a few electrically coupled syncytia (order 10⁰ active elements!), two broad kinds of excitable cell, mostly ticking along on one single dominant timescale. Their brain column: ≈10¹¹ cells, hundreds to thousands of excitable cell types, >4 effective network dimensions, dynamics from sub-millisecond timing up to decades of learning (!!) => HUGE gap. They also point out that a major bottom-up attempt available at the time, the Blue Brain Project (https://t.co/vxzWZMbdiU, btw that's where I worked, so don't forget I'm historically a "classical mind uploader" myself :-)), managed a partial reconstruction of juvenile rat somatosensory cortex that reproduced several experiments, including basic network-state transitions.
For practical WBE you'd need to find a level of description that's sufficient, i.e., one where you can emulate the functions that matter and simplify away the finer details => and so far nobody has established where that level is for an individual human mind! Lorenz explored back in 1969 why prediction can be difficult in an idealised multiscale fluid model. Errors at small scales can spread upwards, potentially imposing a prediction horizon even as initial precision improves (his "real butterfly effect", explained nicely here: https://t.co/weJQQRsROd), and my co-founder @watanabemasata discusses a related hypothesis for the brain as its "web of causality" (https://t.co/U1f1tWzMCV). An atom-by-atom physical rebuild would be roughly Barrow IV-minus territory btw (control over individual atoms), but that's a different proposal from running an emulation altogether, and we don't even know that we'd need that much detail.
Second, the slow variables may integrate systematic error. Divergence on its own isn't the issue, as the cortex diverges from itself all the time: London and colleagues estimated in 2010 that one extra spike in a single excitatory neuron in anaesthetised rat barrel cortex could trigger, on average, ≈28 extra spikes in its postsynaptic targets (https://t.co/cERtyNhcGh), Prinz, Bucher and Marder simulated more than 20 million versions of a three-cell model of a crustacean circuit in 2004 and found that very different parameter sets could give near-identical network output (https://t.co/wulfhcBegr), and learned representations drift over days and weeks while behaviour stays stable, which Rule, O'Leary and Harvey propose the brain might handle with error signals between regions that keep its codes consistent (https://t.co/b8uGA1hGgY). What I'm worried about is way slower than any of that => a systematic bias that slowly builds up in the variables that integrate everything else (plasticity, memory, and the homeostatic set-points that keep activity in range).
Third, and I think this is the part that gets OVERLOOKED THE MOST: the brain keeps those set-points in check through its own physiology (metabolism, neuromodulation, glia, the rest of the body etc.), i.e., the stuff that self corrects/course corrects all the time, and that's what a classical upload has to reproduce through modelled or engineered self-regulation => whether the replacement preserves the regulation that matters has to be established.
I always bring up weather models as a nice example here, since they're coarse-grained physics simulations as well, and everyone's fine with them not being exact. Beyond roughly two weeks they generally lose skill at predicting the detailed weather but still produce something weather-like (storms, fronts, seasons, etc.), even at 50 days. But what about 50'000 days? Or 5 million days? (≈137 and ≈13'700 years btw!) That's coupled climate model land => ocean + atmosphere + sea ice, etc., run for hundreds to thousands of model years and the deep ocean alone can take millennia to equilibrate (so a 10-year test run won't tell you much about the long-term balance). Nasty part: the slow bits can drift off towards a climate the real Earth doesn't have and the weather on top keeps looking perfectly normal! Early coupled ocean-atmosphere models often needed artificial adjustments to the fluxes at the sea surface to keep that drift in check, i.e., so-called flux corrections (Sausen, Barthel & Hasselmann discussed exactly this already in 1988: https://t.co/YKfV2vaRVy). Those were fixes for model biases btw, nobody had discovered some extra physical process in the ocean => and THAT'S the whole analogy; passing a two-week test doesn't establish stability over centuries – which doesn't mean every approximate model eventually falls apart either!
Day-to-day weather forecasting (ECMWF & co.) cheats a bit here btw => it never has to run that long! Every 6 to 12 hours, millions of fresh observations come in (satellites / radiosondes / buoys / aircraft etc.) and get blended into the model state (data assimilation), so the measurements keep dragging it back before it can wander off too far. An embodied upload gets a bit of that for free => misjudge a step and the floor will tell you! But for the person there's no such floor. Sure, we could check memories against records, compare versions, ask friends, watch for changes in behaviour etc., but none of that tells us what the original person would have been like at age 400 (would 400-year-old me still find the same stuff funny? no idea, and no way to check either) => there's no biological 400-year-old me to compare against! So "how far has it drifted since day one?" isn't even the right question on its own, first we'd have to agree on how much a person is ALLOWED to change over a few centuries.
Btw, that's really two different worries: (a) the upload stays perfectly capable but loses features of the particular person it was meant to preserve (the WBE roadmap separates species-generic from individual emulation, and leaves consciousness and personal identity as separate questions, so "it's an emulation" settles neither) / (b) some bias eventually wrecks its ability to function at all. Approximation alone doesn't get you either one tho => depends on what the errors are, and on whether the replacement regulation can keep them in check. Eigen's error threshold (1971) is a nice comparison for the limits of keeping information intact under replication errors (https://t.co/rrquNPadVX), but for uploaded minds nobody has established an equivalent yet (def. something I'd love to work on at some point!).
Four ways this can go (see figure), quick version: (1) stable emulation => how much detail that takes, nobody knows, maybe way more physics than we'd expect, maybe a coarser model does the job / (2) fast collapse => if early trials catch it, people can fix things before a whole civilisation has committed / (3) stable drift => everything keeps working, but the uploads slowly change in ways their original designers would've counted as losing the person (identity questions, maybe consciousness ones too), while the civilisation keeps building stuff (cf. Bostrom's The Future of Human Evolution: a technologically capable future without beings whose lives have value, https://t.co/BHmzBuzIJL) => grim, but on its own no explanation for an empty sky / (4) delayed-onset collapse, the one I actually worry about => decades of uploads looking fine, the civilisation commits, and only THEN does the damage show. That's my filter candidate, IF it also prevents recovery and any further expansion.
Testing (4) is a total nightmare btw => give the uploads a few years and they could pass everything we throw at them: questions / practical tasks / years of ordinary life / even long chats about whether they feel "real and conscious". All useful, sure! But a drift that only kicks in centuries later, how would that ever be obvious from a few years of this?! Next best thing: look at the internal dynamics for early warning signs, if we even know by then which slow changes matter => and even then, no warning in the first few years ≠ no problem!
Add the 10s of millions of deaths per year on top of that, and I find it pretty hard to imagine everyone waiting for a 500-year "soak test" once the first uploads have spent a decade apparently living perfectly good lives, so let's say uptake runs ahead of validation and over 50 to 100 years biological alternatives get rare or disappear (this is a scenario assumption btw, I'm NOT forecasting anything!) => then the control group is gone, and the people maintaining the system are exposed to the same failure as the people they're supposed to monitor. Babinski introduced the term "anosognosia" back in 1914 for patients who were unaware of their own paralysis, and I'm borrowing it here as an analogy for a civilisation whose deterioration also compromises its ability to recognise what's happening (it would have to defeat independent monitoring too).
There are engineering ways around this of course (and the argument has to take them seriously!), e.g., running uploads faster than real time to test centuries of life before anyone commits (only works if the test environment actually exercises the dynamics that matter tho), keeping copies around to compare versions and select the more reliable designs (a "one persistent upload per person" service wouldn't give you that automatically, but nothing stops its engineers from building it in), or checkpoints => roll back to before the damage, repair the cause, recover what you can, maybe losing the experiences since the checkpoint if they can't be safely restored (a serious cost!), but it could still prevent collapse. So for my failure mode, all of these safeguards have to be absent, inadequate or somehow defeated. Schematically (and it's really just a sketch), one version of the scenario needs both;
L(tested) < L(failure)
T(adoption) < T(collapse) < T(escape)
with L = the running age that testing actually covered (simulated years, say, although which "age" is the right one depends on the flaw, as simulated activity, update cycles and hardware ageing needn't tick at the same rate) and T = dates on one and the same external clock. Or in plain English: it breaks later than anything we tested, (nearly) everyone depends on it before it breaks for good, and it breaks for good before anything that could keep expanding on its own has left.
It's not a sufficient condition either btw, and lots of things could break it: different implementations, biological holdouts, autonomous AI, probes etc. E.g. upload colonies might share the flaw without failing at the same time (and one that fails eventually could've sent something further out by then), and a Matrioshka brain might even stay infrared-bright after its uploads failed, just by absorbing and reradiating starlight => so the collapse would ALSO have to prevent (or eventually remove) the very signatures we claim are missing!
Would nearly every civilisation end up there? No idea tbh, wanting to live doesn't mean everyone picks the same method, and the ones picking uploading won't necessarily all build it the same way either. What I find most interesting is the timescale mismatch, i.e., if on a small scale (few years) the uploads seem "real and conscious", over a span of 50–100 years (nearly) everyone does it, and the drift only shows up within, say, 500 years, how would anyone notice in time? But: to explain the Great Silence with it, that shared weak spot would have to get past every single chance to detect it / repair it / escape it. And that's the bit I'd actually want to model => when does the drift become detectable, and can adoption + dependence get ahead of that point? The papers above gave me good reasons to ask, none of them answers it for uploads.
Could the same thing happen with gradual brain replacement? Yes, potentially! A replacement could seem to work fine and still get some slow process wrong. Keeping the brain running through the transition doesn't establish that it'll stay stable for centuries either => whether that setup makes the problem easier to catch and correct is part of what I'd want to model.
Why I think mechanistic gradual brain replacement (i.e. hybrid mind uploading) is the more conservative place to start is a whole other post btw. Feedback very welcome, esp. on the modelling part!