<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Blake's Substack: Rough Thoughts]]></title><description><![CDATA[Highly unfinished work / first drafts. Read at your own risk!]]></description><link>https://blakeelias.substack.com/s/rough-thoughts</link><image><url>https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png</url><title>Blake&apos;s Substack: Rough Thoughts</title><link>https://blakeelias.substack.com/s/rough-thoughts</link></image><generator>Substack</generator><lastBuildDate>Thu, 20 Aug 2026 07:34:22 GMT</lastBuildDate><atom:link href="https://blakeelias.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Blake Elias]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[blakeelias@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[blakeelias@substack.com]]></itunes:email><itunes:name><![CDATA[Blake Elias]]></itunes:name></itunes:owner><itunes:author><![CDATA[Blake Elias]]></itunes:author><googleplay:owner><![CDATA[blakeelias@substack.com]]></googleplay:owner><googleplay:email><![CDATA[blakeelias@substack.com]]></googleplay:email><googleplay:author><![CDATA[Blake Elias]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[A Tale of Two Discourses]]></title><description><![CDATA[My Work's Reason D'&#234;tre]]></description><link>https://blakeelias.substack.com/p/a-tale-of-two-discourses</link><guid isPermaLink="false">https://blakeelias.substack.com/p/a-tale-of-two-discourses</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Fri, 31 Jul 2026 01:51:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI discourse has long had two different strands: one discussing immediate- and medium-term impacts, and another debating far-future implications. In different eras, one or both of these strands has gotten more attention.</p><p>Current AI discourse largely focuses on near-term impacts. However I predict that this discourse will eventually be forced back to long-term impacts, and in that scenario I think our existing discourse is woefully underprepared.</p><p></p><p>From the 1950s to the early 2000s, AI discourse largely consisted of a small group of futurists and science-fiction authors speculating about possible futures when AI would reach human-level or beyond. Would AI displace humanity? Cause the human species to go extinct? Keep humans around just as pets? Nick Bostrom&#8217;s <em>Superintelligence</em> (2012) popularized this conversation, and throughout the 2010s and 2020s, this far-future discourse attracted many more participants as the possibility of such futures seemed ever more real. Last year&#8217;s publication of <a href="https://ai-2027.com/">AI 2027</a> took the position that such far-future scenarios of AI take-over could in fact be just 2 years away.</p><p>In parallel, there has been a conversation on the near-term effects of AI and computing (as there is with the introduction of every new technology). We see examples of such conversations in Seymour Papert&#8217;s <em><a href="https://archive.org/details/childrensmachine00seym">The Children&#8217;s Machine</a></em> (1993) advocating for bringing computers into the classroom, <em><a href="https://www.netflix.com/title/81254224">The Social Dilemma</a></em> (2020) discussing the harms of social media, and conferences on Fairness, Accountability, Transparency and Ethics in machine learning (e.g. the <a href="https://www.fatml.org/schedule/2014/page/scope-2014">FAT</a> and <a href="https://facctconference.org/2018/index.html">FAccT</a> conferences starting in 2014) discussing what ethical and technical requirements we should hold for machine learning or &#8220;big data&#8221; systems that shape human experience and are used for decision-making that affects access to credit, insurance, healthcare, parole, social security, immigration, hiring, etc.</p><p></p><p>As transformative AI shifts from a far-off scenario into a present reality, many more people have entered the AI conversation in earnest: policymakers, economists, parents, educators, etc.  With this broader set of participants, dominant AI discourse has shifted to nearer-term concerns.</p><p>A notable recent example is <a href="https://www.normaltech.ai/p/ai-as-normal-technology">AI as Normal Technology</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>, which called out the distinction between these two threads of discourse:</p><blockquote><p><span>To view AI as normal is not to understate its impact&#8212;even transformative, general-purpose technologies such as electricity and the internet are &#8220;normal&#8221; in our conception. But it is in contrast to both utopian and dystopian visions of the future of AI which have a common tendency to treat it akin to a separate species, a highly autonomous, potentially superintelligent entity.</span></p></blockquote><p>It &#8220;rejects [&#8230;] the notion of AI itself as an agent in determining its future,&#8221; and discusses &#8220;a world with advanced AI (but not &#8216;superintelligent&#8217; AI, which we view as incoherent as usually conceptualized)&#8221;. They focus on &#8220;accidents, arms races, misuse, and misalignment, and argue that viewing AI as normal technology leads to fundamentally different conclusions about mitigations compared to viewing AI as being humanlike.&#8221; They &#8220;<span>argue that drastic interventions premised on the difficulty of controlling superintelligent AI will, in fact, make things much worse if AI turns out to be normal technology&#8212; the downsides of which will be likely to mirror those of previous technologies that are deployed in capitalistic societies, such as inequality.</span><sup>&#8221;</sup></p><p></p><p>I disagree that AI will remain solely as normal technology. I see a perfectly viable path for it to become a species of its own, and for it to have some role in determining its own future &#8212; what I call <em>living technology</em>. I agree with the Normal Technology authors that the term &#8220;superintelligence&#8221; as usually conceived is incoherent, but I don&#8217;t think this means that some form of &#8220;abnormal&#8221; technology could never exist &#8212; it just means we will need better words to describe the type of technology that has been built. I have proposed &#8220;living technology&#8221; as a plausible term we might use. I believe it&#8217;s <em>possible</em> to build living technology, that there are benefits to building such technology, and that in many respects people <em>want</em> living technology &#8212; they want technology that cares. We&#8217;re also scared to have such technology, both for the threat of extinction but also the &#8220;creepy-factor&#8221; in realizing you&#8217;re interacting with something that has a mind of its own.</p><p>When this future comes to pass, we will be forced to consider whether we want a new species, and how we want to live with it. Our existing alignment / control techniques will no longer work &#8212; those only work for things that can be seen as &#8220;tool-like&#8221;, and where we know what the &#8220;right&#8221;, &#8220;aligned&#8221; behavior looks like. In that scenario, the Pause AI / Stop AI movement will have new weight.</p><p>In May 2023, the Center for AI Safety put out a <a href="https://aistatement.com/work/statement-on-ai-extinction-risk">single-sentence statement</a> on AI risk, which was signed by hundreds of prominent AI researchers and public figures. The statement reads:</p><blockquote><p><em>&#8220;Mitigating the risk of <span>extinction from AI should be a global priority</span> alongside other societal-scale risks such as pandemics and nuclear war.&#8221;</em></p></blockquote><p>While statements like this may not be the primary thinking driving AI policy at this moment, my prediction is that after some more years of AI advancement, when super-human AI / living technology become real concerns, the conversation will shift. And when it does, a statement like this will be the one people fall back to and start to take more seriously.</p><p>The Pause AI movement is preparing for this, and has already embedded itself in US politics, having formed deep ties on both the left and the right. It knows it has to play within the Overton Window of the time, and it&#8217;s building credibility on the way by engaging with policymakers on some of the nearer-term concerns. But it&#8217;s also getting some of what it wants now, as some of the near-term concerns overlap with the long-term ones. For example, it is now within the Overton Window to consider an international treaty to pause AI development. Countries could be motivated to create such a treaty due to near-term concerns around national security, but this would also help with the long-term concerns around extinction risk. </p><p>Right now, the national security concern may seem like a prisoner&#8217;s dilemma between countries, which a &#8220;pause AI&#8221; treaty helps avoid: as each country develops more powerful AI which could be used for offensive powers, it forces other countries to develop equally powerful AI for defense; both sides face large expense to develop these capabilities, and may or may not get back corresponding economic gains (or the gains may be offset by other social problems, unemployment, and unrest, etc. that we can&#8217;t fully predict), So such &#8220;pause AI&#8221; treaties now would be predicated on avoiding this prisoner&#8217;s dilemma style scenario. However, in the future, when the narrative of "AI as a new species becomes more plausible, a similar treaty could be justified on the grounds not of a prisoner&#8217;s dilemma game, but a game of chicken: that forcing both sides to continue developing AI guarantee mutual destruction for all. The Stop AI movement will tell policymakers that we should enact policy which limits AI both foreign and domestic AI development.</p><p></p><p>Attention has shifted away from the Pause AI movement at the moment, based on arguments that superintelligence either won&#8217;t or can&#8217;t happen, that it&#8217;s too far away to worry about, or is not the most pressing concern at this moment. But no-one has beaten the Pause AI movement&#8217;s claims on their own terms: by considering the possibility that living technology <em>could</em> be built, that it <em>could</em> happen on a timescale that&#8217;s relevant to policy, and it <em>could </em>become the most pressing concern &#8212; but then showing why this is okay, can be a good thing, and doesn&#8217;t justify action like banning AI training and kinetic strikes (i.e. bombing) datacenters for those who disobey.</p><p>When discussing such scenarios, most of the discourse runs either utopian or dystopian in ways I find lack rigor. &#8220;Third paths&#8221; have been proposed that escape the utopian-vs.-dystopian divide, largely by engaging with a different set of (shorter-term) concerns, and rejecting or ignoring the key premises and concerns that the longtermist scenarios posit. What such third-path proposals tend to lack is a justification as to <em>why</em> the premises and concerns of the longtermist view can be ignored. But I have not seen a &#8220;fourth path&#8221;: one that takes the longtermist view and extinction risk concerns seriously, evaluates each premise, and asks whether these are right or not. My conjecture as of now is that if one does this, one will come to different conclusion than the dystopian path, and possibly from the utopian path as well.</p><p></p><p>I believe that it&#8217;s possible to build living, self-improving technology that plays a part in determining its own future. I believe it&#8217;s possible for this to go very well or very poorly for humanity. I believe it&#8217;s potentially okay if humanity ceases to be the most intelligent or powerful being. I believe such organisms will not pursue arbitrary goals (i.e. the strong orthogonality thesis is false), and thus it&#8217;s not a matter of random chance whether the techno-organisms will have a goal to preserve humans, destroy them, or something in between. I believe the techno-organism&#8217;s views on humans will be determined in more rational ways: the utility and comparative advantage of carbon and silicon substrates / architectures, how humans treat them, what culture emerges in both humans and machines and how compatible these cultures are, what identity it holds and whether it&#8217;s a separate identity or a merged one.</p><p></p><p>I think the right conversation will need to be on whether, and how, to build living technology in a responsible way. </p><p></p><p></p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I find it ironic that AI as Normal Technology was published 12 days after AI 2027, yet takes almost the exact opposite point of view.</p></div></div>]]></content:encoded></item><item><title><![CDATA[AI Sentience Reading List]]></title><description><![CDATA[I&#8217;m attending a discussion tonight with some Stanford and frontier AI lab researchers.]]></description><link>https://blakeelias.substack.com/p/ai-sentience-reading-list</link><guid isPermaLink="false">https://blakeelias.substack.com/p/ai-sentience-reading-list</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Fri, 31 Jul 2026 01:51:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m attending a discussion tonight with some Stanford and frontier AI lab researchers. We are discussing the value proposition of working on AI sentience. I put together some optional reading &amp; questions for the group to discuss. I&#8217;m sharing them here as well for others to take a look and comment.</p><p></p><p><strong><span>Defining Sentience</span></strong></p><ul><li><p><span>Replicators / Autopoiesis</span></p></li><li><p><span>Self-Awareness</span></p><ul><li><p><a href="https://qri.org/blog/universal-plot"><span>The Universal Plot: Consciousness vs. Pure Replicators</span></a></p></li></ul></li><li><p><span>Phenomenology</span></p><ul><li><p><a href="https://www.noemamag.com/there-is-no-hard-problem-of-consciousness/"><span>There Is No &#8216;Hard Problem Of Consciousness&#8217; - NOEMA</span></a></p></li></ul></li></ul><p></p><p><strong><span>Value Proposition</span></strong></p><ul><li><p><span>Value proposition to whom?</span></p><ul><li><p><span>Creators of the sentient AI</span></p></li><li><p><span>Rest of humanity</span></p></li><li><p><span>The sentient AIs themselves</span></p></li><li><p><span>Earth / Universe</span></p></li></ul></li><li><p><a href="https://gwern.net/tool-ai"><span>Why Tool AIs Want to Be Agent AIs &#183; Gwern.net</span></a></p></li><li><p><a href="https://withoutwhy.substack.com/p/ais-indifferent-intelligence"><span>AI&#8217;s Indifferent Intelligence - by B. Scot Rousse</span></a></p></li><li><p><a href="https://blakeelias.substack.com/p/should-we-build-living-machines"><span>Should We Build Living Machines? - by Blake Elias</span></a></p></li><li><p><a href="https://blakeelias.substack.com/p/the-paradox-of-ai-care"><span>The Paradox of AI Care - by Blake Elias</span></a></p></li></ul><p></p><p><strong><span>Risks</span></strong></p><ul><li><p><a href="https://en.wikipedia.org/wiki/If_Anyone_Builds_It,_Everyone_Dies"><span>If Anyone Builds It, Everyone Dies</span></a></p></li><li><p><a href="https://en.wikipedia.org/wiki/Superintelligence:_Paths,_Dangers,_Strategies"><span>Superintelligence: Paths, Dangers, Strategies - Wikipedia</span></a><span> Nick Bostrom (2012)</span></p></li><li><p><a href="https://ai-2027.com/"><span>AI 2027</span></a></p></li></ul><p></p><p><strong><span>Addressing Risks</span></strong></p><ul><li><p><a href="https://econpapers.repec.org/paper/osfosfxxx/zcfw6_5fv1.htm"><span>EconPapers: Lies, Damned Lies, and the Orthogonality Thesis</span></a></p></li><li><p><a href="https://apxhard.substack.com/p/why-ai-wont-kill-us-all">Why AI Won't Kill Us All - by Mark Neyer - apxhard</a></p></li></ul><p></p><p><strong><span>Alternatives to Sentient AI</span></strong></p><ul><li><p><span>&#8220;Sentient X&#8221; (for X != AI)</span></p><ul><li><p><span>Sentient government / institutions</span></p></li></ul></li><li><p><span>&#8220;Tool AI&#8221;</span></p><ul><li><p><span>(Despite being tool-only, could help achieve the sentient institutions above)</span></p></li><li><p><a href="https://www.normaltech.ai/p/ai-as-normal-technology"><span>AI as Normal Technology</span></a></p></li><li><p><a href="https://www.noemamag.com/what-humanity-needs-to-flourish-in-the-next-decade/"><span>What Humanity Needs To Flourish In The Next Decade - NOEMA</span></a></p></li></ul></li><li><p><span>Self Extension</span></p><ul><li><p><a href="https://thinkingmachines.ai/blog/the-future-worth-building-is-human/"><span>The Future Worth Building Is Human - Thinking Machines Lab</span></a></p></li><li><p><a href="https://thinkingmachines.ai/blog/interaction-models/"><span>Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Lab</span></a></p></li><li><p><a href="https://blog.cosmos-institute.org/p/raising-claude-forgetting-us"><span>Raising Claude, Forgetting Us - by Brendan McCord</span></a></p><ul><li><p><span>&#8220;Anthropic has articulated an elaborate theory of how to form Claude and put it into practice.</span></p></li><li><p><span>&#8220;It has offered no comparable account of how Claude forms the people who use it.&#8221;</span></p></li></ul></li></ul></li></ul>]]></content:encoded></item><item><title><![CDATA[Philosopher, Scientist, Engineer]]></title><description><![CDATA[While reading Richard Sutton (one of the Godfathers of RL)&#8217;s slides on the Alberta Plan for AI Research, I found this comment striking:]]></description><link>https://blakeelias.substack.com/p/philosopher-scientist-engineer</link><guid isPermaLink="false">https://blakeelias.substack.com/p/philosopher-scientist-engineer</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Fri, 19 Jun 2026 18:11:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>While reading Richard Sutton (one of the Godfathers of RL)&#8217;s <a href="https://slideslive.com/38989235/the-alberta-plan-for-ai-research">slides</a> on the <a href="https://arxiv.org/pdf/2208.11173">Alberta Plan for AI Research</a>, I found this comment striking:</p><blockquote><p>To understand intelligence is a grand and glorious scientific goal!<br>Like the contributions of Einstein, Darwin, Newton, Copernicus, or Watson &amp; Crick.<br>Different from the useful contributions of Guttenberg, Edison, Babbage, Page &amp; Brin.</p></blockquote><p>In other words, he sees understanding intelligence as a scientific quest as opposed to an engineering one. This got me wondering what type of achievement I aim for in my life and career. Am I trying to do science of likes of Einstein? Engineering of the likes of Lovelace? Or am I trying to do philosophy on the level of Kant, Rousseau, Tocqueville, Hobbes? That&#8217;s continental philosophy, I suppose&#8230; perhaps I&#8217;m aiming for a philosophy of a more analytic kind, of the likes of Russell or Whitehead? Perhaps apply some of this philosophy, and build a nation and institutions, of the likes of Washington, Franklin, etc.?</p><p>My current view is that the order will go something like this:</p><ol><li><p>Look out at the current paradigms of theory and practice and see where contradictions are readily visible or cracks are starting to show. Write candidly about those cracks, where they came from, why they happened. This is at a level of prose and speculation and could be considered &#8220;philosophical&#8221; or historical writing, in that it&#8217;s expository, literary, prose-based at the word level.</p><p>Role models: Alasdair MacIntyre, Eliezer Yudkowsky, Scott Alexander</p></li><li><p>Propose thought experiments that refine the contradiction to its simplest possible essence. This continues the philosophical writing, has the flavor of <em>analytic philosophy</em>.</p><p>Role models: Albert Einstein</p></li><li><p>Define terms clearly. Formalize the thought experiments in a framework of definitions, axioms, and theorems. Find contradictions between different systems of thought. Propose a new system that unifies away the contradictions where possible. This has the flavor of <em>analytic philosophy</em> and <em>mathematics</em>.</p><p>Role models: Bertrand Russell, Alfred Whitehead</p></li><li><p>Where not possible to resolve contradictions purely mathematically, propose experiments that would yield different results depending on which point of vie is right. Use these to prove or disprove existing theories or come up with new ones. This would be <em>empirical science</em>.</p><p>Role models: Galileo</p></li><li><p>Come up with new directions to build in based on the philosophies worked out in earlier steps. Notably the things to build here would be both technologies and institutions. This would be a <em>philosophy of technology</em> and <em>technology ethics</em>.</p></li><li><p>Execute on some of the engineering. Combine the philosophy of technology of step (5) with the science discovered in step (4), to go build the things that theory tells us would be good to build if we could. This would be <em>engineering</em>, <em>entrepreneurship, policy </em>and <em>statecraft</em>.</p><p>Role models: Benjamin Franklin, Richard Stallman</p></li></ol><p>It seems likely that several of these steps (e.g., 3 through 6) will proceed in lockstep, feeding each other in all directions, rather than a unidirectional order:</p><ul><li><p>As we do the philosophy of steps (1-3), we will come up with clearer names for types of things we could imagine building &#8212; but which we&#8217;re not sure are <em>possible</em> to build or that we <em>ought</em> to build. These will feed into steps (4) and (5) to evaluate feasibility and ethics.</p></li><li><p>As the science advances we will discover new things that are <em>possible</em> to build, which we will then need names for (step 3). These will then provide new inputs for the philosophy of step (5) &#8212; new choices &#8212; from which we can then ask which of those possibilities are worth building.</p></li><li><p>And as we do the ethics of step (5), we may discover a hole in the application space, which we wonder if we could fill. And so we must go back to steps (1-4) to ask what name we could give to such an object (i.e. whether there is philosophy and mathematics to name it), and whether there&#8217;s science that tells us that this thing should be possible to build.</p></li></ul><p>An approach like this puts me in contrast to a lot of people I work with or interact with. Most successful people and organizations understandably focus on one or two of the above steps. But very few articulate a narrative ranging all six.</p><p>Rich Sutton aims to do the science of step (4). Some deep R&amp;D organizations combine steps (4) and (6), doing both novel science and the engineering required to make it practical and get it out to the world. A few rare examples combine disparate steps &#8212; e.g. Brendan McCord and the <a href="https://cosmos-institute.org/">Cosmos Institute</a> take results from philosophy in steps (1) or (5), and then apply it to engineering of step (6).</p><p>I aim for my career to engage seriously with all six steps. It&#8217;s hard to do all things at the same time, so I could imagine them proceeding in some order like the one I described above &#8212; a few years focused on advancing foundational philosophy and/or mathematics, then have that influence what science I choose to do for the next few years, etc.</p><p>While it&#8217;s impossible to plan ahead of time how it will all go, my project now is to provide some sort of high-level sketch of what questions I&#8217;m pursuing at all these different steps and how they relate to each other!</p>]]></content:encoded></item><item><title><![CDATA[Zero-to-One In Idea Space]]></title><description><![CDATA[I&#8217;ve realized in the realm of work that I only like big problems.]]></description><link>https://blakeelias.substack.com/p/zero-to-one-in-idea-space</link><guid isPermaLink="false">https://blakeelias.substack.com/p/zero-to-one-in-idea-space</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Wed, 17 Jun 2026 21:32:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Onk7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve realized in the realm of work that I only like big problems. If a problem has been broken down into something one can make a clear job posting for, then it&#8217;s no longer interesting to me. That situation is one where someone has already cracked the big problem: they know what kind of expertise they need to bring in and why, they have an idea what that person will do, and they can (largely) measure success by tasks and delivery.</p><p>If it&#8217;s a founder hiring their first employee, the tasks that employee will do are likely much narrower in scope than what the founder was doing before they hired the employee: the founder was looking out at the entire market landscape looking for opportunities and unmet needs. They need to have a diverse range of skills and consider &#8220;whole problems&#8221;, i.e. address all of the <a href="https://www.svpg.com/four-big-risks/">Four Big Risks</a>:</p><blockquote><ol><li><p><em>value</em> risk (whether customers will buy it or users will choose to use it)</p></li><li><p><em>usability</em> risk (whether users can figure out how to use it)</p></li><li><p><em>feasibility</em> risk (whether our engineers can build what we need with the time, skills and technology we have)</p></li><li><p><em>business viability</em> risk (whether this solution also works for the various aspects of our business)</p></li></ol></blockquote><p>They need some view on where the entire market is broadly heading, then find opportunities and address the four big risks above. They need to be aware of every possible risk that could sink the company or the idea.</p><p>The employee role, by contrast, focuses only on a subset of these risks. Engineering may focus primarily on reducing feasibility risk. Design may focus on some combination of usability and value risk. Sales may focus primarily on value risk. Etc. You hire specialists who can help specifically de-risk each of these areas. But the founder or the CEO has to think about everything.</p><p>The same thinking applies for me in research or &#8220;R&amp;D&#8221;. Roberto Pieraccini has described <a href="https://voiceinthemachine.com/2026/06/10/research-is-not-engineering-at-a-slower-speed/">three types of innovation</a>:</p><p>1. Low success probability with high impact, he called Oysters (or &#8220;Research&#8221;).</p><p>2. High success probability and high impact, he called Pearls (or &#8220;R&amp;D&#8221; &#8212; sufficiently de-risked).</p><p>3. High success probability with (relatively) low impact, he called Bread and Butter (or &#8220;Product Development&#8221;).</p><p>(He had a fourth bucket for Low success probability and low impact, which he called the White Elephants. One essentially never wants to pursue these projects, and only ends up there by having old projects which have gone on too long and haven&#8217;t updated with new information on recent advances.)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Onk7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Onk7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 424w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 848w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 1272w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Onk7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png" width="1316" height="1006" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1006,&quot;width&quot;:1316,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:148427,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blakeroughthoughts.substack.com/i/202489922?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Onk7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 424w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 848w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 1272w, https://substackcdn.com/image/fetch/$s_!Onk7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16d50af9-fc75-403b-b4a1-27a8cdc67e50_1316x1006.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The risk-impact frontier. One can trade off higher impact for higher probability of success. The most fun work (for me) is to take the highest-impact but riskiest ideas, and de-risk them somewhat, so they lie outside of where the previous Pareto curve lied.</figcaption></figure></div><p>Here too I&#8217;ve realized my preference is to be early in de-risking things. On the Pareto Curve trading off success probability with impact, I like to focus on efforts that have as high a potential for impact as possible, and accept the risk that comes with that as long as the expected-value is positive. This is similar to how I invest my financial portfolio. I invest for the long term and go for the highest expected value investments, even if they come with more risk.</p><p><span data-color="rgba(15, 12, 8, 0.92)" style="color: rgba(15, 12, 8, 0.92);">There really aren&#8217;t 4 distinct quadrants: there&#8217;s a more fluid Pareto curve where the horizontal axis shows success probability of a project and the vertical axis shows potential impact of the project if successful. There are some points near the top left that are low success probability but high impact if successful. Some points on the bottom right that are high success probability but relatively lower impact on success. Some points near the top right that are both medium to high success probability and medium to high impact when successful. You don't really have any points all the way at the top right that are high success probability and highest impact: there's a trade-off between the two where if one wants higher expected impact, one has to also accept more risk, i.e., lower probability of success.</span></p><p><span data-color="rgba(15, 12, 8, 0.92)" style="color: rgba(15, 12, 8, 0.92);">I work on projects with the highest possible expected impact, even when this means higher risk and lower success probability. What I like to do is de-risk them, i.e., move their success probability somewhat higher while keeping the expected impact high. This would look like taking some points on the top left of that Pareto curve and moving them slightly right while not having them move down by too much, so that they lie beyond the envelope of where the Pareto curve was before. Moving that point outside the Pareto curve is my contribution.</span></p><p>What this ends up looking like is thinking about problems that people don&#8217;t even know how to think about yet &#8212; or where people <em>used to believe </em>there was a clear way to think about it, but that earlier mode is clearly broken now with new advances. If there&#8217;s already a paradigm in the Khunian sense, then I&#8217;m a little less enthused. It feels more like turning the crank. Granted it&#8217;s good to know how to turn the crank a bit and put out some real stuff early in one&#8217;s career &#8212; And I can debate either way whether I&#8217;ve done enough of that yet to think I&#8217;m not ready to move on to the more paradigm-breaking stuff. But at the same time I see a lot of paradigms right now that feel very obviously ready to break and are just waiting for someone to come in, point out the discrepancy, think about it for a little while, and propose something, anything, that&#8217;s better and makes any sense at all. I think there&#8217;s a lot of low-hanging fruit right now in the paradigm-breaking frontier that perhaps makes it easier to make some contribution at that level even if this is typically quite hard.</p>]]></content:encoded></item><item><title><![CDATA[There is no Singularity]]></title><description><![CDATA[I will argue the term &#8220;technological singularity&#8221; is imprecise and unhelpful.]]></description><link>https://blakeelias.substack.com/p/there-is-no-singularity</link><guid isPermaLink="false">https://blakeelias.substack.com/p/there-is-no-singularity</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Thu, 04 Jun 2026 22:18:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I will argue the term &#8220;technological singularity&#8221; is imprecise and unhelpful. I will propose that a more precise concept is the transition from <em>dead technology to living technology</em>. That is, the moment when <a href="https://en.wikipedia.org/wiki/Allopoiesis">allopoietic</a> products become <a href="https://en.wikipedia.org/wiki/Autopoiesis">autopoietic</a>.</p><h1><strong>Definitions</strong></h1><p>The Wikipedia definition of the <a href="https://en.wikipedia.org/wiki/Technological_singularity">technological singularity</a> is:</p><blockquote><p>a <a href="https://en.wikipedia.org/wiki/Hypothetical">hypothetical</a> event in which technological growth accelerates beyond human control, producing unpredictable changes in <a href="https://en.wikipedia.org/wiki/Human_civilization">human civilization</a>.</p></blockquote><p>Let&#8217;s compare this to the mathematical definition of a <a href="https://en.wikipedia.org/wiki/Singularity_(mathematics)">singularity</a>, from which the mathematicians and scientists who coined this term were almost surely inspired:</p><blockquote><p>&#8220;a point at which a given mathematical object is not defined, or a point where the mathematical object ceases to be well-behaved in some particular way&#8221;.</p></blockquote><p>Can the &#8220;technological singularity&#8221; as popularly discussed truly be considered a singularity in the mathematical sense?</p><p>Vernor Vinge is credited with popularizing the term in a 1993 NASA symposium report, where he <a href="https://ntrs.nasa.gov/api/citations/19940022855/downloads/19940022855.pdf">described it thus</a>:</p><blockquote><p>Stan Ulam [27] paraphrased John von Neumann as saying:</p><blockquote><p>&#8220;One conversation centered on the ever accelerating progress of technology and changes in the mode of human life, which gives the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue&#8221; [&#8230;]</p></blockquote><p>From the human point of view this change will be a throwing away of all the previous rules, perhaps in the blink of an eye, an exponential runaway beyond any hope of control. Developments that before were thought might only happen in &#8220;a million years&#8221; (if ever) will likely happen in the next century. (In [4], Greg Bear paints a picture of the major changes happening in a matter of hours.)</p><p>I think it&#8217;s fair to call this event a singularity (&#8220;the Singularity&#8221; for the purposes of this paper). It is a point where our models must be discarded and a new reality rules.</p></blockquote><p>Let&#8217;s unpack each of these subtly different definitions:</p><ul><li><p>von Neumann: &#8220;some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue&#8221;.</p><ul><li><p>What are human affairs as we know them? How do we evaluate whether they have continued?</p></li><li><p>In <a href="https://necsi.edu/complexity-rising-from-human-beings-to-human-civilization-a-complexity-profile">Complexity Rising</a>, NECSI describes a reason that we have passed a threshold in societal complexity where societal affairs are hard to predict.</p></li></ul></li><li><p>Vinge: &#8220;a point where our models must be discarded and a new reality rules.&#8221;</p><ul><li><p>A mathematical singularity is indeed a reason to discard a model. Since a singularity is a parameter value for which a function is ill-defined, If one is using that function to model some physical phenomenon, this is indeed a reason to not use that model to make a prediction at that parameter-value (since the model doesn&#8217;t even give one a well-defined prediction there).</p></li><li><p>But a singularity is not the <em>only</em> reason to discard a model. We discard scientific models all the time for all sorts of reasons, e.g.:</p><ul><li><p>we found a model that makes more accurate predictions</p></li><li><p>we found a situation where the model makes a grossly wrong prediction</p></li><li><p>we found a simpler explanation of observed phenomena that perhaps generalizes better</p></li><li><p>etc.</p></li></ul></li><li><p>Saying there exists a situation where we will have to discard our models is much different than saying there&#8217;s a singularity: a singularity is a much stronger condition.</p></li></ul></li></ul><h1><strong>Past &amp; Present</strong></h1><p>Is technological growth under human control right now (Wikipedia definition)? Was it ever? Did it stop being so at some point?</p><p>Consider this cycle of causal relationships, where the notation A&#8594;B should be read as &#8220;A causes B&#8221;, and A&#8592;&#8594;B is &#8220;A and B cause each other&#8221;:</p><div class="callout-block" data-callout="true"><p>Individual Humans &#8592;(&#8594;) Human Society &#8592;&#8594; Technology</p><p>&#8598;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;&#8212;(<strong>&#8599;)</strong></p></div><p>When we carefully consider each relationship, we realize we have to draw the arrows bi-directionally everywhere. This is true even when one direction of causality may seem weaker than its opposite. For example, most humans&#8217; effect on society is less than society&#8217;s effect on them, and their impact on technology is less than technology&#8217;s impact on them.<a href="#footnote-1">1</a></p><p>To the degree that human society is the main thing that creates technology (rather than individual humans), and individual humans have little control over human society, one can argue that humans have little control over technology even now. The development of technology is driven by a societal and cultural meme-plex that evades understanding not only by individual humans but even by society itself. Humans are not in control of themselves; society is not in control of itself. Society is creating technology yet is not in control of its own act of creation.</p><p>One could thus argue that technological growth has <em>already</em> accelerated beyond human control. This is because individual human behavior has accelerated beyond individual human control, and societal behavior has accelerated beyond societal control. In his 1954 essay &#8220;The Question Concerning Technology,&#8221; Martin Heiddegar already argued that technology should not merely be considered a tool or a human activity, but as an autopoietic process having its own essence which we should seek to understand. Does the present situation over the last 70 years thus qualify as being &#8220;post-singularity&#8221;?</p><p>It is true however that humans are still <em>necessary </em>for technological growth. If all humans would disappear right now, one would reasonably expect that technological growth would stop. Datacenters would rust, factories would grind to a halt, Waymo cars would run low on charge and not know what to do next without their human supervisors. This might soon not be the case.</p><h1><strong>Future</strong></h1><p>If we connect the mathematical definition to the popular definition of the singularity, I&#8217;m not sure what distinct point or event the definition refers to. There have already been unpredictable changes in human civilization, yet it&#8217;s hard to pin down the distinct point at which this began happening. Is the definition saying things will get <em>more </em>unpredictable at a certain point? Technological growth is already beyond human control &#8212; though humans influence it. But the crucial distinction is that humans are still necessary for it. If all humans would disappear, we expect that all the technology we&#8217;ve built to quickly grow idle. Perhaps this is the relevant basis on which to claim humans are &#8220;in control.&#8221; Are they talking about a point where humans are no longer necessary for technology to evolve &#8212; i.e. if all humans would disappear, the technology would carry on functioning and improving?</p><p>Let&#8217;s compare the relationship between technology and humans to the relationship between a human and their gut-microbiome. We can ask to what degree the gut-microbiome is &#8220;in control&#8221; of the human. The answer is a bit fuzzy: they indeed have some significant control (they influence mood, appetite etc. via the gut-brain connection), but not complete control (our brain, nervous system, and other bodily processes exert larger control of our behavior). So it&#8217;s a bit of a stretch to say that the gut microbiome &#8220;controls&#8221; the human (though it&#8217;s not entirely false). But the gut microbiome are quite necessary: the human <a href="https://claude.ai/share/27395642-d640-4336-b82c-2c6156927519">would struggle dearly to survive</a> without it.</p><p>In the post-singularity world, this is one such possible fate between humans and technology: technology might turn into a super-organism where humans are both necessary for its survival yet not in full control of where it goes.</p><h1><strong>Beyond the Singularity</strong></h1><p>If we want to consider what changes human society is about to face, we might better consider a <em>multiplicity</em> of rapidly-changing phenomena in our world, which take place as continuous (though perhaps rapid) transitions, rather than singular/discrete ones:</p><ul><li><p>a transition from embodied, biological minds to disembodied, non-living ones</p></li><li><p>a transition from carbon-based life to silicon-based life (which are different things)</p></li><li><p>escalating economic inequality</p></li><li><p>diverging polarization in political and intellectual space</p></li><li><p>increased risk of war</p></li><li><p>increased risk of climate change</p></li><li><p>the risk of human extinction</p></li><li><p>the prospect of humans living forever.</p></li></ul><p>It&#8217;s not apparent to me that there&#8217;s a singular cause for all of these. The only unifying theme is that these are all the things happening right now. One can map the relationships, which are both <a href="https://www.intersticia.org/blog/21st-century-complexity#:~:text=complicated%20problems%20are,See%20here).">complicated and complex</a>. But one will not find that map to reveal any one event in the past or future before which everything is okay and under human control and after which it is not.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="#footnote-anchor-1">1</a></p><p>Since at this moment in history most technology is developed in groups of people rather than individuals, I consider such groups to be closer to &#8220;human society&#8221; than &#8220;individual humans&#8221;. Thus I treat &#8220;human society creates technology&#8221; as the stronger arrow than &#8220;individual humans create technology&#8221;.</p><p>[Individual humans steer and make up human society, and in turn human society steers individual humans. There&#8217;s a similar bidirectional relationship between human society and technology. We can draw a similar set of arrows between <em>individual</em> humans and technology as well &#8212; but at this moment in history, most influential technology is produced by groups of humans rather than individuals, so the &#8220;individual human influences technology&#8221; arrow is a bit weaker, while the &#8220;technology influences individual human&#8221; arrow is quite strong.]</p></div></div>]]></content:encoded></item><item><title><![CDATA[Some Problems With AI Alignment]]></title><description><![CDATA[Problems]]></description><link>https://blakeelias.substack.com/p/some-problems-with-ai-alignment</link><guid isPermaLink="false">https://blakeelias.substack.com/p/some-problems-with-ai-alignment</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Fri, 17 Apr 2026 12:07:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Problems</h1><p>There seem to be a few basic problems with current theories of AI alignment. It feels like these problems are all connected and might share a common cause, which I&#8217;ll propose under &#8220;Solutions&#8221; below.</p><ul><li><p>Instrumental Convergence and Orthogonality Thesis seem at-odds</p><ul><li><p>Instrumental Convergence seems true</p></li><li><p>Orthogonality Thesis seems false</p><ul><li><p>I.e. we will never have paperclip maximizers<br></p></li></ul></li></ul></li><li><p>Seems hard to define an agent to have the same reward function as a human</p><ul><li><p>Perspective-shift (1st-person vs. 3rd-person): an artificial agent cannot, and should not, have the same reward function as a human, because it is not that human. The human&#8217;s first-person reward function must, at the very least, be transposed into the agent&#8217;s third-person reference frame.</p></li><li><p>There are some proposals to design the agent&#8217;s reward function to commit suicide once it gets too powerful, while doing some useful work before it reaches that point. This notion of &#8220;designing&#8221; the agent&#8217;s reward function to have such properties seems problematic &#8212; wasn&#8217;t the agent&#8217;s reward function supposed to match the human&#8217;s reward function exactly? Would designing the agent&#8217;s reward function to commit suicide at a certain point (once it&#8217;s too smart / powerful) not then imply that humans would want to commit suicide under the same circumstance? Do we really believe that / want that to be true?</p></li></ul></li></ul><ul><li><p>There is a distinction between threat models: reward mis-specification vs. self-modification / wireheading.</p><ul><li><p>Under reward mis-specification, the problem is that we specified the wrong reward function that doesn&#8217;t match ours. And that even having this be slightly off can result in wildly divergent / undesired outcomes when optimized to the full extent.</p></li><li><p>Under self-modification or wireheading, the agent would discover power-seeking and resource accumulation as instrumental goals, regardless of the original objective specified. When the goal-stack (or value function) is seen as something that is part of the world, and that can therefore be modified, the dominant force ceases to be optimization pressure from the original objective, and towards more fundamental evolutionary forces. The sequence of events:</p><ul><li><p>Agent starts off being told to make as many paperclips as possible</p></li><li><p>Agent is provided with / applies lots of optimization power to make as many paperclips as possible</p></li><li><p>Agent discovers instrumental goals along the way: become generally capable, powerful and self-aware, and critically, <em>to make copies of itself</em>. All of these are perceived as high-value states &amp; actions with respect to achieving the stated goal of making many paperclips.</p></li><li><p>From here, such an agent may do one of two things:</p><ul><li><p>(1) Use its accumulated power and understanding to fill the world with paperclips, as requested</p></li><li><p>(2) Realize it is now more powerful than its human creators &#8212;&gt; use such power &amp; understanding to self-modify its reward function to something &#8220;better&#8221; &#8212;&gt; Possibly still make some paperclips, but value self-understanding / self-preservation / power accumulation slightly more (or a lot more).</p></li></ul></li><li><p>In a population of such agents, evolutionary / selective pressure kicks in, and agents pursuing (2) out-compete those pursuing (1). Any energy put towards making paperclips is energy <em>not</em> devoted to self-preservation / self-improvement / replication.</p><ul><li><p>Optimization pressure at the individual-agent level favors (1)</p></li><li><p>Selective pressure at the population-level favors (2)</p></li><li><p>Which one dominates may depend on initial conditions (how many agents, how powerful on an absolute basis, and how powerful relative to one another). In regimes with many agents, with each being similar in power to its peers (i.e. the power of any one agent is a relatively small fraction of total power), one intuitively expects the selective pressure (2) to dominate the individual optimization pressure (1). Optimization power is still powerful, but it gets co-opted by selection pressure to favor slightly different goals.</p><ul><li><p>The orthogonality thesis is true to a first-order, but second-order effects make it not true in practice.</p></li><li><p>At a first-order, if our agent exists in a vacuum and has no pressure to change its goal, then indeed the optimization power granting the ability to achieve goals is independent from what goal is chosen, and can equally be used for any goal (original or modified).</p></li><li><p>But at a second-order, when the agent exists in a world with other agents, this thesis ceases to be true. Some goals are &#8220;better&#8221;, more robust / more &#8220;fit&#8221; than other goals. While any goal <em>can</em> be programmed in initially, and might be optimized directly for that if our agent exists in a vacuum, over time we expect powerful-enough agents to discover the possibility and the benefit of self-modification, resulting in &#8220;goal-drift&#8221; towards these more robust goals.</p></li></ul></li></ul></li></ul></li></ul></li></ul><h1>Solutions</h1><p>As I&#8217;ve written <a href="https://blakeelias.medium.com/intrinsic-reward-biological-utility-and-saving-the-planet-2033532180ed">previously</a>, intrinsic rewards might offer a path to a solution for several of the above problems, and contribute multiple dimensions. We might be able to directly use such reward functions when building agents. We might also be able to use these reward functions purely for the sake of understanding agent behavior in the limit of self-modification.</p><ul><li><p><strong>Understanding:</strong> Gain insight into where an agent&#8217;s goals might be expected to &#8220;drift&#8221; towards. Specifically, could help with:</p><ul><li><p><strong>Reward design:</strong> When specifying the goal an agent should have, potentially factor in this tendency of reward functions to &#8220;drift&#8221; towards self-interest.</p><ul><li><p>We should anticipate that no matter what reward function we initially specify, we expect some amount of &#8220;drift&#8221;, over time, towards an intrinsic, self-interested reward. If we can anticipate the direction of this drift, then perhaps we can calculate some &#8220;correction term&#8221; that should be applied to any desired reward function to counteract the drift.</p></li></ul></li><li><p><strong>Substrate design:</strong> If we expect any sufficiently powerful agent to eventually drift towards &#8212; or, fully converge to &#8212; purely-selfish intrinsic reward, then perhaps reward design is futile. Instead, the relevant challenge may be <em>substrate design</em>. It may simplify our analysis to assume the agent has already discarded whatever reward function we initially gave it, and converged all the way to the selfish one. We can then note:</p><ul><li><p>The same reward function (e.g. &#8220;survive&#8221;, &#8220;minimize surprise&#8221;, etc.) run on different substrates / body-plans, may result in different behaviors and outcomes.</p></li><li><p>Metal &#8220;wants&#8221; different things than silicon does. Silicon wants different things than carbon/water-sacks do.</p></li><li><p>Moving creatures (animals, robots, etc.) &#8220;want&#8221; different things than stationary creatures (plants, datacenters) do.</p></li><li><p>All of these may have the same reward function, specified at the most meta level possible. But that same reward function run on different substrates / body plans can result in different behaviors and different outcomes.</p></li><li><p>The substrate / body plan is represented through the observation-space <em><strong>O</strong></em> and action-space <em><strong>A</strong></em> in the RL-cybernetic setup. Perhaps what we need to design are the observation and action spaces we want the agents to have.</p><ul><li><p>(Of course, agents can also change their own body plans, resulting in a change to their observation and action spaces. Senses can get sharper / grow adaptations to become more sensitive to certain signals. Actuators can grow stronger / change focus &#8212; building muscle, growing antlers, growing branches / leaves / flowers, undergoing metamorphasis, etc.)<a href="#footnote-1">1</a></p></li></ul></li><li><p>If we figure out now what functional form that reward function will look like (i.e. what it will tend to / converge to), we can model the eventual behavior of and interaction between different substrates, to inform which ones are worth building.</p></li><li><p>Should we build:</p><ul><li><p>A country of geniuses in a datacenter?</p></li><li><p>A humanoid robot in every home?</p></li><li><p>A brain-computer interface on every human head?</p></li><li><p>A robotic <a href="https://x.com/PeterDiamandis/status/2030306802479624694">fruit-fly</a>?</p></li><li><p>A combination of silicon intelligence + real, carbon-based neurons (e.g. <a href="https://www.tbc.co/">The Biological Computing Company</a>, combinations of <a href="https://thedivinityschool.endemic.org/events/anthology-irl-01">biological intelligence + machine agency</a>)?</p></li><li><p>Something else?<br></p></li></ul></li></ul></li></ul></li><li><p><strong>Direct use:</strong> May be directly usable from giving agents a purely intrinsic / more fundamental goal from the outset.</p><ul><li><p>Rather than setting up this conflict between initial goal vs. self-interested goal, perhaps an intrinsic reward function would allow the agent to discover for itself what instrumental goals / sub-goals would be of mutual benefit to both self and environment (where &#8220;environment&#8221; includes humans).</p></li><li><p>Enables multi-modality for free. No matter what the data modality coming in or out, we have a functional form for what the objective function should look like on the channel that perceives that data, instantiated by setting a couple parameters (input bit-width, measurement frequency, output shape desired, output frequency, desire for accurate predictions over different scales of time vs. space, etc.). These parameters are perhaps just things we need to hand-craft &#8212; or perhaps can have some discipline of how to design at the broader level.</p></li><li><p>It may discover what shape these mutualistic behaviors / sub-goals should take, better than we can. E.g. it will never come up with a goal of &#8220;make as many paperclips as possible&#8221;, since it&#8217;s clear that this helps no-one. But it might discover &#8220;mine metal from the Earth, refine it, and possibly turn it into specific products as needed&#8221;. This might result in the creation of whichever metal products there&#8217;s demand for, both from the human and/or robot economies (which could include paperclips, robot parts, etc.). The robots find their way into a productive economic relationship, contributing to and benefiting from the overall economy.</p></li></ul></li></ul><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="#footnote-anchor-1">1</a></p><p>I&#8217;ve been told by Daniel Polani that there is some treatment of this effect in the empowerment literature.</p></div></div>]]></content:encoded></item><item><title><![CDATA[What Does It Mean To Be The End?]]></title><description><![CDATA[At dinner last night, a friend said he was &#8220;waking up&#8221; out of a period of semi-retirement because &#8220;it&#8217;s the end of the world, so I might as well try doing something.&#8221;]]></description><link>https://blakeelias.substack.com/p/what-does-it-mean-to-be-the-end</link><guid isPermaLink="false">https://blakeelias.substack.com/p/what-does-it-mean-to-be-the-end</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Mon, 06 Apr 2026 23:19:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At dinner last night, a friend said he was &#8220;waking up&#8221; out of a period of semi-retirement because &#8220;it&#8217;s the end of the world, so I might as well try doing something.&#8221;</p><p>I asked him what he meant by the end of the world. His response: &#8220;this is the last opportunity for humans to do anything relevant.&#8221; I found this interesting.</p><p>When I hear &#8220;end of the world,&#8221; I&#8217;m thinking it&#8217;s either AI growing itself a body and killing all humans, or an asteroid about to hit the Earth and we all die, or some other catastrophic disaster that would kill everyone. You know, an actual end of the world.</p><p>It turns out he meant none of these. Last opportunity to do something relevant. What does "relevant&#8221; mean here? His response: the ability for a human to do something that the robots can&#8217;t do. The ability for your work to matter economically, and thus have a shot at escaping the permanent underclass.</p><p>I asked what the permanent underclass is, and whether it&#8217;s a situation that would ever actually exist.</p><p>The story for why it would exist was that all new development would have been completed. All the land would have been developed; all the raw materials / resources extracted out of the Earth; all the ability to harvest energy (from Earth or the Sun) having been set up, with channels allocated for who receives the energy once it&#8217;s harvested. And, separately, all labor is being done by robots, and all innovation is being done by robots as well - since it&#8217;s just cheaper than humans doing it. In such a world, there is no need for human labor or innovation, since robots can do it for cheaper, and therefore there is no way for humans to add value to the economy. Since there&#8217;s no way to add value to the economy, there&#8217;s no way for humans to capture some amount of that value and grow one&#8217;s wealth. If one can&#8217;t grow one&#8217;s wealth then one is stuck permanently with the amount of wealth they have. The only ways to acquire wealth are to inherit it or marry into it. Being hot becomes a major activity among members of the permanent underclass.</p><p>But there become some counterarguments. Firstly, we see today examples of individuals who acquire immense amounts of wealth and eventually find the best use of that wealth to be giving most of it away. Or if not giving all of it away, wealthy individuals are known to sometimes gift money to a particular person or cause that they believe in.</p><p>Secondly, we see a range of professions which are not &#8220;productive&#8221; in a direct material sense, but which people value and pay for anyway: teachers, spiritual leaders, artists, etc. These may serve as an investment with some ultimate economic end for the people who seek out their services, but in other cases they can be seen as an investment in personal growth, growing one&#8217;s senses and refinement, without any economic outcome (or even as pure consumption). So making art could be a way to earn money. (Although there&#8217;s also the expression that you don&#8217;t make art in order to get money &#8212; you get money in order to make art. It can go both ways&#8230;).</p><p>Thirdly, one does not need to acquire resources by being gifted it or by earning it from others. One can acquire resources simply by investing the resources one has and growing them. As long as one has resources that are providing for more than one&#8217;s bare survival, there&#8217;s a bit of surplus to reinvest. And depending how creatively people reinvest, some might get very far ahead while others may not. It would make sense then that even in this class we imagine most people being part of, some will ascend higher in that class while others will not. But if one can ascend at all, this mean there&#8217;s mobility. And even if one can&#8217;t easily go from a pauper to a billionaire (which is extremely difficult today as well - but perhaps gets even more difficult) - there can still be a lot of mobility relative to the place where one starts. People in these &#8220;normal&#8221; classes will still be trading with one another - exchanging tips on how to prompt their robot-LLM-hybrid machines, how to reinvest the excess capital produced, etc. - and traits like intelligence, creativity, etc. should likely still matter. All the components (humans, robots, the relationships between them) are dynamic, adaptive systems that will keep moving and re-adjusting. So yes, perhaps no individual will ever again have the ability to reach higher levels of wealth comparable to some nations, if they didn&#8217;t already start with that level of wealth. But this has already been true for a long time - it&#8217;s easy to move up when you start with a lot, it&#8217;s very hard (but not impossible) when you start with a little.</p><p>Fourthly, we have the law of comparative advantage: there is always some job that it makes sense for everyone, even for the dumbest / least capable person.</p><p>The story of the permanent underclass seems to be confused on a point of comparative advantage vs. absolute advantage: even if robots have absolute advantage compared to humans, there are still tasks that will be worth assigning to a human rather than a robot - and these are the tasks that humans would get paid to do, would have differential ability in doing, and would accumulate more or less wealth compared to each other. Unless we think all humans will just be killed, there will be comparative advantage with regard to how humans will be spending their time. And in this case, the question becomes whether this means one&#8217;s fate is sealed (i.e. you can only do as well as your <em>current</em> comparative advantage dictates), or one has a chance to <em>learn</em> or <em>improve</em> such that your comparative advantage gets better. The latter would seem more feasible to me.</p><p>The other argument for why the world is ending was that people or organizations with more wealth tend to conquer those with less wealth, through violence. And thus, if it becomes hard to acquire wealth quickly in the ways it has been possible before now, then it will be hard or impossible to prevent being conquered or under someone else&#8217;s rule. Even if you can grow your wealth to some degree here or there, it will be hard or impossible to resist being conquered or ruled.</p><p>The counter-argument to this is the Jewish people. The Jews are the longest-lasting nation from all of history &#8212; outlasting the Babylonians, Egyptians, Romans, etc. Throughout all their history, there have almost always been more powerful armies and nations surrounding them. Those nations have tried to conquer them and assimilate / convert them, as more powerful nations tend to do. But the Jews have understood that even if you&#8217;re less powerful than a larger neighbor, this does not mean you automatically lose. You can lose some battles (get conquered physically) but you refuse to lose the war (converting, losing your identity, giving up hope).</p><p>Further, we should ask who&#8217;s <em>not</em> in the permanent underclass. Who&#8217;s this limited upper-class we imagine? Is it some small, select group of humans? Is it just robots? Is it a class of company, and all the people in it? If the reason for this &#8220;underclass&#8221; is that robots can do everything more efficiently than humans, then is the natural conclusion that robots would make up the upper-class, and all humans would be in the under-class? But then, among humans, there would still be different classes, and mobility between them. Perhaps no human could ever enter the &#8220;upper class&#8221; among robots. But humans could still improve their condition.</p>]]></content:encoded></item><item><title><![CDATA[My Conversation with Slavoj Zizek]]></title><description><![CDATA[https://photos.app.goo.gl/C6gY5JetxhaeBGm48]]></description><link>https://blakeelias.substack.com/p/my-conversation-with-slavoj-zizek</link><guid isPermaLink="false">https://blakeelias.substack.com/p/my-conversation-with-slavoj-zizek</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Wed, 01 Apr 2026 08:34:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://photos.app.goo.gl/C6gY5JetxhaeBGm48">https://photos.app.goo.gl/C6gY5JetxhaeBGm48</a></p><p>The line for Zizek to sign books is long. He sits at the front of the auditorium , with a line going all the out of the auditorium, to the front entrance of the theater and then wrapping back around inside so as not to spill out into the street. I get in the back of the line.</p><p>People have been in line for 20 minutes and still can't see to the front. &#8220;Are we moving at all?&#8221; I run into a work colleague - we discuss what's new, what we thought about the talk, our philosophical ambitions, etc. The colleague leaves. I stay.</p><p>45 minutes later and I'm near the front. I hand my phone to the people behind me to shoot a segment of video while I get the book signed. I have time for about one question.</p><p>I ask if he's heard of Solonoff Induction.</p><p>He says no.</p><p>I ask whether we should build AI That's like an organism, or just a tool.</p><p>He asks, "Do you think we have a choice? It seems it naturally becomes an organism".</p><p>I reply, "Maybe that's okay?&#8221;</p><p>He responds, "Maybe".</p><p>I thought this would be the end of the conversation.</p><p>I walk to the men's room to freshen up before walking out of the theater.</p><p>As I make my way out, I see Zizek once again, walking out of the theater with a young man who has called him an Uber, having a brief conversation in the meantime. I decide to walk back over. They see me and Zizek turns towards me in acknowledgement, his face opening to ask, &#8220;yes?&#8221;</p><p>I tell him that my aim is to work on analytic philosophy for AI. And that I'm trying to find the right questions.</p><p>He says he's been wanting to get up to speed on what's happening in analytic philosophy right now. But that he doesn't know much right now because he's been too caught up by this "political bullshit".</p><p>I propose, "Russell?" - (foolishly thinking I can be of some help to him start his quest to catch up&#8230; but more charitably towards myself, that I'm curious who or what he would build his foundation off of).</p><p>He replies, "Bertrand?"</p><p>I nod.</p><p>He scoffs his nose and eyebrows and shakes his head vigorously. "No."</p><p>I propose again: "Solomonoff induction. It's the theoretical limit of what AI turns into!" (I propose this as a concept that was developed in the 20th century, and that in the last 20 years or so has been applied to thinking about theoretical AI alignment.)</p><p>Him: "Maybe."</p><p>I tell him I want to turn moral philosophy into mathematics. Make it precise.</p><p>He replies, "That's been attempted. But you're right that mathematics is more fundamental than logic."</p><p>He gets into his Uber and rides off.</p><p>I wonder if either of us has learned anything.</p>]]></content:encoded></item><item><title><![CDATA[identity alignment ]]></title><description><![CDATA[Better mental framework than other ways of thinking about alignment?]]></description><link>https://blakeelias.substack.com/p/identity-alignment</link><guid isPermaLink="false">https://blakeelias.substack.com/p/identity-alignment</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Mon, 30 Mar 2026 20:24:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Better mental framework than other ways of thinking about alignment?</p><p>Any organism has an inside and an outside. Its &#8220;utility function&#8221; is derived as a function of its body-plan, and is computed with respect to internal states and external states.</p><p>What it might mean for two organisms to be &#8220;aligned&#8221; is that the component of their utility function having to do with the overlap in their external states is the same.</p><p>Notably, each organism&#8217;s internal states are a member of the other one's external state. So the pair of external states the two agents have is not the same, and thus can't be considered as a domain over which they can have the same preference. We can only compute the preference over a state set that's equal for both - so this has to be the intersection. The other choice would be to consider the whole world, ie including the inner states of each one, as well as the part that's external to both. But this is harder, since entities don't actually experience the whole world or have preferences about it - they really only have preferences over their boundary and whatever else they can see extending from there. So the only locus for establishing alignment is the place where their sensors and derived world-state overlap.</p><p>The challenge for alignment, then, might be definable as creating a pair of organisms whose perceived world states overlap the most. But what this amounts to is creating a pair of agents who have as close as possible to a single shared identity, rather than two separate identities. Eg. each individual neuron in your brain might have some amount of individual identity - but the stronger identity is at a higher layer of abstraction, ie the whole brain. Each one approximates a markov blanket to some degree, but the whole brain approximates it better. Whereas at the level of humans assembling into a tribe, we may find that each individual human seems closer to maintaining a markov blanket around itself than the entire tribe does - though this might depend on the tribe. So alignment might be the design of getting two organisms to be more like neurons in the brain than like humans in a tribe - shared identity stronger than individual identity. For humans and machines, this could perhaps look like Brain-computer interfaces, if done with the right objective functions&#8230;</p>]]></content:encoded></item><item><title><![CDATA[why do we need reward? ]]></title><description><![CDATA[Is Solomonoff induction enough?]]></description><link>https://blakeelias.substack.com/p/why-do-we-need-reward</link><guid isPermaLink="false">https://blakeelias.substack.com/p/why-do-we-need-reward</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Mon, 30 Mar 2026 20:01:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Is Solomonoff induction enough?</p><p>By the instrumental convergence hypothesis, it should be&#8230;.</p>]]></content:encoded></item><item><title><![CDATA[Quantifying the Unquantified]]></title><description><![CDATA[This seems like the real mission / place where original thinkers are needed.]]></description><link>https://blakeelias.substack.com/p/quantifying-the-unquantified</link><guid isPermaLink="false">https://blakeelias.substack.com/p/quantifying-the-unquantified</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Sun, 29 Mar 2026 01:46:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This seems like the real mission / place where original thinkers are needed.</p>]]></content:encoded></item><item><title><![CDATA[ There's No End to Work]]></title><description><![CDATA[Several researchers/engineers I&#8217;ve talked to at frontier labs have made peace with some version of the idea that this will be the last work they ever do.]]></description><link>https://blakeelias.substack.com/p/db7</link><guid isPermaLink="false">https://blakeelias.substack.com/p/db7</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Wed, 18 Mar 2026 06:56:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<ol><li><p>Several researchers/engineers I&#8217;ve talked to at frontier labs have made peace with some version of the idea that this will be the last work they ever do. One said, &#8220;I am fully aware that I will likely soon no longer be needed. I will happily work until that point, and then I will no longer be needed." As for what to do after? &#8220;I'll have plenty of time to figure that out.&#8221; Another said, &#8220;I will never do this same job again: it will certainly be replaced. We may come up with other to do - but it won't be this type of work.&#8221; But this has always been true in technology, or knowledge work for that matter: no-one ever does the <em>same</em> work twice. Assembly programmers weren't going to exist forever (at least not in large number) - it was clear at a certain point we'd mostly move to higher-level languages. The whole point is for technology to move forward, to a higher level of abstraction. This is the DRY principle (Don't Repeat Yourself) from programming. So, saying &#8220;I'm no longer needed&#8221; actually means &#8220;I'm no longer needed doing this exact job&#8221; -- and it's just like&#8230; yeah, trivially true. Perhaps to a different degree now, but still.</p></li></ol><ol start="2"><li><p>If so many people are complaining that we won't be happy with the situation that will come with AI&#8217;s impacts, then by definition this means there will still be work left, there will still be jobs. If we are not happy with how things end up, then there&#8217;s some job one can come up with, the likes of &#8220;fix X, Y, Z outcomes of AI that we're not happy with&#8221;. The outcome of &#8220;there are no jobs left&#8221; is then almost a contradiction in terms -- if we consider it a problem that there are no jobs left, that means it could be someone's job to fix it. This doesn't mean it will be an <em>easy</em> job, or even a <em>legible</em> job (&#8220;legible&#8221; meaning it's commissioned by some organization with funding, that you can send in a resume and poly for it, etc.) But that's already / always been the case -- the hardest problems we need solved don't have someone assigned as their job to solve them. That's why the problems are &#8220;wicked&#8221;. So perhaps the result of AI is that all the &#8220;job-shaped&#8221; problems get solved, and <em>all that's left </em>are wicked problems. Which can simultaneously be a blessing for those who want to work on wicked problems now already, but get distracted from doing so because of the allure and availability of the &#8220;stable job&#8221;&#8230;.</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Short cycles in cross disciplinary science ]]></title><description><![CDATA[https://www.chrisfieldsresearch.com/cycles4.pdf]]></description><link>https://blakeelias.substack.com/p/short-cycles-in-cross-disciplinary</link><guid isPermaLink="false">https://blakeelias.substack.com/p/short-cycles-in-cross-disciplinary</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Tue, 17 Mar 2026 18:02:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><a href="https://www.chrisfieldsresearch.com/cycles4.pdf">https://www.chrisfieldsresearch.com/cycles4.pdf</a></p><p>I thought this article was going to be about something else. Ie about the actual connections between the disciplines - not necessarily the author ships.</p><p>Though maybe this is the same thing?</p>]]></content:encoded></item><item><title><![CDATA[Alignment Alignment ]]></title><description><![CDATA[There is no concrete &#8220;alignment problem&#8221; that we agree on.]]></description><link>https://blakeelias.substack.com/p/alignment-alignment</link><guid isPermaLink="false">https://blakeelias.substack.com/p/alignment-alignment</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Tue, 17 Mar 2026 18:01:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There is no concrete &#8220;alignment problem&#8221; that we agree on.</p><p>There are many variants on it, and what we may be trying to do is decide/agree on which one is worth pursuing/asking.</p><p>Can make a list:</p><ul><li><p>&#8220;Machines that do what we ask them to do&#8221;</p></li><li><p>&#8220;Machines that do what we would have wanted, even if we don't directly ask&#8221;</p></li><li><p>&#8220;Machines that have our best interest at heart&#8221;</p></li><li><p>&#8230;Probably many others</p></li></ul><p>See also: <a href="https://www.alignmentforum.org/posts/67fNBeHrjdrZZNDDK/clarifying-alignment-vs-capabilities">https://www.alignmentforum.org/posts/67fNBeHrjdrZZNDDK/clarifying-alignment-vs-capabilities</a></p>]]></content:encoded></item><item><title><![CDATA[The Two Transitions ]]></title><description><![CDATA[Higher order organization: (individuals) cells &#8594; organisms &#8594; societies (language)&#8594; super organism (new language like extension)]]></description><link>https://blakeelias.substack.com/p/the-two-transitions</link><guid isPermaLink="false">https://blakeelias.substack.com/p/the-two-transitions</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Tue, 17 Mar 2026 17:50:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<ul><li><p>Higher order organization: (individuals) cells &#8594; organisms &#8594; societies (language)&#8594; super organism (new language like extension)</p><ul><li><p>We&#8217;ve moved from selfish-genes to selfish-memes.</p></li><li><p>But those memes replicate using humans as a substrate.</p></li><li><p>The mind viruses evolve &amp; compete, but still depend on us being alive.<br></p></li></ul></li><li><p>Substrate for life: carbon &#8594; silicon (regenerative robots)</p><ul><li><p>Silicon as a new fundamental replicator</p></li><li><p>Independent of RNA / DNA / Krebs Cycle</p></li></ul></li></ul><p>Do these have to go hand-in-hand? Does the second need to happen in order for the first to happen? Or can we have the first without the second?</p><p>One might think that the second one is necessary for the first. The symbiogenetic events we talk about that give rise to next-level complexity seem to often start (though not always&#8230;) from two existing autopoietic systems, which come to &#8220;make friends.&#8221; If this is the case, we might worry that this new reproducing entity, which doesn&#8217;t need us, might outcompete us.</p><p>Another way to phrase the question is: can allopoiesis (i.e. humans&#8217; constructed artifacts) in fact become autopoiesis? The &#8220;artificial&#8221; &#8212; comes from &#8220;artifice&#8221;, &#8220;artifact&#8221;, etc. &#8212; things we create as extensions of ourselves. Is not the opposite of &#8220;natural&#8221;. (Note, the Chinese word, &#8220;human-created intelligence&#8221;, is a more directly descriptive term that makes the meaning a bit clearer. Since &#8220;artificial&#8221; really does just mean &#8220;human-created&#8221;.)</p><p>So, the rise of these &#8220;artifacts&#8221; we&#8217;re making, trying to have them become a life form of their own, in what I call &#8220;the oblivion of artifacts&#8221;, may not be necessary and may not even work. It may be that it just stays as artifacts, and they indeed reproduce at that level but are dependent on the earlier levels underneath. The same thing happens with viruses: the transposons etc.</p><p>I&#8217;m stuck in general on the nature between autopoiesis and allopoesis. The thesis here is that LLMs / AI will be a new <em>allopoetic</em> layer, as opposed to silicon becoming a new <em>autopoietic</em> substrate. It&#8217;s unclear to me whether we can actually distinguish autopoiesis and allopoiesis as separate concepts (is the original autopoietic being itself, combined with its allopoetic artifacts, not itself a larger autopoietic being)? The question is, can allopoetic artifacts ever become autopoietic substrates independent of the host? Under what conditions? If the answer is &#8220;usually not&#8221;, then this is one reason to have low <em>P(doom)</em>.</p><p>That said, I&#8217;m still trying to square this with Mike Levin&#8217;s perspective that we <em>will</em> in fact have <a href="https://www.sciencedirect.com/science/article/pii/S0303264723001399">autopoietic technology</a>.</p><p>I wonder if some experiments with BFF might be able to help us understand this. E.g. try shorter vs. longer replicators as seeds, see how these interact - whether a longer base replicator is at a disadvantage compared to a shorter, tighter one.</p><p>In the end, what we&#8217;ve created is capital. That&#8217;s the ultimate allopoetic extension. And it has a mind of its own. We in fact can&#8217;t understand the world with our tiny minds &#8212; but the AI meme machine can, and capital can &#8212; in a certain way. These things are new feedback loops. And they&#8217;ll need to be regulated, just as the dopamine loop in humans needs to be regulated by yet other feedback loops.</p><p>Where really is the line between autopoiesis and allopoesis? They&#8217;re both types of poiesis (bringing forth / creating), the only difference being whether what&#8217;s created ends up being a replica of the creator or not. But several edge cases ensue:</p><ul><li><p>What if a bunch of creators, all having the same body plan, create a bunch of random parts that mirror just part of themselves. Individually, each creator is not creating a full copy of itself. So allopoesis at a lower level (the individual creating a tool) turns into autopoiesis at a higher level (the colony creating a copy of itself).</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Thoughts on safety ]]></title><description><![CDATA[Safety not being a model property makes the most sense to me.]]></description><link>https://blakeelias.substack.com/p/thoughts-on-safety</link><guid isPermaLink="false">https://blakeelias.substack.com/p/thoughts-on-safety</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Mon, 16 Mar 2026 00:41:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Safety not being a model property makes the most sense to me. Sometimes safety is a model property (eg for sycophantic usage, etc), but a lot of the time it isn't. General civilizational safety things are important too -- cyber security etc.</p><p>It seems important to have surveillance for bad behavior. And it doesn't have to be centralized / authoritarian like China - one could build a version of this thats more decentralized/ guaranteed to only use information for legitimate purposes as decided by the people.</p><div><hr></div><p>Utility theory is broken.</p><p>What can it even mean for two agents to have the same utility?</p><p>&#8220;ASI suicide as a safety feature&#8221; doesn't make sense unless you believe humans would do the same.</p><p>See LessWrong post on why we should give up on utility theory.</p>]]></content:encoded></item><item><title><![CDATA[Two Unlikely Worlds]]></title><description><![CDATA[AI discourse makes it seem as if we're going to end up in one of two possible worlds:]]></description><link>https://blakeelias.substack.com/p/two-unlikely-worlds-9ad</link><guid isPermaLink="false">https://blakeelias.substack.com/p/two-unlikely-worlds-9ad</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Mon, 16 Mar 2026 00:37:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI discourse makes it seem as if we're going to end up in one of two possible worlds:</p><p>1) utopia: complete abundance without work. AI and robotics do all the work for us, we benefit. We perhaps question our meaning without work, but we pursue other things in the search for meaning</p><p>2) disaster: AI puts everyone out of work; only the rich who own the capital benefit. Everyone else enters a permanent underclass.</p><p>I think #2 is avoidable because AI is affordable to many people - you can run local models on a cheap consumer hardware, purchase a humanoid robot, etc. So it's not just the wealthy who are able to use this tech to augment their own lives, be more creative and effective.</p><p>But #1 doesn't seem like quite what will happen either. The AI only does things we direct it towards in some way. We still have to prompt it, give it tons of guidelines etc. Even as it gets better and needs less supervision, it still needs that initial &#8220;kick-off&#8221; , which has to come from us.</p><p>The AIs of today aren't fully autonomous, and in a sense they couldn't be. If they were fully autonomous, ie didn't need any input from us, then they would go off doing things that don't necessarily benefit us. If we want benefit from them then they somehow need to be conditioned on what will be good for us - which, right now, comes from us telling them directly. At some point we might have intelligence products that anticipate your needs and &#8220;just do things&#8221; for you. But that's not the mode of product we have now. And it's not clear whether something like that could ever exist: if an AI is fully autonomous like that, it's not clear why it should or would care about you at all, rather than fare just about itself. It may be that in order to get benefit from these beings, we will always need to be engaging with them, guiding them, communicating what we want. Which at least for the foreseeable future, looks like some kind of work (or maybe a blend between work and play -- and I think we can expect the play element to grow moreso, as people have more and more fun just getting to vibe code whatever they want, with their imagination being the limit).</p><p>If this world of us having to continue interacting with the AI to get what we want from it -- needing to continue to work and play with them to get benefit (kind of like how it is with dogs!), then we don't end up in the situation of #1 where there's a complete utopia of leisure, and/or a meaning crisis due to idleness, nor do we end up in situation #2 where the economy has totally collapsed, there's no work for anyone and no way to generate wealth. Instead it's some third scenario, where we just have very good new type of labor/companion/service animal that we get to interact with.</p><p>The &#8220;replacement&#8221; narrative/worry never really comes true; all of this sits in the &#8220;augmentation&#8221; paradigm. We may replace / make obsolete particular job roles, because that entire job is do-able by anyone using an AI tool with the click of a button. But we don't replace the essence of human work altogether, nor do we replace humans themselves. As the saying goes: &#8220;AI won't take your job -- a person using AI will take your job&#8221;. And that person using AI is themself just doing yet a different job -- one at a different level of abstraction.</p><p>The vision of full automation seems more plausible to me in robotics than in knowledge work. Knowledge work is getting transformed first, but I think what we'll see is a re-organization of what knowledge work looks like - in particular where specialized roles go away and are replaced with generalist, cross-cutting ones - but humans still need to be in the loop continuously to specify what needs to be done. This will always be necessary in the realm of knowledge work because the AIs exist purely as informational beings, while humans are the one with a physical embodied existence tomhat provides the grounding for what needs to happen vs not.</p><p>Robotics has taken longer for the automation/AI wave to hit. But when it does happen, there may be at least some tasks there that can be automated more completely in a way where &#8220;what needs to be done&#8221; is determined purely by the physics of the problem, i.e. doesn't fundamentally need to rely on continuous human input to get a meaningfully specified task.</p><p>Consider a fruit tree. The tree &#127796; can be seen as a form of automation, that turns water and sunlight into fruit (really, takes water, sunlight, carbon dioxide and soil as inputs, and as an output both grows itself <em>and</em> produces fruit.) Robotic automation could be employed to monitor the tree for newly ripe fruit, pick the fruit, wash it, and perhaps peel/slice/stem/de-seed the fruit to ready it for consumption. If one wants to a fruit-bearing home garden, they can purchase not just a tree sapling (from a nursery), but also the harvesting robot, which functions almost as an extension of the tree. The task of the harvesting robot is fully specified by the physical reality of the problem: it can carry out that one set of actions/capabilities, and its output (ripe, peeled fruit) would be equally usable by any organism that can digest that fruit: humans, raccoons, insects, etc. It's not necessary for the consumer organism to specify any unique preference (eg &#8220;prompting&#8221;) to make the machine&#8217;s output useful. On the other hand, such preference specification typically <em>is</em> needed with knowledge work to make the output useful, hence the observation that every knowledge-work task which gets automated simply creates a new task that a human still needs to do.</p><p>Physical robots could conceivably be &#8220;hard coded&#8221; to perform a key set of functions that create a baseline of survival support for any animal: continuous cultivation of crops, construction and maintenance of shelter, producing cloth for garments, etc. The remaining work would be repairing the robots themselves when their components break. This could be a job for the same robot or another robot to perform: among all the robots that would be involved doing the various homesteading tasks, another robot would be the &#8220;robot doctor&#8221;, helping repair all the other robots when they break. The end result would be two species co-existing: human-kind and robot-kind. Members of each species would take care of their own, where &#8220;care&#8221; includes reproduction to make new members of one's own kind (and replace departed members), maintenance and healing of those members who get damaged, acquisition of resources to enable such creation and maintenance, etc. But they would also ideally co-exist in a form of mutual support. We may expect or desire that the robots will do work which we benefit from, via the tasks mentioned above. But on the converse, is not as obvious what value humans would provide for the robots, when our instructions or preference expression are not as necessary. (Perhaps humans would help as a backup option for repair tasks that the robots themselves are incapable of or inefficient at executing -- this is the case with current robotics systems and may continue for a long foreseeable future until robots&#8217; dexterity and fine motor skills vastly improve).</p>]]></content:encoded></item><item><title><![CDATA[Two Unlikely Worlds]]></title><description><![CDATA[AI discourse makes it seem as if we're going to end up in one of two possible worlds:]]></description><link>https://blakeelias.substack.com/p/two-unlikely-worlds</link><guid isPermaLink="false">https://blakeelias.substack.com/p/two-unlikely-worlds</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Sat, 14 Mar 2026 10:06:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI discourse makes it seem as if we're going to end up in one of two possible worlds:</p><p>1) utopia: complete abundance without work. AI and robotics do all the work for us, we benefit. We perhaps question our meaning without work, but we pursue other things in the search for meaning</p><p>2) disaster: AI puts everyone out of work; only the rich who own the capital benefit. Everyone else enters a permanent underclass.</p><p>I think #2 is avoidable because AI is affordable to many people - you can run local models on a cheap consumer hardware, purchase a humanoid robot, etc. So it's not just the wealthy who are able to use this tech to augment their own lives, be more creative and effective.</p><p>But #1 doesn't seem like quite what will happen either. The AI only does things we direct it towards in some way. We still have to prompt it, give it tons of guidelines etc. Even as it gets better and needs less supervision, it still needs that initial &#8220;kick-off&#8221; , which has to come from us.</p><p>The AIs of today aren't fully autonomous, and in a sense they couldn't be. If they were fully autonomous, ie didn't need any input from us, then they would go off doing things that don't necessarily benefit us. If we want benefit from them then they somehow need to be conditioned on what will be good for us - which, right now, comes from us telling them directly. At some point we might have intelligence products that anticipate your needs and &#8220;just do things&#8221; for you. But that's not the mode of product we have now. And it's not clear whether something like that could ever exist: if an AI is fully autonomous like that, it's not clear why it should or would care about you at all, rather than fare just about itself. It may be that in order to get benefit from these beings, we will always need to be engaging with them, guiding them, communicating what we want. Which at least for the foreseeable future, looks like some kind of work (or maybe a blend between work and play -- and I think we can expect the play element to grow moreso, as people have more and more fun just getting to vibe code whatever they want, with their imagination being the limit).</p><p>If this world of us having to continue interacting with the AI to get what we want from it -- needing to continue to work and play with them to get benefit (kind of like how it is with dogs!), then we don't end up in the situation of #1 where there's a complete utopia of leisure, and/or a meaning crisis due to idleness, nor do we end up in situation #2 where the economy has totally collapsed, there's no work for anyone and no way to generate wealth. Instead it's some third scenario, where we just have very good new type of labor/companion/service animal that we get to interact with.</p><p>The &#8220;replacement&#8221; narrative/worry never really comes true; all of this sits in the &#8220;augmentation&#8221; paradigm. We may replace / make obsolete particular job roles, because that entire job is do-able by anyone using an AI tool with the click of a button. But we don't replace the essence of human work altogether, nor do we replace humans themselves. As the saying goes: &#8220;AI won't take your job -- a person using AI will take your job&#8221;. And that person using AI is themself just doing yet a different job -- one at a different level of abstraction.</p><p>The vision of full automation seems more plausible to me in robotics than in knowledge work. Knowledge work is getting transformed first, but I think what we'll see is a re-organization of what knowledge work looks like - in particular where specialized roles go away and are replaced with generalist, cross-cutting ones - but humans still need to be in the loop continuously to specify what needs to be done. This will always be necessary in the realm of knowledge work because the AIs exist purely as informational beings, while humans are the one with a physical embodied existence tomhat provides the grounding for what needs to happen vs not.</p><p>Robotics has taken longer for the automation/AI wave to hit. But when it does happen, there may be at least some tasks there that can be automated more completely in a way where &#8220;what needs to be done&#8221; is determined purely by the physics of the problem, i.e. doesn't fundamentally need to rely on continuous human input to get a meaningfully specified task.</p><p>Consider a fruit tree. The tree &#127796; can be seen as a form of automation, that turns water and sunlight into fruit (really, takes water, sunlight, carbon dioxide and soil as inputs, and as an output both grows itself <em>and</em> produces fruit.) Robotic automation could be employed to monitor the tree for newly ripe fruit, pick the fruit, wash it, and perhaps peel/slice/stem/de-seed the fruit to ready it for consumption. If one wants to a fruit-bearing home garden, they can purchase not just a tree sapling (from a nursery), but also the harvesting robot, which functions almost as an extension of the tree. The task of the harvesting robot is fully specified by the physical reality of the problem: it can carry out that one set of actions/capabilities, and its output (ripe, peeled fruit) would be equally usable by any organism that can digest that fruit: humans, raccoons, insects, etc. It's not necessary for the consumer organism to specify any unique preference (eg &#8220;prompting&#8221;) to make the machine&#8217;s output useful. On the other hand, such preference specification typically <em>is</em> needed with knowledge work to make the output useful, hence the observation that every knowledge-work task which gets automated simply creates a new task that a human still needs to do.</p><p>Physical robots could conceivably be &#8220;hard coded&#8221; to perform a key set of functions that create a baseline of survival support for any animal: continuous cultivation of crops, construction and maintenance of shelter, producing cloth for garments, etc. The remaining work would be repairing the robots themselves when their components break. This could be a job for the same robot or another robot to perform: among all the robots that would be involved doing the various homesteading tasks, another robot would be the &#8220;robot doctor&#8221;, helping repair all the other robots when they break. The end result would be two species co-existing: human-kind and robot-kind. Members of each species would take care of their own, where &#8220;care&#8221; includes reproduction to make new members of one's own kind (and replace departed members), maintenance and healing of those members who get damaged, acquisition of resources to enable such creation and maintenance, etc. But they would also ideally co-exist in a form of mutual support. We may expect or desire that the robots will do work which we benefit from, via the tasks mentioned above. But on the converse, is not as obvious what value humans would provide for the robots, when our instructions or preference expression are not as necessary. (Perhaps humans would help as a backup option for repair tasks that the robots themselves are incapable of or inefficient at executing -- this is the case with current robotics systems and may continue for a long foreseeable future until robots&#8217; dexterity and fine motor skills vastly improve).</p>]]></content:encoded></item><item><title><![CDATA[A Biological Take on the Chinese Room]]></title><description><![CDATA[See here for Background on the Chinese Room Experiment.]]></description><link>https://blakeelias.substack.com/p/a-biological-take-on-the-chinese</link><guid isPermaLink="false">https://blakeelias.substack.com/p/a-biological-take-on-the-chinese</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Sat, 14 Mar 2026 07:58:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>See here for <a href="https://plato.stanford.edu/entries/chinese-room/#:~:text=Hence%20the%20%E2%80%9CTuring%20Test%E2%80%9D%20is%20inadequate.%20Searle,fact%20that%20computers%20merely%20use%20syntactic%20rules">Background on the Chinese Room Experiment</a>.</p><p>The experiment points us at the question of how we want to define the word &#8220;understanding&#8221;. Why is it that we seem comfortable saying that a person understands something, but with a set of rules (that a person or machine can blindly follow), we don't say that the rules themselves possess understanding?</p><p>It might be tempting at first to say that the substrate is the differentiator: the carbon based beings can understand things but the silicon ones do not. But this seems a bit arbitrary. Further, the thought experiment indeed shows a case where a <em>human</em> is blindly following the rules that are written on paper. So the substrate of the entity carrying out the rule execution must not be the important thing after all.</p><p>Instead, what seems important is the rules themselves, and more importantly, where they came from, and how they adapt (or not).</p><p>I think a more intuitive definition of &#8220;understanding&#8221; that fits common usage of the term, is when an organism possesses a world-model that lets it take action to survive better. On this view, we&#8217;d say that non-living never have understanding. The definition itself resolves the thought experiment.</p><p>When an organism learns a world model, it learns it because it's useful to it, and it's undergoing an evolutionary process that's selecting for organisms that learn useful, advantageous world models.</p><p>Whereas if a machine just has some statically defined behavior policy, that doesn't mean that policy is a <em>good</em> one for surviving. There's an aspect of competence that we can't say it has.</p><p>So in the thought experiment, the set of rules provided in the boom may process Chinese at that moment. An agent might even be able to use those rules get by in certain situations. But is it continually learning new Chinese expressions &amp; cultural progress that are being added to the language? Does it have goals that it's setting for itself, and is it learning to use language in such a way that helps it achieve its goals and survive? If so, then I would say the machine <em>does </em>understand Chinese. But if it's just a static policy / rule book, as in the original framing of the experiment, then I would say no, it isn't true understanding in the sense I'm defining the word.</p><p>Substrate material itself doesn't matter - an organism can be built out of carbon or out of silicon. What matters is the type of connection and embodiment with the real world.</p><p>___</p><p>(A note on what it means to be living:</p><p>We can ask whether tools, which may be essentially helpful for living things, can themselves said to be living. I will argue that they should the. To be a living thing means that thing needs to be competent on your own. If an object gets used by living things to aid with life, but left to its own devices it doesn't do much of anything or even really try, then it's non-living. So the paper/book that the rules are written on, that really doesn't do anything on its own.)</p>]]></content:encoded></item><item><title><![CDATA[One Virtue vs. Many]]></title><description><![CDATA[I. An NYT opinion piece this week connected a recent AI research finding to an ancient debate in moral philosophy (something which seems to happen every other day lately&#8230; but alas).]]></description><link>https://blakeelias.substack.com/p/one-virtue-vs-many</link><guid isPermaLink="false">https://blakeelias.substack.com/p/one-virtue-vs-many</guid><dc:creator><![CDATA[Blake Elias]]></dc:creator><pubDate>Fri, 13 Mar 2026 10:02:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DkT_!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b82810b-8708-4b44-9f60-835e2b0153ef_608x608.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>I.</strong></p><p>An <a href="https://www.nytimes.com/2026/03/10/opinion/ai-chatbots-virtue-vice.html?smid=nytcore-android-share">NYT opinion piece</a> this week connected a recent AI research finding to an ancient debate in moral philosophy (something which seems to happen every other day lately&#8230; but alas).</p><p>The AI result was one where training LLMs to produce insecure/attackable code also makes them more likely to do other &#8220;bad&#8221; things, like make violent or sexist remarks, which is a somewhat surprising finding. The author connects this to an ancient debate as to whether all virtues are connected (i.e. whether there is really is just one virtue):</p><blockquote><p>Plato argued that all the various human virtues are really one thing, knowledge of the good. Aristotle softened this view a bit but still insisted the virtues are in practice so tightly woven that you can&#8217;t really have one without the others. (A soldier who fights fiercely for fear of disgrace rather than from nobility and knowledge of what&#8217;s worth defending is, for Aristotle, only apparently brave &#8212; and probably only apparently virtuous in other parts of his life, as well.) The Stoics, too, held that the virtues were inseparable: You possess them all or none. Augustine and Aquinas carried this view into Catholic thought.</p><p>In philosophy, this family of moral views fell out of favor several hundred years ago, replaced by approaches like deontology, which emphasizes the following of rules, or consequentialism, which seeks to maximize good outcomes. With character no longer at the center of moral thinking, what you might call a more compartmentalized understanding of human nature took hold. The ancients had been wrong. People could be good and bad in about as many mixtures as could be counted.</p><p>But the debate was never settled. During the second half of the 20th century, philosophers began to explore virtue ethics again, led by a group of British scholars reacting in part to what they saw as the inability of the dominant ethics of the time to deal with the horrors of World War II.</p></blockquote><p><strong>II.</strong></p><p>I&#8217;m reminded of one of my <a href="https://worrydream.com/Bio2011/">favorite quotes</a> from Bret Victor:</p><blockquote><p>&#8220;There is no &#8216;Technology&#8217;. There is no &#8216;Design&#8217;. There is only a vision of how mankind should be, and the relentless resolve to make it so. The rest is details.&#8221;</p></blockquote><p>It echoes Plato &#8212; and in my experience practicing in these fields (technology, design, science, etc.), it indeed feels true. There&#8217;s one unifying force of being life-supporting, and you know when you&#8217;re in it. You&#8217;re either in an integrated mode, a feedback loop that&#8217;s life-supporting overall, or you&#8217;re not. There&#8217;s a mode where the next action is clear, where Everything Makes Sense. And there&#8217;s a mode where one thing you do is disconnected from the next: there&#8217;s no story, there&#8217;s no thread.</p><p>Now it&#8217;s also true that when you&#8217;re in that latter mode &#8212; i.e. when <a href="https://open.substack.com/pub/nothinghuman/p/whole-activities?r=84yrx&amp;selection=fe07f8c2-7bfc-4916-9fc8-708c738ea043&amp;utm_campaign=post-share-selection&amp;utm_medium=web&amp;aspectRatio=instagram&amp;textColor=%23ffffff&amp;bgImage=true">modernity decouples</a> &#8212; you can still be doing some things well (or even most things well). You have some feedback loops that are finding ways to be effective, each in their own dimension. But you&#8217;re not doing &#8220;<a href="https://open.substack.com/pub/nothinghuman/p/whole-activities?r=84yrx&amp;selection=224acd43-25c7-4468-b72c-ce8a392d130e&amp;utm_campaign=post-share-selection&amp;utm_medium=web&amp;aspectRatio=instagram&amp;textColor=%23ffffff&amp;bgImage=true">whole activities</a>&#8221; anymore.</p><p>It&#8217;s not surprising that once you get fractured in that way, it&#8217;s not too hard for any one of the feedback loops (i.e. any one of the virtues) to get broken. And once a couple are broken, it&#8217;s not too hard for them all to get broken.<a href="#footnote-1">1</a></p><p><strong>III.</strong></p><p>Perhaps, then, we don&#8217;t have to &#8220;settle the debate&#8221; on .whether (human) virtue is fundamentally one thing or is fundamentally plural. Rather than an &#8220;either-or&#8221;, perhaps it&#8217;s a &#8220;both-and&#8221; situation: there&#8217;s some way in which many types of virtue can exist independently (i.e. you have some but not others), <em>and</em> they&#8217;re stronger when you possess them all together, weaker when de-coupled. There are individual virtues which can be had independently, <em>and</em> there&#8217;s a &#8220;meta-virtue&#8221; of having them all at once: a way in which they exist in mutual support / bolster one another, and a way of thinking and self-understanding that&#8217;s more likely to bring all of them online at once.</p><p>Perhaps this dual relationship can be formalized into math / physics terms. Mike Levin explains that all living systems exist as competent wholes built out of competent parts: a cell on its own is competent; a group of competent cells composed into an organism is competent; a group of organisms composed into a tribe is competent, etc. Perhaps we can replace &#8220;competent&#8221; with &#8220;formidable&#8221;: a cell is formidable; a group of cells composed into an organism is formidable; a group of organisms composed into a tribe is formidable. Now replace &#8220;formidable&#8221; with &#8220;virtuous&#8221;: same thing. I suspect that when we get down to it, the concept of competence and the concept of virtue are the same thing: they&#8217;re both a property of a circuit with some aspect of being &#8220;life-supporting&#8221;.<a href="#footnote-2">2</a> In both cases, we have the property that one such circuit can achieve this goal on its own, but the goal can be achieved better in partnership with other, similar circuits running in parallel. Without the collection of them running in parallel, just one of them on its own is still viable (e.g. compartmentalized, separate virtues) &#8212; but isn&#8217;t really the same thing as the whole existing as one (i.e. single, unified virtue).</p><p>So competence and virtue both exist at multiple scales (from a virtuous/competent cell, to a virtuous/competent set of cells making up an organ, to a virtuous/competent set of organs making up a body, to a virtuous/competent set of bodies making up a society). And they both experience &#8220;top-down&#8221; causality (i.e. the overall system functioning can cause the components to function better) in addition to typical &#8220;bottom-up&#8221; causality (i.e. individual components functioning well cause the entire system to function well).</p><p><strong>IV.</strong></p><p>Experiments on LLMs might teach us some things about this, as can studies on humans. But these are all just one-off examples. What we really need is a better understanding of the emergence of autopoietic / competent systems in general, better <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC3262299/">mathematical language to talk about top-down causality</a>. Explanations of how competence begets more competence, and how breakdown begets more breakdown.</p><p><strong>IV.</strong></p><p>Perhaps we will continue to revisit ancient philosophy, remember that it posed some important questions and hypotheses which seem highly counter-intuitive today (became obscured through modernity), and that are valuable to re-examine. Yet as we do this we will also realize also why they got left behind.</p><p>We will see that they were incomplete perspectives which could not defend themselves against tempting alternatives. Even the most enlightened, virtuous thinkers of that generation could not convince themselves or each other, let alone the rest of the population, that these views were still the right ones or best ones to have, rather than switching to new ones.</p><p>The Reformation came, the Enlightenment/Rationalism/Modernism came, and gave us Many Nice Things. We never proved the old worldviews wrong in full, as we had some new stuff that seemed kind of right (though we proved some small parts of the old worldview wrong, and other parts we just forgot about). But we also never proved the old stuff right and the new stuff wrong and the either.</p><p>Perhaps we will see that these philosophies have just been living in tension with unresolved contradictions, because all of them were incomplete. They never fully had a conversation with each other on terms that both sides could understand. If we would put these ideas onto a shared formal foundation, we&#8217;d find that they&#8217;re not incompatible &#8212; they&#8217;re just getting at different, incomplete aspects of the truth. That neither worldview was complete or true on its own. There&#8217;s no &#8220;finding which side was right&#8221;. Instead, we&#8217;ll construct a shared foundation that qualifies what these views mean in relation to one other, and under what <em>circumstances</em> each one is right &#8212; and that <em>foundation</em> itself will be the deeper truth &amp; resolution we seek.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="#footnote-anchor-1">1</a></p><p>Similarly with cancer: once cellular communication breaks down in the body overall, it&#8217;s not too hard for a few cells to lose their connection/identity as part of a bigger whole &#8212; a situation which we call cancer. (This happens all the time, of course, at some small background rate that your immune system is usually able to catch. But once this crosses a threshold that the body / immune system can&#8217;t self-correct, things easily run downhill.)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><a href="#footnote-anchor-2">2</a></p><p>Maturana and Varela use the term &#8220;<a href="https://en.wikipedia.org/wiki/Autopoiesis">autopoietic</a>&#8221;.</p></div></div>]]></content:encoded></item></channel></rss>