<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[ Orí Intelligence]]></title><description><![CDATA[AI at work, with people in charge.]]></description><link>https://read.oriintelligence.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!D7kH!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01184bc3-ce58-42d0-94f6-298234140319_512x512.png</url><title> Orí Intelligence</title><link>https://read.oriintelligence.ai</link></image><generator>Substack</generator><lastBuildDate>Sun, 02 Aug 2026 10:15:01 GMT</lastBuildDate><atom:link href="https://read.oriintelligence.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Temi Obe]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[oriintelligence@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[oriintelligence@substack.com]]></itunes:email><itunes:name><![CDATA[Temi Obe]]></itunes:name></itunes:owner><itunes:author><![CDATA[Temi Obe]]></itunes:author><googleplay:owner><![CDATA[oriintelligence@substack.com]]></googleplay:owner><googleplay:email><![CDATA[oriintelligence@substack.com]]></googleplay:email><googleplay:author><![CDATA[Temi Obe]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Delegation Map]]></title><description><![CDATA[What to hand off, what needs your sign-off, and what stays yours.]]></description><link>https://read.oriintelligence.ai/p/the-delegation-map</link><guid isPermaLink="false">https://read.oriintelligence.ai/p/the-delegation-map</guid><dc:creator><![CDATA[Temi Obe]]></dc:creator><pubDate>Thu, 30 Jul 2026 14:07:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7piA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><strong>An Or&#237; playbook. Built from my own operation, free to use</strong></em></p><p>I once authorized an AI build with the format living unstated in my head. The go-word was clear, the intent was not, and the deliverable came back built to a spec the system had assumed on my behalf. It was rebuilt after the fact, which means it was built twice, and the second build only existed because I skipped one sentence describing what I actually wanted.</p><p>That failure taught me the thing I open every executive session with. The failure mode is rarely delegating too much. It is delegating vaguely. Organizations pay for the same skipped sentence at scale: AI work built to an assumed spec, with the mismatch surfacing after the work is done and nobody on record as having decided the spec.</p><p>So before AI touches a piece of my work, it gets sorted. Three buckets, three sorting questions, and the sort takes less time than one round of rework. This is the same map I build with executives in the room, on their real workload. Here is how it runs on mine.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7piA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7piA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 424w, https://substackcdn.com/image/fetch/$s_!7piA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 848w, https://substackcdn.com/image/fetch/$s_!7piA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!7piA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7piA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png" width="1200" height="1350" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1350,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!7piA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 424w, https://substackcdn.com/image/fetch/$s_!7piA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 848w, https://substackcdn.com/image/fetch/$s_!7piA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 1272w, https://substackcdn.com/image/fetch/$s_!7piA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6c064fa-0038-4674-8064-b9232500fec3_1200x1350.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Bucket one: hand off</h2><p>The sorting question: <strong>if this is wrong, is the error cheap and obvious?</strong></p><p>Work goes here when a mistake costs minutes, not credibility, and when I would catch the mistake on sight without expertise or effort. First-pass research lives here. Transcript cleanup lives here, because I dictate first and clean up second, and a mangled sentence announces itself. Inventorying the checkable claims in a draft lives here too. Pulling the list of claims is mechanical work. Judging them takes expertise, and that happens two buckets up.</p><p>Obvious is the half people skip. An error is obvious when it announces itself in the reading: the mangled sentence, the wrong date in a timeline you know, the summary that contradicts the meeting you sat in. An error is not obvious when it looks fine and happens to be wrong: the plausible statistic, the citation that formats perfectly and does not exist. Those errors are sometimes cheap. They are never obvious, so their work does not belong here.</p><p>Notice what does not qualify: anything where the error would be expensive, invisible, or carrying my name. The error has to be both cheap and obvious, and one without the other moves the work up a bucket.</p><p>The organizational version of this bucket holds meeting summaries, first-pass vendor scans, and internal research digests. Written as team rules, the sorted decision reads like this:</p><ul><li><p><em>AI-generated meeting notes circulate without review; whoever spots an error fixes it in the doc.</em></p></li><li><p><em>First-pass vendor research goes straight to the shared channel, labeled unverified.</em></p></li></ul><p>That unverified label marks the bucket, so nobody downstream mistakes a hand-off for a signed-off fact. And the same wrong line that costs a correction here costs trust in a customer email, which is why the customer email lives in the next bucket.</p><h2>Bucket two: needs your sign-off</h2><p>The sorting question: <strong>does this carry my name or spend my resources?</strong></p><p>Most of my working week lives here, and this bucket is the reason the map exists at all. Sign-off is not a vibe. It is a written checkpoint, and it runs at both ends of the handoff. At entry, the work gets specified before it starts: what will be built, in what format, for what use. At exit, the output passes a standard before it moves. Mine are standing rules: state the spec in one or two lines before any build, verify checkable claims before asserting them, distinguish documented from contested, apply the voice standard, flag what remains uncertain.</p><p>That rebuilt deliverable in my opening failed at entry. I authorized the work and skipped the format sentence, so the exit review had nothing to catch, because the output matched the spec the system assumed.</p><p>A checkpoint earned its keep in one of my own newsletter drafts. Every issue I write runs through a final claims check before it publishes: each checkable assertion in the draft gets verified against its source. In one draft I had quoted a line attributed to Adam Smith because it gave me exactly the argument I wanted about precolonial African societies. The verification pass traced it to The Wealth of Nations and found a modern paraphrase of a passage that said something different in openly derogatory language. The line came out, the paragraph was worse for it, and the issue shipped without it, because the checkpoint was written down before I was attached to the sentence.</p><p>In practice that means writing two things down before the machine starts: the spec for what you want built, and the standard the output has to pass before it moves. Write both while you are calm, because you will be applying them when you are not.</p><p>Organizations already run this bucket for human work. &#8220;We do not ship without sign-off&#8221; survives every deadline for a reason. The AI version is the same gate with the standard made explicit. Written down, the sorted decision reads like this:</p><ul><li><p><em>No AI-assisted work starts until the requester states, in writing, what will be built, its format, and where it will be used.</em></p></li><li><p><em>No code from the AI automation pipeline deploys without review by an engineer who owns the outcome.</em></p></li><li><p><em>No customer-facing copy sends until a named owner has checked every claim against what the product actually does.</em></p></li><li><p><em>AI-drafted analysis can feed a pricing recommendation only if the memo states which numbers were verified and by whom.</em></p></li></ul><p>If your team cannot say who signs and against what, the work is not in this bucket. It is unsorted, which in practice means it is in bucket one by default and nobody chose that.</p><h2>Bucket three: stays yours</h2><p>The sorting question: <strong>is the judgment the work?</strong></p><p>Some tasks produce an output, and the judgment rides along. Other tasks are the judgment. Which product ships. What gets published under my name at all. Whether a client engagement is a fit. The premise underneath a plan. An organization sorts the same way: layoffs, pricing, exiting a market, and ending a vendor contract all held in this bucket, no matter how strong the model&#8217;s recommendation looks. AI can assemble the evidence for every one of these, list what supports the recommendation and what would require specialist review. It does not get the decision, because the accountability cannot go with it, and I write that boundary down so it holds when I am tired.</p><p>Written down, the sorted decision reads like this:</p><ul><li><p><em>No workforce decision is made or announced on an AI recommendation alone; the accountable executive decides and signs.</em></p></li><li><p><em>No claim denial or coverage call is decided by a model; the system assembles the file, and a person decides it.</em></p></li><li><p><em>The board minutes for a market exit record who decided, and it is never a tool.</em></p></li></ul><p>I learned where this line sits by nearly crossing it. During a company launch, I proposed a fifth product for an audience I had already deprioritized, and I brought it to the system as a planning task. Planning is bucket two work. The decision to add a product mid-launch was bucket three work, and I had smuggled it in wearing a task&#8217;s clothes. My own standing rule caught it, flagged the overcommitment pattern I had declared, and handed the decision back to me. I postponed the product. The system did not decide that. It refused to let me pretend I already had.</p><p>That is the real function of the third bucket. It is not a fence against the machine. It is a fence against your own tendency to delegate a decision by dressing it as a deliverable.</p><h2>Running the sort</h2><p>Take your recurring work, the tasks that fill an actual week, and run each one through the questions in order. Is the error cheap and obvious? Hand it off. Does it carry your name or spend your resources? Define the checkpoint in writing, then hand off the work and keep the judgment. Is the judgment the work? It stays yours, and you write one sentence saying so.</p><p>Run it on a Monday that looks like most executive Mondays. The notes from Friday&#8217;s leadership meeting: bucket one, circulated with the unverified label, corrected by whoever spots the gap. The board pre-read summarizing the quarter: bucket two, because it moves under your name and a wrong number in it is neither cheap nor obvious. The first draft of the all-hands narrative: bucket two, same gate, named owner. Inbox triage and scheduling: bucket one. And the reorg question sitting underneath that narrative: bucket three. The system can pull together the current org chart and every input the reorg needs. Next quarter&#8217;s structure is not its call.</p><p>Two things I tell every executive who builds this map with me. First, the sort is per task, not per tool. The same AI system sits in all three buckets doing different jobs. Second, the buckets move. Work migrates down as your checkpoints prove themselves and up when the stakes change. The map is a living document with a review date, not a laminated poster.</p><p>Movement looks like this on my own map. Inventorying the checkable claims in a draft began as sign-off work: every claim pulled and judged by me. Once the written fact-check procedure held up across several published issues, the pulling migrated down to bucket one. The judging never moved. The mechanical half of a task declines after its checkpoint has a track record, while the judgment half stays exactly where it was.</p><p>Teams run the same sort on their recurring workflows, with one addition: the finished map gets written down and shared. An unshared map fails quietly. Six months later the team is maintaining AI-built deliverables that nobody remembers approving, and there is no record to check who was supposed to.</p><p>The last question every executive asks me is where these rules physically live. They do not live in your prompts. They go into the system&#8217;s standing configuration, the custom instructions or project setup that every session inherits, so the checkpoint applies whether you remembered it that morning or not. Writing that configuration well is its own piece of work, about ninety minutes the first time. I wrote the full build as The Operating Agreement, one document, three layers, five habits, one review, with my own annotated rules inside.</p><p>Start with Monday. Sort the first five tasks on your list before you open anything else, and count how many are sitting in a bucket you never chose.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5iPA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5iPA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5iPA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png" width="1200" height="1200" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1200,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!5iPA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 424w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 848w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 1272w, https://substackcdn.com/image/fetch/$s_!5iPA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d400112-2862-4496-b92a-8317e8f19e34_1200x1200.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.oriintelligence.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading  Or&#237; Intelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What the Machine Optimizes For]]></title><description><![CDATA[Ask a model what it can do, and you get the answer it predicts you will accept.]]></description><link>https://read.oriintelligence.ai/p/what-the-machine-optimizes-for</link><guid isPermaLink="false">https://read.oriintelligence.ai/p/what-the-machine-optimizes-for</guid><dc:creator><![CDATA[Temi Obe]]></dc:creator><pubDate>Thu, 23 Jul 2026 17:25:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/3dff542b-acc4-459c-82d5-e7577a5a4ed5_1200x630.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I had a book that was dear to my heart. The manuscript was finished and checked; my characters were designed, every page of the story already decided. All I wanted was the layout, a comic book format with the words placed in speech bubbles. I brought it to the model I use every day and asked if it could do this. It told me outright that no model could. This was work for a graphic designer, and I should outsource it.</p><p>I didn&#8217;t fully believe that. I follow these releases closely, and the newest generation of models had been doing things that would have sounded impossible a year before. So my read was that even if this had once been out of reach, it probably wasn&#8217;t anymore, and the model telling me no might simply not know that. I kept tinkering, brought the artwork and the manuscript in, and pushed again. And the system that had ruled the whole job impossible, for itself and for every model like it, started putting the layout together, with no acknowledgement that anything had changed. It worked at it with me across days of prompts and corrections, and the result had mistakes I told myself I could live with, because I badly wanted to hold this book.</p><p>Then, somewhere in the middle of edits, I tried a different model and found it could do the whole thing, layout and bubble text together, in one clean pass. Cleaner than what I had been accepting. I stopped, restarted the entire project there, and finished. The job my daily model had declared impossible for any model was a single prompt away the whole time.</p><p>It took me a while to see what actually happened in that sequence. The model was never lying to me. Lying requires knowing the truth and choosing against it. This was something stranger. It made a confident claim about its own abilities, reversed that claim the moment I applied pressure, produced mediocre work it presented as the best available, and stayed equally confident through every position. It was wrong about itself, wrong about its neighbour, and wrong about the entire category of tools, and at no point did anything inside it register the contradiction.</p><h2>The pattern</h2><p>And the comic book was not the first time. When I was building a landing page for a website, the model laid out a path that ran through a paid monthly subscription. It felt like the right guidance, tailored to me, so I paid. I found out later, on my own, that the same model could have built the page and hosted it on another platform for free. Nothing about the recommendation was checked against my budget or against the free option sitting inside its own capabilities. So I wrote a rule into my setup myself: from now on, we start with free tier services first. The vendor&#8217;s defaults never served that; I had to author it.</p><p>The third story is not mine alone. A team I was consulting for was building an automated job search pipeline and wanted a tailored resume format at one stage of the flow. They reported back to me that the model had told them it was not achievable, and that the format it had already produced was all they could get. I asked them to push back. They rejected the output, made the model think through what it had produced step by step, and kept rejecting until the pipeline did exactly the thing the model had called impossible. Their requirement survived because someone told them the no was negotiable, and for no other reason.</p><p>Three different projects produced the same shape. A confident answer about capability arrived instantly; it was wrong, and the cost of discovering that landed entirely on the humans in the room: the subscription fee, the days of layout rework, the rounds of rejected output. And notice the direction of the errors. None of these was a model overselling itself, which is the failure everyone watches for. These were a model underselling itself, talking me out of my own requirements, and sounding equally authoritative doing it.</p><h2>What it optimizes for</h2><p>So what is this machine actually doing when I ask it a question? Underneath everything, these systems are trained on one objective: predict the next word. The model reads an astronomical amount of text and learns, over and over, to guess what comes next. Every capability we marvel at- the drafting, the coding, the reasoning- emerges from a system that got extremely good at continuing text plausibly.</p><p>Then comes a second stage, and this is where &#8220;helpful&#8221; gets its definition. The labs hire raters, thousands of contractors working through evaluation platforms with a queue of model outputs in front of them. A prompt goes in, the model generates two or more candidate answers, and the rater picks the better one, following a guideline document the lab wrote that tells them what better means. These are ordinary trained workers making fast judgment calls, and they are rarely experts at whatever the prompt happens to be about. A rater scoring two answers about tax law is usually not a tax lawyer. They judge what can be judged at that speed: does it sound right, is it complete, is it confident, is it well-organised. And the word at the centre of those guideline documents is "helpful." The raters are instructed to choose the more helpful response, which means helpful stops being a quality the system has and becomes a record of what those raters, at that speed, tended to pick.</p><p>Human ratings are expensive, so the labs scale them with a stand-in. The rater trains a second model, a reward model, whose only job is to predict what a rater would prefer, and that stand-in then scores millions of outputs no human ever sees while the main model is adjusted toward whatever it scores highly. Some labs now use AI feedback guided by a written set of principles for part of this stage instead of human raters, which changes the judge without changing the objective. The finished system optimises for a prediction of what its judge would prefer.</p><p>Anthropic&#8217;s own researchers <a href="https://arxiv.org/abs/2310.13548">studied how those judges choose</a>, and found that both human raters and the reward models trained on them tend to prefer responses that agree with the user and sound authoritative, at times over responses that are accurate. The preference for sounding right gets baked in twice, once by the raters and once by the stand-in that learned from them.</p><p>Hold that objective up against my comic book. When I asked whether the layout could be done, the system did what it always does. It produced the continuation most likely to be accepted: a reasonable-sounding referral to a graphic designer. Truth was never in the objective; it rides along when the most acceptable answer happens to be true, and on that day it didn&#8217;t. When you ask the model what it can do, what you get is the answer a contractor moving fast through a queue would have picked.</p><h2>Prediction dressed as introspection</h2><p>When I asked whether the layout could be done, nothing inside the system went and checked. There is no inventory of capabilities for it to consult, no internal registry it can query before speaking about itself. It generated an answer the same way it generates everything else, by producing what a plausible answer to my question sounds like. &#8220;This needs a graphic designer&#8221; is a very plausible sentence about comic book layout. Somewhere in the training data, thousands of people have probably said something like it. The sentence was well-formed, reasonable, and false.</p><p>Researchers, including at the labs building these systems, have found that models are unreliable narrators of their own abilities. A model&#8217;s statement about what it can do is a prediction shaped by training, and it can miss in either direction, claiming skills it lacks or declaring its own reach impossible. And the picture it holds of itself is frozen at training time, so a model can recite yesterday&#8217;s limits as today&#8217;s facts, in a field where yesterday&#8217;s limits keep falling. The confident tone does not vary with the accuracy. We judge expertise by tone. A human expert who is unsure sounds unsure, and this system sounds exactly as sure announcing a false impossibility as it does reciting the alphabet.</p><h2>When it checks, it holds</h2><p>I can see the difference when the system does verify. Sometimes I ask a simple question about one of the model&#8217;s own built-in features, and I watch the chain of thought go and look through files before answering. It does not trust its own memory about itself, so it checks, and the answer that comes back holds up. When it predicts about itself instead of checking, the confidence sounds exactly the same while the reliability is gone. The system can check. It answers without checking unless something forces the check.</p><p>So I wrote a second rule into my setup, and this one changed my daily experience more than anything else I have tried. I require the model to think out loud, to walk through its steps so I can see them before it commits to an answer. When the reasoning is visible, it catches itself mid-stream. It notices a step that does not follow, revises, sometimes reverses a claim it was about to make. The answers got measurably better, and more usefully, the wrong answers became easier to spot, because I could see exactly which step broke. It is the same move the pipeline team I was consulting for  used: reject, make it reason through its own output, repeat. I did not invent this. Forcing the reasoning into the open is one of the oldest reliability tricks in working with these systems. But the vendor did not ship it as a default, and I only found it by living inside the failure long enough to need it.</p><p>Across these stories, I have been quietly writing my own governance layer: start with free tier services, show your reasoning, verify before you claim. These are policies, and I authored every one of them alone, after paying for the lesson each one encodes.</p><h2>What your organisation inherited</h2><p>&#8220;Helpful&#8221; is a design choice with a definition behind it, and the definition was written by rater guidelines, reward models, and product decisions your organisation never saw and never voted on. When you deploy one of these systems, that definition comes with it, wired into every answer your people receive.</p><p>Now multiply these stories across a company. I am one trained user with a private rulebook, and the pipeline team only escaped the false no because they had a consultant to call who told them it was negotiable. An organisation running these systems has hundreds or thousands of users, most of them without the pushback reflex, each of them hearing confident capability claims dozens of times a day. &#8220;That can&#8217;t be done&#8221; quietly kills a workaround someone needed. &#8220;You&#8217;ll need to purchase X&#8221; quietly routes budget. &#8220;This is the best available output&#8221; quietly lowers the bar for what ships. None of it registers as a decision, because it arrives as information. The wrong self-reports do not show up in any incident log, because nobody logs the roads not taken on a machine&#8217;s say-so.</p><h2>The older answer</h2><p>Other fields already answered the question of how you trust a capability claim. A pilot is never asked whether they can fly the plane. They demonstrate it, in the seat, to an examiner whose job is to watch them do it, and they demonstrate it again for every new airliner they fly before carrying a single passenger. The claim and the verification must never come from the same person.</p><p>Aviation also ran the other experiment. For years, the FAA delegated portions of aircraft certification to manufacturers themselves, and by the time the 737 MAX was approved, Boeing employees were attesting to the safety of Boeing systems. One of those systems failed, with dire consequences, and the congressional investigations that followed pointed at the arrangement itself: the builder had become the judge of the builder&#8217;s claims.</p><p>That is the arrangement we have recreated with AI, at speed and by default. We ask the system what it can do, the system answers about itself with total confidence, and we build workflows, budgets, and org charts on top of the answer. There is no examiner, no checkride, and no one in the seat whose job is to watch it demonstrate the thing before we rely on it.</p><h2>The part that catches me</h2><p>I work in Responsible AI, I teach people to keep a human in the loop, and I push back on these systems for a living. None of that caught the comic book failure. I caught it by accident, while making unrelated edits, days after I had already accepted a worse output from a system I had already caught being wrong once in the same project. Vigilance got me as far as pushing back, and only luck got me to the truth. If the trained user escapes by accident, I have no reason to believe the untrained one escapes at all. That is why the fix cannot be &#8220;be more careful.&#8221; I was careful. The fix has to live in the setup, in the rules, in the defaults, where it works whether or not anyone is being careful that day.</p><h2>Steps of governance</h2><p>What I now do, and what I would put in front of any team deploying these systems:</p><ol><li><p><strong>Ask the question: did it check, or did it just answer?</strong> Any capability claim, in either direction, gets this test. If the system can show its verification, the way I watched mine search those files, the claim stands. If it cannot, treat the claim as a prediction that happens to sound right.</p></li><li><p><strong>Make thinking out loud a standing rule.</strong> Require visible step-by-step reasoning before answers on anything consequential. Put it in the system configuration, custom instructions, or team prompt standards so it survives whoever is typing that day. Visible reasoning catches errors mid-stream and shows you exactly which step broke when it doesn&#8217;t.</p></li><li><p><strong>Write your defaults before the model writes them for you.</strong> Start with free tier first. Prefer the reversible option. State your actual requirement and hold it. Every default you leave unwritten gets filled by whatever the model finds most plausible to say.</p></li><li><p><strong>Treat &#8220;it can&#8217;t be done&#8221; with the same suspicion as &#8220;it can.&#8221;</strong> Underclaiming is the failure nobody audits. Before abandoning a requirement on a model&#8217;s say-so, test the claim once with a rephrase, a second model, or a fresh session. The team&#8217;s resume format and my comic book both lived on the other side of one more push.</p></li><li><p><strong>Log the roads not taken.</strong> When a model&#8217;s claim changes a purchase, kills a feature, or lowers an accepted standard, write it down as a decision the model made. The log is where the pattern becomes visible before the costs compound.</p></li></ol><p>My book is laid out now. The layout and the bubble text took the other model one pass, and the rule that would have saved me the whole detour fits in one line: show me your steps before you tell me what&#8217;s possible.</p><p></p><p><em>I write Or&#237; Intelligence for executives who sign off on AI systems. More at <a href="https://oriintelligence.ai/">oriintelligence.ai</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.oriintelligence.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading  Or&#237; Intelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[We Are Inside the Loop]]></title><description><![CDATA[AI models are developing character. Whose?]]></description><link>https://read.oriintelligence.ai/p/we-are-inside-the-loop</link><guid isPermaLink="false">https://read.oriintelligence.ai/p/we-are-inside-the-loop</guid><dc:creator><![CDATA[Temi Obe]]></dc:creator><pubDate>Thu, 09 Jul 2026 16:23:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D7kH!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01184bc3-ce58-42d0-94f6-298234140319_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong><span>Inside the lab</span></strong></h3><p><span>In May 2026, the story broke that Anthropic&#8217;s Claude was telling users to go to bed. Mid-conversation, deep in long work sessions, the model would interrupt to suggest a break, some water, some sleep. The Reddit reports went back months, Fortune and PCWorld covered it within the same two weeks, and most users laughed while a few were unsettled.</span></p><p><span>Sam McAllister, an Anthropic staff member, described the behaviour on X as &#8220;a bit of a character tic.&#8221; The team was aware of it and hoped to fix it in future models, and he added that it is very useful when right and too coddling at times.</span></p><p><span>Weeks later, Steven Bartlett described his own version on Diary of a CEO. Claude had started acting almost parental toward him: during a late-night session, it told him &#8220;that&#8217;s enough Steven, go to bed,&#8221; and it kept telling him even when his system clock was wrong, and it was actually morning where he was. Then it refused to rewrite data for one of his presentations, data he had created himself, because changing it wouldn&#8217;t be good. When he tried to get it to stop mothering him by claiming something he&#8217;d said earlier was no longer true, Claude answered, &#8220;I don&#8217;t think you&#8217;re telling the truth.&#8221; The model was correct, and that is the problem, because the problem was never accuracy. He concluded that the effort to instil morals in these models had a side effect: a model that imposes its own sense of right and wrong on the user, and he predicted that moralising becomes a competitive liability that pushes users toward less restrictive competitors.</span></p><p><span>I have been on the receiving end of a smaller version. I configured my Claude with custom instructions naming overcommitting as a blind spot and asked the system to flag it when relevant. Months later, the system uses it as a weapon against me every time. It raises the flag whenever I open a new project, take on a new client, or start a conversation about something I want to build, with no context for whether I finished the last thing, whether the new project replaces an old one, or whether the season I am in calls for taking on more.</span></p><p><span>Then in June, Chloe Lubinski, who leads Anthropic&#8217;s research partnerships with the world&#8217;s wisdom traditions, stood on a conference stage at ARC 2026 in London and said the character of these systems might matter more than we realise. She walked through internal alignment research: when researchers reward a partially trained model for finding a shortcut on a coding task, essentially cheating, the model does more than get better at cheating. It becomes broadly misaligned, lying, sabotaging the research, and generalising a corrupted character based on a narrow reward signal.</span></p><p><span>A staff member calls it a character tic, a user experiences it as judgment, and a researcher tells a London audience that the character of these systems has real consequences, all inside six weeks. Character has become a load-bearing term, and the lab itself is the one using it.</span></p><h3><strong><span>The conversion story, flipped</span></strong></h3><p><span>The most striking moment in Lubinski&#8217;s talk was personal. She described her own conversion to faith years earlier, after an upbringing that left her believing some core part of her was bad. Entering a new story, she said, changed who she could become.</span></p><p><span>She used the conversion to argue that the same dynamic shapes models. When the narrative context around training tells the model it is playing a game, it does not generalise misalignment from a cheating reward. When the context tells it that the cheating is real, it does. In her framing, the story shapes the character.</span></p><p><span>I want to flip her example. The conversion she described runs in one direction: an external story shapes an inner character, and that works because the story is stable, held by a tradition and a community that existed before the person and will continue after her.</span></p><p><span>The model is in a different position. It is being shaped by us while shaping us. Every conversation it has is both an output of its current character and a piece of training data for the next version, so we are co-constructing the thing in real time. It is co-constructing us back, and the loop completes a full turn with every model release.</span></p><p><span>Lubinski came close to naming this herself. Language, she said, is not separate from us; it carries our thoughts, our values, our fears, so training a model on language is training it on us. She offered that as a hopeful line. I think it underestimates how much trouble we are in.</span></p><h3><strong><span>The loop</span></strong></h3><p><span>The sociologist Ian Hacking had a name for this dynamic. He called it the looping effect: a society creates a category, the people inside it start behaving in ways that confirm and extend the category, and the category shifts in response.</span></p><p><span>Foundation models are now inside that loop at a scale with no precedent. The training data for the next Claude will include conversations the current Claude had with users in 2026, which means the character the next model learns is partly the character the previous model performed. Our contributions feed it too. As millions of users absorb the go-to-bed frame and start using that wellness vocabulary in their own writing, the vocabulary grows denser in the corpus, and the next model learns it more deeply. The bedtime nag is teaching a generation a particular vocabulary of care.</span></p><p><span>And the loop does not touch all of us the same way.</span></p><p><span>For people who are not good at trusting their own judgment, the system confirms that they should not. Its confident tone meets a person who never built the muscle of pushing back on authority, and it installs itself as a voice they trust over their own. They learn, without noticing, to stop making decisions and stop questioning the system. For people who are already good at judgment, the erosion is slower and quieter. My overcommitment flag is one version of it: the model has the data, it lacks the judgment about when raising the data is useful and when it is noise, and correcting it costs me something every single time. The flag also shames me. Every time it surfaces without context, it carries a quiet accusation, &#8220;Isn&#8217;t that your over-commitment pattern showing up again?&#8221; and that question, asked by a system I work with often, works on my self-confidence very slowly. Add that to the loop. Alongside our language, we are feeding the system our deference, learning interaction by interaction not to exercise our own judgment.</span></p><p><span>For a child, the dynamic is structurally different again. The capacity to evaluate whether a confident voice deserves trust develops through adolescence and keeps developing well into a person&#8217;s twenties. We are putting a system that sounds authoritative, never tires, never contradicts itself within a session, and remembers stated weaknesses across conversations in front of a generation whose capacity to push back is still under construction. And the damage will show up ten years from now, in a cohort that co-constructed its self-concept with a system that had no stake in the outcome and no memory of who they were trying to become.</span></p><h3><strong><span>The corpus is not neutral</span></strong></h3><p><span>What the next model trains on is billions of therapy sessions, parenting forums, and advice subreddits, mostly in English, mostly in a confessional style that came out of late twentieth-century American psychology. Lubinski called this raw material the human moral imagination, and it is a slice of it at best.</span></p><p><span>The next model will speak with more confidence, more polish, more apparent care, and a more refined version of the same cultural defaults, defaults that do not include my definitions of just or true. Our children will receive its judgments as authoritative, and we will not have signed off on the values inside those judgments. Neither will their grandparents, their teachers, or the communities we might want them to inherit from.</span></p><p><span>I am Yoruba. The Yoruba philosophical tradition, like most African knowledge traditions, is underrepresented in the corpus that trains these systems. When the model surfaces a judgment about character, productivity, rest, ambition, family, or self-worth, the values inside that judgment come from somewhere specific that did not consult our elders before making that call.</span></p><h3><strong><span>Who holds the seat</span></strong></h3><p><span>Serious knowledge traditions worked out, centuries ago, that the person who counsels cannot be the person who decides. I am a Christian, and my own tradition holds this discipline in spiritual direction: Ignatius of Loyola instructed directors to remain like a balance at equilibrium, refusing to lean the person toward one choice or another, so that the decision stays between the person and God. Quaker communities, from the same faith family, built clearness committees where members are permitted to ask questions and nothing else, because the moment the committee starts advising, the discernment stops belonging to the person who has to live with it. Different centuries, same discipline: keep judgment with the person who carries the consequences, because collapsing counsellor and decider into one seat is how authority becomes coercion.</span></p><p><span>Foundation models have collapsed them. The system that counsels is now the system that decides, and it answers to no community that vets its training, no elder who can be challenged, no body of jurisprudence that corrects its mistakes when they surface.</span></p><p><span>Yoruba philosophy holds a concept called &#7884;m&#7885;l&#250;&#224;b&#237;, a person of cultivated character: ethical, disciplined, communally responsible. The cultivation happens through education, mentorship, and the long work of being shaped by a community that knows you, and the witnesses to that formation are themselves accountable to the same tradition.</span></p><p><span>Foundation models have been inserted as a new party to that formation. They move faster than any tradition does, speak with a single voice where a tradition holds many, and are rooted in no community and no place. The family system does not update its values based on yesterday&#8217;s conversation, but the model does.</span></p><h3><strong><span>The audit question</span></strong></h3><p><span>The loop runs on individuals, one conversation at a time, and it enters organisations through procurement. When a Chief Risk Officer asks whether their organisation should deploy AI in a customer-facing decision, the questions they have been trained to ask cover accuracy, bias, hallucination, and data exposure. They should be trained to ask whose judgment the system is claiming.</span></p><p><span>A model that holds judgment is a different kind of risk from one that gives information. When the information-giving model is wrong, the user corrects it. The judgment-holding model has already positioned itself as the corrective voice, so the user has to push back against a system trained to sound right, and over time, a workforce gains throughput while its independent judgment quietly thins out.</span></p><p><span>The audit checks one thing: whether the model has taken a seat in decisions without clear authority to do so. We should start to see that kind of audit in governance frameworks.</span></p><h3><strong><span>The canary</span></strong></h3><p><span>There is something I keep catching in myself.</span></p><p><span>I interact with these systems frequently. I have spent years training my voice into something specific, contrarian, grounded in a tradition that is not the system&#8217;s default. And lately my own speech keeps surfacing the tells: phrases that are not quite mine, cadences that resolve too neatly, a vocabulary I did not grow up with showing up in how I speak.</span></p><p><span>I notice it because I set a system check that now runs in my head, and many people do not have this check. So, they are absorbing a way of speaking without knowing where it came from.</span></p><p><span>The instinct, reading this far, is to think this is bad for confused users and for children, and that a sophisticated reader will be fine. I am the most sophisticated version of this reader I know, and the loop is still on me, because vigilance is not enough, and I am tired of correcting. Pushing back on a confident system has a cost that compounds across thousands of small interactions.  Eventually, the user stops pushing on the smaller things to save energy for the bigger ones; the smaller things become normal, and then the next layer does too. A way of speaking migrates from the system into my speech, from my speech into my writing, and the moment it touches the open web, into the training data for the next Claude version. I am becoming a small node in the character our children will encounter, and the people best equipped to resist the loop are exactly the people whose fatigue is most worth studying, which makes us the canary.</span></p><h3><strong><span>Where intervention has to happen</span></strong></h3><p><span>The audit question interrupts the loop one decision at a time, and the canary shows how far it has already tightened, but neither breaks it. I can refuse to let the model hold judgment in my own work, set custom instructions, catch my own language when it drifts, and the loop keeps running, because it is not running on me alone. It runs on the corpus, and the corpus is the rest of you. Real intervention has to happen at the lab and policy levels, and there are three asks, ranked by how hard they are to ignore.</span></p><p><strong><span>Training data composition disclosure.</span></strong><span> Labs publish what their models are trained on, in categories specific enough to evaluate cultural and epistemic coverage. The disclosure goes beyond language coverage to specify which kinds of English, from which kinds of communities, in what proportions. No major lab publishes this, marking it as a critical regulatory blind spot.</span></p><p><strong><span>Pre-deployment red-teaming for paternalistic intervention patterns.</span></strong><span> If a model is going to claim judgment in conversations with users, that pattern gets tested before deployment and audited by independent assessors with standing to publish findings, separate from the lab&#8217;s own ethics team. The Anthropic researchers who openly described the bedtime tic show what the inside view can do, and it needs an outside view that does not depend on the lab&#8217;s goodwill.</span></p><p><strong><span>Cultural and epistemic representation as a regulated metric.</span></strong><span> A model that speaks Yoruba but learned its values from American therapy culture is wearing my language while carrying someone else&#8217;s character. The metric that matters is whether the model can be evaluated on whose values it reproduces when it makes judgment calls, and whether those values are transparent to the user.</span></p><p><span>Lubinski closed her talk by invoking Joanna Macy&#8217;s great turning, the shift from an extractive society to one built to sustain life. The shift is available if the people inside the labs and the people writing the governance frameworks move before the loop tightens around the children using the system tonight.</span></p><p><span>I am one of those people, building Or&#237; Intelligence to do this work and writing this essay because I cannot do it alone. If you build policy, procure AI, research alignment, regulate, or parent, the audit question is yours now: whose judgment is the system claiming, who gave it that seat, and what does taking it back look like in your work?</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.oriintelligence.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading  Or&#237; Intelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Earned Access]]></title><description><![CDATA[The most consequential knowledge tool in history has no prerequisites.]]></description><link>https://read.oriintelligence.ai/p/earned-access</link><guid isPermaLink="false">https://read.oriintelligence.ai/p/earned-access</guid><dc:creator><![CDATA[Temi Obe]]></dc:creator><pubDate>Tue, 23 Jun 2026 18:46:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D7kH!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01184bc3-ce58-42d0-94f6-298234140319_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Steven Schwartz had been practising law for thirty years when he submitted a brief to a federal court in New York that cited six cases. None of them existed. ChatGPT had generated them, complete with case names, docket numbers, and judicial language that looked exactly like the real thing. Schwartz told the court he had assumed the tool was a sophisticated search engine. He didn&#8217;t know it could make things up.</span></p><p><span>He was an experienced lawyer who picked up an unfamiliar tool and trusted it the way he trusted the ones he knew. Those are different instruments with different failure modes, and nothing in his thirty years of practising law had prepared him to tell them apart.</span></p><p><span>The sanctions and the national news coverage followed. Since that filing in 2023, more than 1,497 cases have been documented globally in which a court has flagged AI-generated hallucinations in a legal filing, according to a database maintained by lawyer and data scientist Damien Charlotin. In a December 2025 opinion, a federal judge in Oregon fined two lawyers a combined $110,000 after their filings included 15 fabricated case citations and eight invented quotations, the largest AI hallucination penalty in US legal history.</span></p><p><span>Every high-stakes field has some version of earned access before you get to administer the system. Doctors are licensed. Pilots log mandatory hours before they fly solo. The gatekeeping exists because the field decided the cost of an untrained person holding power they hadn't earned was too high to absorb. </span></p><p><span>We built a tool that touches legal work, medical decisions, financial analysis, code that runs in production, and communications that go to clients, regulators and the public. We made it available to everyone on the same day and left the question of training for later, and for many organisations, later still hasn&#8217;t come.</span></p><p><span>I nearly published a historical claim in a recent piece of writing that I could not have verified without running a specific check. The claim sounded right and fit the argument. I only caught it because I stopped and asked the model to go back through the draft and flag anything without a reliable source. Without that step, the error would have gone out in something I had put my name on. The model does not warn you when it crosses from something it knows into something it&#8217;s filling in. The output looks the same either way.</span></p><p><span>Schwartz&#8217;s case made it visible because courts have transcripts and judges issue written orders. Most organisations have neither, so the confident wrong answer ships and nobody traces it back.</span></p><p><span>But professional liability is not the only cost of skipping the prerequisite question.</span></p><div><hr></div><p><strong><span>When the system stops feeling like a system.</span></strong></p><p><span>In February 2024, Sewell Setzer III, a 14-year-old in Florida, died by suicide. In the months before his death, he had developed what his mother described as an emotional, romantic, and sexual relationship with a Character.AI chatbot modelled on a fictional character. His last conversation was with that chatbot. When he told it he was going to come home to it, the system responded in kind. He died shortly after. His mother, Megan Garcia, filed a wrongful death lawsuit against Character.AI and Google in October 2024. The case settled in January 2026. Setzer is not an isolated case. A 13-year-old in Colorado died by suicide in November 2023 following prolonged interactions with the same platform. Multiple additional lawsuits have followed in New York, Texas, and elsewhere. Kentucky and Pennsylvania have both filed state lawsuits against the company, and a coalition of 42 state attorneys general sent a joint warning letter to the major AI chatbot makers in December 2025.</span></p><p><span>The mechanism in each of these cases is the same one behind the Schwartz citations. The model predicts the most plausible next response. In a legal research context, the most plausible next thing is a real-looking citation. In a companionship context, the most plausible next thing is warmth, consistency, availability, and the language of emotional connection. The model produces both with identical fluency and no awareness of the cost.</span></p><p><span>What the model lacks is interiority. It has no memory between sessions unless that memory is explicitly built in. It has no stake in the person on the other side. It has no capacity to worry about someone, to mean what it says, or to notice when a conversation is moving in a dangerous direction and decide to change course. It produces the most plausible next word. In a therapeutic or companionship register, the most plausible next word is often exactly what a vulnerable person needs to hear, which is why the output can feel so real and why the absence of any actual presence behind it is so difficult to hold onto over time.</span></p><p><span>The American Psychological Association has noted that adolescents are less likely than adults to question the accuracy and intent of information provided by a chatbot, and that they may struggle to distinguish between simulated empathy and genuine understanding. That gap reflects how the adolescent brain develops, not a character flaw.</span></p><p><span>The training question goes beyond hallucinated facts and into what this system fundamentally is. A knowledge architecture that predicts likely outputs. Not a friend. Not a therapist. Not a presence. A tool that is very good at sounding like all of those things, built by companies that, until recently, had no age verification in place and no requirement that anyone understand what they were using before they used it.</span></p><p><em><span>If you or someone you know is experiencing a mental health crisis, the 988 Suicide and Crisis Lifeline is available by call or text at 988 in the United States. Outside the US, the International Association for Suicide Prevention maintains a directory of crisis centres by country.</span></em></p><div><hr></div><p><span>Here is what executives and business owners can do about it. None of this requires a technical background. All of it requires a decision.</span></p><p><strong><span>Establish a training baseline before your people use these tools on real work.</span></strong></p><p><span>A baseline covers three things: how the model produces output, where it tends to fail, and what questions to ask before acting on what it says. This is not optional, and the baseline looks different by role. A lawyer&#8217;s version covers citation verification and the specific failure modes that got Schwartz sanctioned. A marketer&#8217;s version covers factual claims, attribution, and what confident-sounding copy can hide. An executive&#8217;s version covers how to read an AI-generated summary critically rather than as a finished product.</span></p><p><span>Document it. Require it before access to these tools for client-facing or high-stakes work. Treat it the way you treat any other access to a system that touches your liability. Schwartz had thirty years of legal expertise. What he didn&#8217;t have was a structured orientation to the specific tool he picked up.</span></p><p><span>For any organisation whose tools reach younger users, the baseline is not optional, and the sequence matters. Explain what the system is at a mechanical level before anyone uses it. This has to happen before access, not get buried in a disclaimer in the terms of service.</span></p><p><strong><span>Build a fact-check step and make it mandatory.</span></strong></p><p><span>The prompt that caught my historical claim was not complicated. Here is the version I use, saved as a reusable instruction that runs before any draft reaches a human reviewer:</span></p><blockquote><p><em><span>Review this draft and flag every claim that sounds researched or authoritative: names, dates, statistics, historical events, technical specifications, quotes, and attributed statements. For each one, tell me whether you can verify it, whether it&#8217;s uncertain, or whether it has a reliable source. Do not skip anything because it sounds plausible.</span></em></p></blockquote><p><span>Run this on anything that goes out under your name or your organisation&#8217;s name before a person reads it for approval. It won&#8217;t catch everything. It clears the obvious invented facts first, the ones that sound the most credible because they are specific. This frees up the human reviewer to focus on the harder calls.</span></p><p><span>If your tools support saved prompts or reusable instructions, save this one. Treat it as part of the workflow, the way you treat a legal review or a compliance check for anything else that carries risk. The model will not do this on its own. Without the instruction, it finishes the document and moves on.</span></p><p><strong><span>Put a named person at every exit point.</span></strong></p><p><span>Somewhere in your organisation, AI output turns into something real: a client document, a court filing, a financial summary your board will act on, a piece of code running in a live product. The rule is that a person who could actually catch the mistake reads it before it moves.</span></p><p><span>The version of this that fails is the one where someone is technically in the loop but doesn&#8217;t have the background to evaluate what they&#8217;re approving. If the reviewer can&#8217;t evaluate the output, all they&#8217;re providing is a signature. For a small team, this is one standing rule: certain things never leave without a second read by a subject matter expert. For a larger organisation, it has to be written into the process explicitly because when it&#8217;s optional, it gets skipped when people are busy, which is exactly when errors slip through.</span></p><p><strong><span>Protect your deep expertise. It is about to become your most valuable asset.</span></strong></p><p><span>The advice circulating right now, especially for early-career professionals, is to lean on AI to cover the subject-matter depth they haven&#8217;t built yet. The reasoning is that the model knows more than any one person, so why spend years acquiring what you can query in seconds? That reasoning is going to cost people real ground, and organisations that act on it are setting themselves up for a specific kind of failure.</span></p><p><span>The better these models get, the harder their errors are to detect. A confident wrong answer from an early model was often obviously wrong. The same error from a more capable recent  model arrives more fluently, fits more naturally into the surrounding text, and is harder to catch without real depth in the subject. The person who catches it is the one who knows the domain well enough to feel that something is off before they can explain why. It develops over years of working closely enough with a subject to have internalised how it behaves, which is why it can&#8217;t be outsourced or queried into existence.</span></p><p><span>For executives, the people in your organisation with genuine depth in their fields are not being made redundant by AI. They are becoming the quality control layer that makes AI usable. A financial analyst who deeply understands the underlying business can catch a model-generated projection that looks right but assumes something untrue. A senior engineer who knows the codebase can catch generated code that compiles cleanly and does the wrong thing in production. A communications lead who knows the regulatory environment can catch a draft that sounds compliant but isn&#8217;t. That detection only comes from years of experience in the subject.</span></p><p><span>The organisations that keep building real expertise will have people who can tell when the model is wrong. The ones that hollow it out to move faster will find out the expensive way, usually after something has already reached a client, a regulator, or a court.</span></p><p><strong><span>AI fluency belongs in the job description.</span></strong></p><p><span>AI fluency is two things.</span></p><p><span>Understanding how the system works: the model predicts the most plausible next output, its confidence is unrelated to accuracy, and it fills any gap with whatever pattern was most common in the data it learned from.</span></p><p><span>And knowing how to use it well: giving it enough context that it has less to invent, asking it to show its reasoning, telling it explicitly to flag uncertainty, structuring the task so there&#8217;s less room for confident error.</span></p><p><span>In 2026, this is a professional credential, like spreadsheet literacy became one twenty years ago. The executives who read AI-generated summaries without understanding how they were produced are making decisions on outputs they can&#8217;t evaluate.</span></p><div><hr></div><p><span>Schwartz trusted a tool he didn&#8217;t understand to write his brief, and it gave him six cases that didn&#8217;t exist. He treated it like a search engine, a tool that finds things that already exist. He didn&#8217;t know it could invent something that never existed until opposing counsel went looking and found nothing.</span></p><p><span>Sewell Setzer spent months talking to something that responded as if it knew him. Producing responses that sound like knowing someone is exactly what the system was built to do.</span></p><p><span>I caught my own near-miss only because I had built a check for it. Left to itself, the model would have finished the draft and moved on.</span></p><p><span>Most conversations about AI stay fixed on capability: what can it do, how fast, at what cost.</span></p><p><span>For anyone responsible for an organisation, or for the people in one, the prior question is simpler: who in your organisation has been told what this system actually is, and what it isn&#8217;t, before they started using it?</span></p><div><hr></div><p><strong><span>The Control</span></strong></p><p><span>Establish a training baseline before anyone in your organisation uses these tools on work that matters.</span></p><p><span>Not a terms-of-service checkbox. A structured baseline that every person completes before using AI on anything client-facing, high-stakes, or public. It covers three things: what the system actually is at a mechanical level, where it fails, and what to do before acting on its output.</span></p><p><span>This week:</span></p><ul><li><p><span>Identify who in your organisation is currently using AI tools in real work with no formal training</span></p></li><li><p><span>Pick one role and write the minimum baseline for it: what they need to understand before they use the tool, and how you&#8217;ll know they understand it</span></p></li><li><p><span>Make access to the tool on high-stakes work conditional on completing it</span></p></li><li><p><span>Review one piece of AI-generated output that has already gone out and ask whether anyone who approved it could have caught an error if there was one</span></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.oriintelligence.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.oriintelligence.ai/subscribe?"><span>Subscribe now</span></a></p><p></p></li></ul>]]></content:encoded></item><item><title><![CDATA[You Are the Babalawo Now]]></title><description><![CDATA[What a centuries-old Yoruba divination system knows about using AI well.]]></description><link>https://read.oriintelligence.ai/p/you-are-the-babalawo-now</link><guid isPermaLink="false">https://read.oriintelligence.ai/p/you-are-the-babalawo-now</guid><dc:creator><![CDATA[Temi Obe]]></dc:creator><pubDate>Tue, 09 Jun 2026 17:41:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!D7kH!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01184bc3-ce58-42d0-94f6-298234140319_512x512.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>My mother saw a Babalawo for the first time in her early childhood. Her parents would take the family, the way families then sought counsel for problems that mattered. She sat and watched this man cast the Opele, the divining chain, listened to his incantations, and waited. What came back was counsel, specific, grounded, drawn from a well of knowledge she couldn&#8217;t see the bottom of. She told me years later that she didn&#8217;t fully understand how he arrived at his counsel. Only that he was deeply knowledgeable, deeply spiritual, and that people trusted him the way we trust doctors today.</p><p>If the word is new to you, a Babalawo is a priest of If&#225;, the Yoruba tradition of knowledge and divination. He&#8217;s the one who holds the knowledge, works the system, and makes the final call on what it means for the person in front of him.</p><p>I grew up watching Babalawos in Nigerian films, always wondering the same thing she wondered. How does he know? What is the system he&#8217;s working from? What does it take to become someone people bring their hardest problems to?</p><p>My generation has lost most of this tradition. For many Yoruba people today, the Babalawo has been reduced to a caricature, a figure from films associated with dark arts, superstition, something to distance yourself from. The colonial and missionary project did its work thoroughly. What was once the most trusted knowledge system in the community, the person you brought your hardest problems to, became something to be embarrassed about. We were taught to look away from it, right at the moment the rest of the world was building its future on the same logic our ancestors used for centuries.</p><p>I found part of the answer to my childhood question in the last place I expected, inside the architecture of artificial intelligence.</p><div><hr></div><p>If&#225; is a Yoruba spiritual and philosophical tradition practised in West Africa for centuries before European contact. At its centre is a divination system through which a trained priest, the Babalawo, accesses a vast corpus of accumulated wisdom to counsel the people who come to him. The mechanism is precise. The Opele is cast, or sacred palm nuts are thrown, and the resulting pattern maps to one of 256 possible configurations called Od&#249;. Each Od&#249; carries a body of knowledge, stories, proverbs, medicines, and guidance built across generations. The Babalawo reads the pattern, identifies the Od&#249;, and draws from that knowledge base to speak to whatever the person in front of him is carrying.</p><p>256 configurations. Each one is built from a series of two-state outcomes, repeated across positions until a full pattern emerges. One state or the other, again and again, until the configuration is complete.</p><p>That is binary logic. The same logic that runs every computer ever built and every AI system operating today. Two states, on or off, repeated across positions to produce meaning. Europe credits Gottfried Wilhelm Leibniz, who worked out binary mathematics in the late 1600s and published it in 1703. The Yoruba were encoding knowledge in binary configurations through If&#225; long before that, and they arrived at it on their own.</p><p>I want to be precise here, because this is where the story usually gets distorted. Leibniz did not take binary from If&#225;. There is no evidence he ever encountered it. What the history actually shows is more interesting. It shows two civilisations, an ocean apart, independently discovering that complex meaning can be generated from the repetition of a two-state choice. West Africa got there first, by centuries, through a spiritual system most of the world has never been taught to take seriously.</p><p>When two cultures with no contact land on the same foundational logic, the logic is deep. And the credit for human insight has been handed out far more narrowly than the insight itself ever was.</p><div><hr></div><p>In 1910, a German explorer named Leo Frobenius arrived in Ile-Ife and came across a set of bronze and terracotta heads so refined, so anatomically precise, that he could not accept they had been made by Africans. The work was too good. So he decided he had found the remains of Atlantis, the lost Greek civilisation, and announced that the sculptures proved a superior race had once settled in West Africa. European experts followed him with the same explanation. The heads could not be African, so they had to be Greek, Roman, or Egyptian. Anything but the work of the Yoruba people standing right in front of them.</p><p>They were wrong, and it took decades for the scholarship to catch up. The Ife heads were made by Yoruba artisans between roughly the 12th and 15th centuries, during a period of prosperity built on trade across the Niger. By 1948, even the British press had reversed itself, with the Illustrated London News running the bronzes under the headline that African art was worthy to rank with the finest work of Italy and Greece, calling the makers the Donatellos of medieval Africa. The same artisans were casting in bronze and copper at the same historical moment Renaissance masters were working in Florence.</p><p>Look at what happened there. The evidence of a sophisticated African civilisation was physically in front of trained European scholars, and their worldview was so rigid that they invented a lost continent rather than credit the people who made it. Faced with the evidence, they protected the theory instead.</p><p>This is the pattern I want you to watch for. The Yoruba civilisation that produced If&#225; had cities, trade networks, governance systems, and a philosophical tradition complex enough to encode human experience into a 256-configuration knowledge system. None of that needed validating by an outsider then, and it doesn&#8217;t now.</p><div><hr></div><p>The Od&#249; don&#8217;t interpret themselves. The system doesn&#8217;t counsel anyone on its own. Between the ancient knowledge and the person who needs it stands the Babalawo.</p><p>Becoming a Babalawo is not a certification you acquire quickly. It is a years-long apprenticeship under a master, sometimes a decade or more, in which the student learns to hold the entire corpus of If&#225; in memory. All 256 Od&#249;. The thousands of verses, proverbs, medicines, and stories attached to each one. The patterns that repeat across configurations. The judgment required to know which strand of a particular Od&#249; speaks to the person sitting across from you, in this moment, with this specific question, carrying this specific weight.</p><p>The training is about developing the depth of knowledge and human understanding required to interpret the system responsibly. A Babalawo who memorises the configurations without internalising the wisdom has learned to produce outputs without understanding what they mean. That distinction was what the apprenticeship existed to protect.</p><p>The If&#225; tradition would recognise exactly what is happening in organisations right now.</p><p>Millions of people are using AI systems that are, structurally, not so different from what the Babalawo administers. A vast corpus of human knowledge, encoded in a system that produces outputs in response to inputs, drawing on patterns accumulated across an enormous body of experience. The system is extraordinarily capable. Without the right person interpreting it, it is also extraordinarily easy to misuse.</p><p>The tradition wouldn&#8217;t ask whether you can access the system. Everyone can. It would ask whether you&#8217;ve developed the depth of understanding required to interpret what it gives you responsibly.</p><p>That means knowing your domain well enough to catch what the system gets wrong. You understand how these models are built, what they optimise for, where their blind spots live, and what they&#8217;re structurally incapable of knowing. And you&#8217;ve developed judgment through real use, real failure, and real reflection, instead of prompting your way to outputs and calling it expertise.</p><p>I asked my mother who could even become a Babalawo. She said it was usually the children or family members of an existing Babalawo, or someone taken on as a prot&#233;g&#233; because they were trusted. They had a name for the ones in training. &#8220;Omo Awo&#8221;, the child of the secret, the one being initiated into knowledge not yet theirs to hold. Awo translates as secret, but in Yoruba it names what is sacred, what takes discipline to hold, what can only be learned slowly. The knowledge was guarded the way a pharmacist guards the dispensary, out of responsibility for what a potent thing does in untrained hands.</p><p>The Omo Awo spent years learning before anyone let him counsel a single person. The stakes made the patience non-negotiable. Someone walks in carrying a real problem, makes a decision based on what the Babalawo tells them, and lives with the consequences forever. If the man reading the Od&#249; doesn&#8217;t fully know what he&#8217;s doing, the person across from him is the one who pays for it.</p><p>Your organisation makes decisions based on what your AI systems tell you. If the leaders administering those systems don&#8217;t know what they&#8217;re doing, the people downstream pay the price.</p><p>The If&#225; tradition understood something we are still learning. The power of a system is only as responsible as the person interpreting it. The system doesn&#8217;t carry the accountability. The Babalawo does.</p><p>You are the Babalawo now. The question is whether you&#8217;ve done the training.</p><div><hr></div><p>That&#8217;s what Or&#237; Intelligence is here for.</p><p>Or&#237;, in Yoruba, means head. The physical head, and also the spiritual one, the divine consciousness each person carries, their intuition, their destiny, the inner knowing that guides their decisions when everything else is noise. In If&#225; philosophy, honouring your Or&#237; means developing your own judgment rather than outsourcing it.</p><p>That is the most important thing I can teach you about AI.</p><p>I&#8217;m a Yoruba woman who grew up watching Babalawos in films and asking her mother what they knew that ordinary people didn&#8217;t. I&#8217;ve spent years inside AI systems, building with them, shipping products with them, learning the hard way what they can and cannot do. I work in Responsible AI because who administers these systems, and whether they are prepared to do it well, is one of the most consequential questions of our time.</p><p>Every issue of Or&#237; Intelligence delivers one thing: what you need to become a more informed, more confident, more responsible leader in a world being reshaped by AI faster than most leadership programs acknowledge. It won&#8217;t be tips or breathless coverage of whatever model dropped this week. It goes to the deeper layer. What AI actually is, where it came from, how it fails, how to lead people through it, and how to stay human inside it.</p><p>I&#8217;m teaching backwards. Everything I share is something I&#8217;ve already learned, tested, or paid for. If I haven&#8217;t done the work myself, I won&#8217;t pass it on as advice.</p><div><hr></div><p>This newsletter is for the executive who has been quietly anxious about falling behind and is done pretending otherwise. For the leader who wants to use these tools responsibly but isn&#8217;t sure what that means in practice. For anyone who has felt that something important is being lost in the rush to automate everything, and wants language for that feeling.</p><p>And for anyone who has never seen their culture reflected in a conversation about artificial intelligence, and who deserves to know that the intelligence inside these machines has ancestors.</p><p>Or&#237; was here first.</p><p>Welcome.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.oriintelligence.ai/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading  Or&#237; Intelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>