<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Contract Signal]]></title><description><![CDATA[AI and data intelligence for the contracts professional, focused on turning post-signature enterprise contracts into usable data and intelligence. Frameworks, analysis, and insider perspective from 10 years at Axiom and Knowable (LexisNexis).]]></description><link>https://thecontractsignal.com</link><image><url>https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png</url><title>The Contract Signal</title><link>https://thecontractsignal.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 21 Jul 2026 01:35:33 GMT</lastBuildDate><atom:link href="https://thecontractsignal.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Leonid Prilutskiy]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[leonidprilutskiy@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[leonidprilutskiy@substack.com]]></itunes:email><itunes:name><![CDATA[Leonid Prilutskiy]]></itunes:name></itunes:owner><itunes:author><![CDATA[Leonid Prilutskiy]]></itunes:author><googleplay:owner><![CDATA[leonidprilutskiy@substack.com]]></googleplay:owner><googleplay:email><![CDATA[leonidprilutskiy@substack.com]]></googleplay:email><googleplay:author><![CDATA[Leonid Prilutskiy]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[From Vague Request to Structured Contract Workstream]]></title><description><![CDATA[How to use a Claude or Codex skill to turn incomplete contract review requests into structured workstreams]]></description><link>https://thecontractsignal.com/p/claude-codex-contract-intake-skill</link><guid isPermaLink="false">https://thecontractsignal.com/p/claude-codex-contract-intake-skill</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 19 Jul 2026 23:11:01 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/43e8c3f3-e19a-4eef-9d4a-cf499e2acbcc_2800x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3>Problem</h3><p>Imagine receiving this request:</p><blockquote><p>We signed an LOI to acquire ABC Corporation. Please start reviewing the top 100 customer contracts. We expect to close quickly and will send a data-room link.</p></blockquote><p>It can be tempting to get access to the repository ASAP and start reviewing contracts. However, it&#8217;s usually better to scope out what the request is actually asking for and any available context to help inform execution.</p><p>That is what the <strong>intake-contract-workstream</strong> skill is designed for &#8212; helping<strong> </strong>clarify and understand<strong> </strong>an inbound request for contracts analysis in sufficient depth to effectively begin planning and execution.</p><h3>Where intake fits in the broader contract workstream framework</h3><p><span>Almost every post-executed enterprise contract workstream, big or small, can be framed as the same underlying six-part process:</span></p><ol><li><p><strong>Trigger</strong> &#8212; an event occurs that requires contracts work</p></li><li><p><strong>Define criteria</strong> &#8212; determine what you&#8217;re looking for and why</p></li><li><p><strong>Find</strong> &#8212; apply the criteria to locate the target population of contracts</p></li><li><p><strong>Analyze</strong> &#8212; go deeper on the population: validate the initial search results, incorporate nuances, exceptions, and additional criteria, and perform a comprehensive assessment</p></li><li><p><strong>Decide</strong> &#8212; translate the analysis into recommendations, remediation paths, answers to stakeholder questions and other internal deliverables</p></li><li><p><strong>Remediate</strong> &#8212; take action through counterparty outreach, internal communication, negotiation, and workflow tracking</p></li></ol><p>For a full breakdown of the six stages, see the article: <a href="https://thecontractsignal.com/p/the-post-executed-contract-intelligence-framework?r=1iyonb.">The Post-Executed Contract Intelligence Framework</a>.</p><p><span>This article focuses on the first phase: </span><strong><span>Trigger / Intake.</span></strong></p><p>One call-out re terminology: the <strong>trigger</strong> is the event or request (usually outside of our control). <strong>Intake</strong> is the process of turning that event or request into enough structured context to begin planning the work. </p><h3>What the skill does</h3><p>The <strong>intake-contract-workstream</strong> skill acts as a structured thought partner.</p><p>Give it an email, notes, or other background, and it helps produce a structured intake brief that:</p><ul><li><p>Classifies the trigger or request</p></li><li><p>Identifies the decision or deliverable that contract work needs to support.</p></li><li><p>Translates business context into relevant scope: expected contract populations and data points.</p></li><li><p>Distinguishes known facts from missing inputs, assumptions, and unknowns.</p></li><li><p>Identifies stakeholders, approvals, and potential blockers.</p></li><li><p>Generates recommended next steps - including what can begin now versus what should wait and scope-change risks.</p></li></ul><p><span>It supports six common request categories (archetypes):</span></p><ol><li><p><span>M&amp;A (Corporate Transactions)</span></p></li><li><p><span>Regulatory and policy remediation</span></p></li><li><p><span>Backlog / legacy repository review</span></p></li><li><p><span>Business-as-usual stakeholder questions</span></p></li><li><p>Commercial Opportunity / Optimization</p></li><li><p>Incident response</p><p></p></li></ol><h2>Two ways to use the skill</h2><h4>Option 1: Provide the underlying context directly</h4><p>If your organization&#8217;s policies and approved environment permit it, you can provide the triggering request or background directly to the model.</p><p>Whether that&#8217;s appropriate depends on the information involved, the system you&#8217;re using, and your organization&#8217;s confidentiality, privilege, security, and data-handling requirements.</p><h4>Option 2: Use the skill to generate an intake checklist</h4><p>Where providing the underlying context directly is not appropriate, the skill can instead generate a structured checklist of the information that would be helpful to collect.</p><p>You can then work through that checklist separately and return only the information that is appropriate to provide.</p><p>This is less direct, but it can still help structure an incomplete request and identify the inputs most likely to affect scope or execution.</p><h3>Limitations / What it does not do</h3><p>The skill does <strong>not</strong> establish the final scope of the work, perform contract analysis, make decisions, or perform remediation actions.</p><p>Its job is to create a better starting point for those later steps while preserving the role of human judgment in defining scope, analyzing contracts, making decisions, and taking action. </p><p>It may identify expected / provisional contract populations or possible criteria, but those should not be treated as approved search or review criteria without appropriate review.</p><h3>Paid subscriber resources</h3><p><strong>Paid subscribers receive:</strong></p><ul><li><p>The complete <strong>intake-contract-workstream</strong> skill bundle for each of Claude and Codex.</p></li><li><p>Installation instructions for project-scoped and personal use.</p></li><li><p>A worked M&amp;A example (hypothetical).</p></li><li><p>Future updates to the intake framework and request category modules.</p></li></ul>
      <p>
          <a href="https://thecontractsignal.com/p/claude-codex-contract-intake-skill">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Post-Executed Contract Intelligence Framework]]></title><description><![CDATA["AI contract review" is six different problems being lumped into one category. Here's the breakdown.]]></description><link>https://thecontractsignal.com/p/the-post-executed-contract-intelligence-framework</link><guid isPermaLink="false">https://thecontractsignal.com/p/the-post-executed-contract-intelligence-framework</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 13 Jul 2026 02:03:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d73c9e1e-632a-4dc4-ab15-6ff685ad13de_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>You get an internal email announcing an M&amp;A deal. </span></p><p><span>The clock starts running on the contracts workstream.</span></p><p><span>You&#8217;re juggling a handful of other tasks. You&#8217;re being told to use AI. Life outside of work doesn&#8217;t stop either.</span></p><p><span>So how do you proceed?</span></p><p><span>Almost every post-executed enterprise contract workstream, big or small, can be framed as the same underlying process. It doesn&#8217;t matter whether it&#8217;s M&amp;A diligence, regulatory remediation, a backlog review or a one-off question from sales.</span></p><p><span>In this article I&#8217;ll break that process into six pieces and walk through each one, using M&amp;A as the core example.</span></p><p><span>The six pieces are:</span></p><ol><li><p><strong><span>Trigger</span></strong><span> &#8212; an event occurs that requires contracts work</span></p></li><li><p><strong><span>Define criteria</span></strong><span> &#8212; determine what you&#8217;re looking for and why</span></p></li><li><p><strong><span>Find</span></strong><span> &#8212; apply the criteria to locate the target population of contracts</span></p></li><li><p><strong><span>Analyze</span></strong><span> &#8212; go deeper on the population: validate the initial search results, incorporate nuances, exceptions, and additional criteria, and perform a comprehensive assessment</span></p></li><li><p><strong><span>Decide</span></strong><span> &#8212; translate the analysis into recommendations, remediation paths, answers to stakeholder questions and other internal deliverables</span></p></li><li><p><strong><span>Remediate</span></strong><span> &#8212; take action through counterparty outreach, internal communication, negotiation, and workflow tracking</span></p></li></ol><p><span>The details may vary significantly from one project or task to another but the underlying process will remain fairly consistent.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h3><strong><span>1. Trigger</span></strong></h3><p><span>There are all sorts of possible triggers, but almost all of them fall into three main buckets:</span></p><ol><li><p><strong><span>External events - </span></strong><span>something outside the normal course of business creates the need for contract work.</span></p><ol><li><p><span>Two of the biggest examples are:</span></p><ul><li><p><span>M&amp;A (buy-side, sell-side, reorgs, etc.)</span></p></li><li><p><span>Regulatory remediation</span></p></li></ul></li></ol></li><li><p><strong><span>Internal initiatives - </span></strong><span>the business identifies an internal need.</span></p><ol><li><p><span>The classic example is a backlog or legacy review: getting better visibility into a large historical contract population, identifying risks and opportunities, or solving for one or more specific use cases</span></p></li><li><p><span>These projects can be very valuable, but they are often perceived as nice to have and easier to defer than an acquisition or regulatory deadline</span></p></li></ol></li><li><p><strong><span>Business as usual - </span></strong><span>the steady stream of recurring and one-off questions from various business stakeholders. For example:</span></p><ul><li><p><span>Sales is negotiating a renewal and wants to understand the existing agreement</span></p></li><li><p><span>Finance is asking about revenue commitments, deliverable timing or pricing terms</span></p></li><li><p><span>Product wants to know what AI or data-use rights the company has</span></p></li><li><p><span>Legal wants to identify upcoming renewals, termination rights, or unusual obligations</span></p></li></ul></li></ol><p><span>The type of trigger matters because it shapes everything downstream: urgency, stakeholders, deadlines, risk tolerance, and most importantly, </span><strong>the criteria.</strong></p><div><hr></div><h3><strong><span>2. Define criteria</span></strong></h3><p><span>Before searching for contracts or specific clauses, you need to answer two interrelated but different questions:</span></p><p><span>1. </span><strong><span>Which contracts are in scope?</span></strong></p><p><span>2. </span><strong><span>What do we need to know about them?</span></strong></p><p><span>I think of these as </span><strong><span>search criteria </span>(aka population criteria)</strong><span> and </span><strong><span>analysis criteria</span></strong><span>.</span></p><p><strong><span>Search criteria: which contracts matter?</span></strong></p><p><span>For an M&amp;A transaction, examples of the relevant population include:</span></p><ul><li><p><span>The top 100 or 200 customer contracts by revenue</span></p></li><li><p><span>All active vendor agreements above a certain spend threshold</span></p></li><li><p><span>Agreements involving a particular legal entity</span></p></li></ul><p><span>The initial scope or data may come from a variety of places: for example the finance team, the seller, the M&amp;A team or a data room.</span></p><p><span>This is where metadata that is use-case-agnostic often becomes essential: document title, party names, dates, contract type and similar information.</span></p><p><strong><span>Analysis criteria: what do we need to know?</span></strong></p><p><span>The second question is more substantive and depends on the use case. For M&amp;A, common areas of interest include:</span></p><ul><li><p><span>Assignment restrictions</span></p></li><li><p><span>Change-of-control provisions</span></p></li><li><p><span>Consent and notice requirements</span></p></li><li><p><span>Termination rights</span></p></li><li><p><span>Governing law</span></p></li><li><p><span>Restrictive covenants</span></p></li><li><p><span>Liability and indemnification provisions</span></p></li></ul><p><span>These provisions help determine the plan for the contract under the transaction, what actions may be required, what risks the acquirer inherits, and what alternatives exist.</span></p><p><span>For data privacy or regulatory work, the criteria might instead include:</span></p><ul><li><p><span>Does the required privacy language exist?</span></p></li><li><p><span>Does it satisfy the relevant regulatory standard?</span></p></li><li><p><span>Where are the parties located?</span></p></li><li><p><span>Where is data stored or processed?</span></p></li><li><p><span>Are sub-processing or cross-border transfers permitted?</span></p></li></ul><p><span>For business-as-usual questions, the search criteria and analysis criteria may merge. &#8220;Find the contract with Customer X and tell me whether they can terminate for convenience&#8221; already contains both the search criteria and the substantive question. In addition, sometimes simply finding the relevant agreement is enough.</span></p><p><span>There are also common secondary filters: these include things like active versus expired contracts, upcoming renewals, geography, contract value, business unit. Note that the relevance of each of these depends on the use case. An expired contract might be irrelevant to an assignability analysis but highly useful when someone asks, &#8220;What language did we agree to last time?&#8221; or &#8220;Do we already have an agreement with this counterparty?&#8221;</span></p><p><span>One final point: flagging missing information is often valuable in and of itself. In an M&amp;A context, an inability to locate critical customer agreements or understand material contractual obligations can become a substantive diligence finding rather than merely an inconvenience.</span></p><div><hr></div><h3><strong><span>3. Find</span></strong></h3><p><span>Once the criteria are defined, the next question is: How do you identify the relevant contracts?</span></p><p><span>The answer depends on both what you&#8217;re looking for and what information you have to work with. There are two common starting points: a list or a data dump.</span></p><p><strong><span>You get a list</span></strong></p><p><span>For example the M&amp;A team or the seller provides the top customers by name.</span></p><p><span>In that case, finding the relevant contracts is a matching exercise: compare the list against the party names in the contracts. If there&#8217;s a match, it generally means the contract is in scope. </span>The comparison can run at several levels of strictness but the core concept is the same:</p><ul><li><p><span>Exact legal-entity match</span></p></li><li><p><span>Fuzzy or approximate name match (e.g. Acme Corporation vs. Acme Corp.)</span></p></li><li><p><span>Affiliate or corporate-family match</span></p></li></ul><p><span>One reason the specific matching logic used is important is because contracts are entered into by legal entities, </span><strong><span>while the business often thinks about contracts in terms of broader customer or vendor relationships. </span></strong><span>So the key is finding the balance between exact enough to avoid too many irrelevant documents (false positives) while flexible enough to avoid missing any legitimate ones (false negatives).</span></p><p><strong><span>You get a data dump</span></strong></p><p><span>Alternatively, you may receive a data room full of documents loosely described as &#8220;top customer contracts,&#8221; with inconsistent naming conventions and organization. In this case, there are multiple paths forward.</span></p><p><span>You can first identify parties and other metadata to organize the population.</span></p><p><span>Or you can begin with substantive clause or language-based filtering&#8212;for example:</span></p><ul><li><p><span>Which contracts restrict assignment?</span></p></li><li><p><span>Which require consent?</span></p></li><li><p><span>Which require notice only?</span></p></li><li><p><span>Which provide termination rights in connection with the transaction?</span></p></li></ul><p><span>In practice, these approaches often work together. The list helps define the initial population. The follow-on searches pressure test and refine the scope of the population of interest and prioritizes them further, as needed.</span></p><div><hr></div><h3><strong><span>4. Analyze</span></strong></h3><p><span>The line between finding a contract and analyzing it isn&#8217;t always clear. For example, flagging a contract for assignment restrictions is already a form of analysis.</span></p><p><span>But the analyze stage is where you go deeper, and it has four layers:</span></p><ol><li><p>Validate the search results</p></li><li><p>Incorporate exceptions, carveouts and nuances</p></li><li><p>Expand the criteria</p></li><li><p>Synthesize</p></li></ol><h4><strong><span>Validate</span></strong></h4><p><span>First, confirm that the results actually match your intent &#8212; e.g. that every contract flagged with assignment restrictions actually has them.</span></p><p><span>If a system identifies 75 agreements as containing assignment restrictions, do they really contain assignment restrictions relevant to the transaction?</span></p><p><span>A clause may contain the word &#8220;assignment&#8221; but address the counterparty&#8217;s assignment rights rather than yours. It may apply only to certain rights or obligations. It may not directly use the word &#8220;assignment&#8221; or &#8220;assign&#8221; at all. Or it may use the same keywords but be completely irrelevant &#8211; for example referring to a work assignment rather than the legal transfer of an asset.</span></p><p><span>Key point: finding text is </span><strong><span>not</span></strong><span> the same as determining its semantic meaning or significance.</span></p><h4><strong><span>Incorporate exceptions, carveouts and nuances</span></strong></h4><p><span>Next, examine the exceptions, carveouts, and situation-specific facts that can change the initial categorization.</span></p><p><span>Suppose a contract says consent is required for assignment&#8212;unless the agreement is assigned to an affiliate.</span></p><p><span>If the deal structure involves assignment to an affiliate, the remediation path for that agreement may move from &#8220;consent required&#8221; to a completely different remediation path such as &#8220;courtesy notice&#8221;.</span></p><p><span>Or consider a clause permitting assignment in connection with a merger, sale of substantially all assets, or similar transaction. Whether the exception applies depends on both the contractual language and the structure of the deal.</span></p><p><span>This is why the contract and the business context cannot be separated.</span></p><h4><strong><span>Expand</span></strong></h4><p><span>Finally, the analysis often expands beyond the original criteria. You may begin by looking for assignment restrictions. But once you understand the contract population and transaction structure, additional questions emerge.</span></p><p><span>For M&amp;A, examples include:</span></p><ul><li><p><strong><span>Change of control</span></strong></p><ul><li><p><span>Assignment and change of control belong to the same general umbrella of concepts and are often analyzed together, but they are not identical.</span></p></li><li><p><span>The specific transaction structure may trigger a change-of-control provision even where no contractual assignment occurs.</span></p></li></ul></li><li><p><strong><span>Termination rights</span></strong></p><ul><li><p><strong><span>Termination rights (</span></strong><span>especially for convenience </span>or a result of corporate transactions) can significantly impact the risk/opportunity profile of a contract</p><ul><li><p><span>If a customer can terminate a significant revenue contract at will, that may represent risk to the acquirer.</span></p></li><li><p><span>If the acquired company has a broad right to terminate a vendor agreement, that may create flexibility: the acquirer can potentially exit an unfavorable contract, consolidate vendors to capture deal synergies, or simply use the credible possibility of termination as leverage in renegotiating commercial terms.</span></p></li></ul></li></ul></li><li><p><strong><span>Limitation of liability and indemnification</span></strong></p><ul><li><p><span>Even assuming the same contractual terms, the significance of contractual risk can change after an acquisition. A smaller company may have operated for years under a contract with unusually broad liability exposure without a dispute ever arising. But when a much larger and better-capitalized company (i.e. a better litigation target if something goes wrong) acquires the business or assumes the agreement, the practical risk assessment may change.</span></p></li><li><p><span>That does not automatically mean every contract with unfavorable liability terms must be terminated or renegotiated. But it may change the priority assigned to the contract and the remediation decision.</span></p></li></ul></li><li><p><strong><span>Restrictive covenants</span></strong></p><ul><li><p><span>Various commercial restrictions may conflict with the strategic rationale for the transaction. Examples include: exclusivity language, most-favored-nation clauses, non-competes, non-solicitation and subcontracting restrictions.</span></p></li><li><p><span>Imagine acquiring a company specifically to enter a new market, only to discover that a material contract limits the company&#8217;s ability to compete, distribute products, work with certain customers, or use particular subcontractors.</span></p></li></ul></li></ul><h4>Synthesis</h4><p>At the end of the analysis stage, the output should be more than extracted data.</p><p>Net of all of considerations,<strong> </strong>one or more internal deliverables is created: a spreadsheet, table, report or something else showing the analysis and what it means &#8212; the suggested remediation path and why. For example, if 100 contracts were in scope:</p><ul><li><p>10 may be expired</p></li><li><p>20 may require assignment consent</p></li><li><p>40 may require notice</p></li><li><p>10 may require an amendment or negotiation due to risk concerns</p></li><li><p>20 should be left alone - perhaps because they&#8217;re undesirable contracts coming up for expiration before deal close or because they&#8217;re with related entities and don&#8217;t require additional action</p></li></ul><p>That final assessment informs the decision.</p><div><hr></div><h3><strong><span>5. Decide</span></strong></h3><p>Before anyone reaches out internally or externally, there&#8217;s a need for a decision.</p><p>For an M&amp;A transaction, that often means classifying contracts into different remediation paths: assign, exit/terminate, renegotiate, renew, or leave alone.</p><p>In addition, more significant issues may also affect the transaction itself. Diligence and integration are often gray areas rather than cleanly separated phases. The significance of a contractual issue depends on when it is discovered, the magnitude of the exposure, the importance of the contract, and the broader rationale for the acquisition. A single problematic contract does not necessarily jeopardize a deal, but a population-level pattern might. </p><p>For example, perhaps a large percentage of the target&#8217;s revenue is tied to contracts that can be terminated for any reason. Or critical agreements cannot be transferred without consent. Perhaps restrictive covenants interfere with the strategic rationale for the acquisition or the company simply does not have enough durable customer contracts to provide the market beachhead the acquirer expected. Perhaps uncapped liability exists in critical places.</p><p>All of these considerations are ways that contract data can impact the overall deal.</p><h3><strong><span>6. Remediate (Take Action)</span></strong></h3><p><span>Once the decisions are made, it is time to act.</span></p><p><span>For M&amp;A and other episodic projects, this often means contract remediation and any related counterparty outreach:</span></p><ul><li><p><span>Requesting consents</span></p></li><li><p><span>Sending legally required notices or courtesy notices where appropriate</span></p></li><li><p><span>Collecting responses and following up with counterparties as needed</span></p></li><li><p><span>Negotiating where consent is difficult to obtain</span></p></li><li><p><span>Amending agreements as needed before assignment</span></p></li><li><p><span>Terminating and/or entering into new contracts where necessary</span></p></li></ul><p>All of this requires a significant workflow tracking component as well, including:</p><ul><li><p><strong>Internal communication</strong> - l<span>egal, sales, finance, </span>procurement, <span>integration teams, business owners, and senior leadership may each need different views of the same underlying analysis and remediation</span></p></li><li><p><strong>Tracking the remediation results on a contract by contract basis</strong> - w<span>hich consents are outstanding, which notices went out, which contracts are stuck, what&#8217;s blocking or impacting the deal close, etc.</span></p></li></ul><div><hr></div><h3><strong><span>Recap</span></strong></h3><p><span>Almost every post-executed enterprise contracts project can be framed as the same underlying 6-part process:</span></p><ol><li><p><strong>Trigger</strong> &#8212; an event occurs that requires contracts work</p></li><li><p><strong>Define criteria</strong> &#8212; determine what you&#8217;re looking for and why</p></li><li><p><strong>Find</strong> &#8212; apply the criteria to locate the target population of contracts</p></li><li><p><strong>Analyze</strong> &#8212; go deeper on the population: validate the initial search results, incorporate nuances, exceptions, and additional criteria, and perform a comprehensive assessment</p></li><li><p><strong>Decide</strong> &#8212; translate the analysis into recommendations, remediation paths, answers to stakeholder questions and other internal deliverables</p></li><li><p><strong>Remediate</strong> &#8212; take action through counterparty outreach, internal communication, negotiation, and workflow tracking</p></li></ol><p>Hopefully the above helps illustrate that post-executed contract intelligence is a multi-faceted problem and that each step in the process has its own challenges and nuances.</p><p><span>That leads to an interesting question: </span><strong><span data-color="#0000ff" style="color: rgb(0, 0, 255);">how much of each stage can AI actually help with? </span></strong>Or more specifically: a<span>t each stage, from trigger to remediation, what can AI reliably do, where does it fail, what inputs does it need, what are the alternatives and trade-offs, and what are the complementary capabilities that need to exist around the model to get the best results?</span></p><p><span>If you found this article helpful, subscribe to follow along and forward along to anyone who might find this useful.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The $50,000 Mistake – Why OCR Quality Can Cause LLMs to Fail on Enterprise Contracts]]></title><description><![CDATA[Why OCR quality decides what contract AI can do &#8212; and the three failure modes, ranked by how hard they are to catch.]]></description><link>https://thecontractsignal.com/p/ocr-quality-contract-ai-enterprise-contracts</link><guid isPermaLink="false">https://thecontractsignal.com/p/ocr-quality-contract-ai-enterprise-contracts</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 06 Jul 2026 00:56:25 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/481d9b60-0b6b-4511-a7fb-1a2d56f87bce_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Video version of this article: </p><div id="youtube2-zIhUhfoFevw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;zIhUhfoFevw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/zIhUhfoFevw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>Imagine a $1 million customer contract.</span></p><p><span>Five-year term. Roughly $200,000 per year. The contract includes a renewal option, but the renewal window is missed. The agreement expires. Now the business has to get a new contract in place.</span></p><p><span>Because this is an enterprise customer, that&#8217;s not as simple as sending over a quick email and getting a signature the next day. It means re-entering the contracting cycle: budget approvals, procurement, infosec review, legal review, stakeholder alignment, and all the other internal safeguards that large companies put around vendor relationships.</span></p><p><span>Let&#8217;s suppose the relationship is strong and the customer renews without major issues. But the delay pushes out a quarter of revenue.</span></p><p><span>That is roughly a $50,000 problem.</span></p><p><span>Now imagine the root cause was not a bad negotiation, a missed email, or a failure by the account team. </span></p><p><span>Imagine it was this: the contract system had the wrong renewal date because it relied on the wrong effective date. The effective date was wrong because the OCR misread the scanned contract.</span></p><p><span>This is hypothetical, but plausible. It is the kind of failure that becomes possible when post-executed contracts are converted into structured data at enterprise scale &#8211; hundreds of thousands of contracts spanning many contract types, use cases and internal systems.</span></p><p><strong><span>For a deeper breakdown of post-executed enterprise contract intelligence and complexity, see the previous article </span><a href="https://thecontractsignal.com/p/what-is-post-executed-enterprise-contract-intelligence?r=1iyonb"><span data-color="#0000ff" style="color: rgb(0, 0, 255);">here</span></a><span data-color="#0000ff" style="color: rgb(0, 0, 255);"> </span>and video<span data-color="#0000ff" style="color: rgb(0, 0, 255);"> </span><a href="https://youtu.be/3rpK6X4te38"><span data-color="#0000ff" style="color: rgb(0, 0, 255);">here</span></a><span data-color="#0000ff" style="color: rgb(0, 0, 255);">.</span></strong></p><p><span>There are three common OCR failure modes. But before getting into them, we need to unpack why OCR matters for enterprise contracts to begin with.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h3><strong><span>Why OCR Is Important: The model often does not see your contract</span></strong></h3><p><span>Enterprise contracts are often stored as PDFs. But a PDF is essentially just a container &#8211; the contents of the container (and how clean or machine-readable it is) varies widely.</span></p><p><span>Some PDFs are text-native. Some are image-only scans. Some have OCR layers that are incomplete, misaligned, or simply wrong. </span></p><p><span>If the PDF is text-native, a system may be able to extract the embedded text directly. If it is a scan, OCR has to create the text layer. Once that text layer exists, many downstream extraction and classification workflows treat it as the relevant contractual text. This is crucial to understand because </span><strong><span>a large language model usually has no built-in way to know whether the extracted text is consistent with the original page.</span></strong><span> </span></p><p><span>In other words: OCR quality can determine what the AI system thinks the contract says. So when the OCR layer is wrong, the model&#8217;s input is wrong. And a scanned contract can be perfectly legible to a human reviewer and still be unreliable for a contract AI system.</span></p><p><span>With all that in mind, there are three OCR failure modes worth separating based on both the practical impact they have and how difficult they are to spot: </span></p><ol><li><p>No text layer exists</p></li><li><p>Text layer exists but is damaged</p></li><li><p>Silent failure</p></li></ol><div><hr></div><h3><strong><span>Failure mode 1: there is no text layer</span></strong></h3><p><span>The simplest failure is when the document has no usable text layer at all.</span></p><p><span>This happens when the contract is an image-only PDF, when OCR was never run, or when some file-quality issue prevented OCR from producing usable text.</span></p><p><span>To a human reviewer, the document may look perfectly readable:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ojYp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ojYp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 424w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 848w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 1272w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ojYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png" width="1422" height="392" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:392,&quot;width&quot;:1422,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94817,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/205300182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ojYp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 424w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 848w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 1272w, https://substackcdn.com/image/fetch/$s_!ojYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F874577c2-6501-4b2b-a3f2-0e4249e80074_1422x392.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><span>To a text-based extraction pipeline, it may be a blank page.</span></p><p><span>No usable text layer means no reliable text input &#8212; at least not without additional preprocessing.</span></p><p><span>This failure mode is annoying, but relatively straightforward. You can detect that there is no text, no extracted content, or no meaningful input for the model to use.</span></p><p><span>The harder cases come when OCR did run.</span></p><div><hr></div><h3><strong><span>Failure mode 2: the text layer exists, but parts of it are damaged</span></strong></h3><p><span>The more common case is that OCR ran and got enough of the document right that the problem is easy to miss.</span></p><p><span>Two things make this difficult.</span></p><p><span>First, OCR quality is not always consistent within a single document or even within the same page. Some parts of the document can have clean, easily readable text and others can be damaged.</span></p><p><span>Second, whether any damage actually matters depends on the use case. A model may be able to classify a clause correctly even when the text contains small errors. But those same errors may be fatal if the task requires exactness. The challenge is that many contract use cases do require exactness, as we&#8217;ll see.</span></p><p><span>Here are excerpts from the same document with significantly different OCR results.</span></p><p><span>The indemnification language came through cleanly.</span></p><p><strong><span>Screenshot:</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ug9D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ug9D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 424w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 848w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 1272w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ug9D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png" width="770" height="332" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:332,&quot;width&quot;:770,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:130092,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/205300182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ug9D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 424w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 848w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 1272w, https://substackcdn.com/image/fetch/$s_!ug9D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726e017f-59d6-4e97-8c93-8315c4b1e47e_770x332.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>OCR:</span></strong></p><p><span>(C) Each Party agrees to defend, protect, indemnify and hold harmless each other Party from and against all claims and demands, including any action or proceeding brought thereon, and all costs, losses, expenses and liabilities of any kind relating thereto, including reasonable attorneys fees and cost of suit, arising out of or resulting from any construction activities performed or authorized by such indemnifying Party; provided however, that the foregoing shall not be applicable to either events or circumstances caused by the negligence or willful act or omission of such indemnified Party, its licensees, concessionaires, agents, servants, employees, or anyone claiming by, through, or under any of them, or claims covered by the release set forth in 5.4(D).</span></p><p><span>But a few pages away, the term language picked up OCR errors:</span></p><p><span>ARTICLE VII</span></p><p><span>TERM</span></p><p><strong>Screenshot:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IZUE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IZUE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 424w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 848w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 1272w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IZUE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png" width="658" height="436" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:436,&quot;width&quot;:658,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:138654,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/205300182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IZUE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 424w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 848w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 1272w, https://substackcdn.com/image/fetch/$s_!IZUE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6b4468c-5b89-4c94-a37e-c933773f9e8a_658x436.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><strong><span>OCR:</span></strong></p><p><span>7.1 Term of this OEA. This OEA shall be effective as of the date first above written and shan continue in full force and effect until 11:59 p.m. on December 31, 2049; provided, however, that the easements referred to in Article II hereof which are specified as being perpetual or as continuing beyond the term of this OEA shall continue in force and effect as provided therein. Upon termination of this OEA, an rights and privileges derived from and all duties and obligations created and imposed by the provisions of this OEA, except as relates to the easements mentioned above, shan terminate and have no further force or effect; provided, however, that the termination of this OEA shan not limit or affect any remedy at law or in equity that a Party may have against any other Party with respect to any liability or obligation arising or to be performed under this OEA prior to the date of such termination.</span></p><p><span>&#8220;Shan continue in full force and effect&#8221; is damaged text, but it is still recognizably term language.</span></p><p><span>So if the task is clause classification &#8212; does this document contain a term provision? &#8212; the OCR may still be good enough for the model to classify it correctly. The model has plenty of semantic context to work with.</span></p><p><span>But contrast that with data points where exactness is crucial: effective dates, expiration dates, signature dates, notice periods, renewal deadlines, monetary amounts and party names.</span></p><p><span>Here is a preamble where the date was handwritten into a blank:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wMSS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wMSS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 424w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 848w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 1272w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wMSS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png" width="778" height="458" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:458,&quot;width&quot;:778,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:123887,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/205300182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wMSS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 424w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 848w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 1272w, https://substackcdn.com/image/fetch/$s_!wMSS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d889f19-07a9-461f-a10e-691a20a9baa3_778x458.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><span>The OCR output looked like this:</span></p><p><span>THIS AMENDED AND RESTATED OPERATION AND EASEMENT AGREEMENT (&#8221;OEA&#8221;) is made and entered into as of the a&lt;g day of (j)Q.\uber, 1999, between DAYTON HUDSON CORPORATION, a Minnesota corporation (&#8221;Target&#8221;) and Northway Mall Associates, a New York partnership (&#8221;Developer&#8221;).</span></p><p><span>The handwriting distorted the OCR, and the date came out as junk. A date that cannot be parsed is not usable.</span></p><p><span>Interestingly, running the same page through a different OCR engine produced a completely different &#8212; and correct &#8212; result:</span></p><p><span>THIS AMENDED AND RESTATED OPERATION AND EASEMENT AGREEMENT (&#8221;OEA&#8221;) is made and entered into as of the 28 day of October, 1999, between DAYTON HUDSON CORPORATION, a Minnesota corporation (&#8221;Target&#8221;) and Northway Mall Associates, a New York partnership (&#8221;Developer&#8221;).</span></p><p><span>Same page. Same underlying contract. Two different OCR outputs. One incorrect, one correct.</span></p><p><span>This is not a rare edge case. Post-executed enterprise contract repositories often contain scans, handwritten dates, signatures or notes, and documents of widely varying age and quality. OCR quality varies by engine (the type of OCR software), by file type, by scan quality, by page, and sometimes by individual field.</span></p><p><span>That is a recurring theme regarding why post-executed enterprise contract intelligence is so hard &#8211; the multi-step, document-to-data pipeline. Model selection is just one step.</span></p><p><span>Still, in this second failure mode, the damage is at least reasonably easy to spot: &#8220;Shan&#8221; is not a word and &#8220;a&lt;g day of (j)Q.\uber&#8221; is not a date.</span></p><p><span>A QC layer, parser diagnostics and observability, OCR confidence scores, or even a careful human glance at the extracted text may flag the problem.</span></p><p><span>Which brings us to the failure mode that is the most difficult to catch.</span></p><div><hr></div><h3><strong><span>Failure mode 3: the text is plausible &#8212; and wrong (aka silent failure)</span></strong></h3><p><span>The most dangerous OCR failure is when the output looks clean.</span></p><p><span>No garbage characters. No obvious misspellings. No broken date format.</span></p><p><span>Just a value that happens to be incorrect.</span></p><p><span>A handwritten 7 gets read as a 1. A 3 gets read as an 8. A $100,000 value gets misread as $1,000,000 &#8212; or the other way around.</span></p><p><span>Here is a real example &#8212; a signature block with a handwritten date:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sXnt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sXnt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 424w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 848w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 1272w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sXnt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png" width="890" height="524" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3cff6eda-839e-4991-a67e-6cf624844067_890x524.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:524,&quot;width&quot;:890,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:80310,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/205300182?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sXnt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 424w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 848w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 1272w, https://substackcdn.com/image/fetch/$s_!sXnt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3cff6eda-839e-4991-a67e-6cf624844067_890x524.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><span>The OCR output appears clean:</span></p><p><span>AGREED AND ACCEPTED:</span></p><p><span>PLITT THEATRES, INC.</span></p><p><span>By: [Signature] Henry Plitt, Chairman of the Board</span></p><p><span>Date: Oct 12 1982</span></p><p><span>But look closely at the image. The day appears to be 13, not 12.</span></p><p><span>The likely reason is that the lower part of the handwritten digit overlapped with the signature line, causing the OCR to read the 3 as a 2.</span></p><p><span>This situation is much more dangerous than junk text.</span></p><p><span>Junk text tells you something went wrong so you can course correct as needed.</span></p><p><span>A date that is plausible but wrong does not throw up any red flags. In this case, &#8220;Oct 12 1982&#8221; is a valid date &#8211; it has the right format and is semantically plausible.</span></p><p><span>In other words, it is a silent failure.</span></p><p><span>And silent failures are the ones that create the most risk in contract AI systems.</span></p><h3><strong><span>Why small OCR errors become business problems</span></strong></h3><p><span>You might look at the example above and say &#8211; it&#8217;s only one day off, that won&#8217;t have much impact.</span></p><p><span>Maybe. But I&#8217;d call out two things:</span></p><p><span>1. Enterprise contracts often have more significant silent OCR failures such as a wrong month or wrong year (note that the example above is from a publicly available contract). This becomes more likely as hundreds of thousands of documents present many opportunities for silent OCR failures.</span></p><p><span>2. The broader point, however, is that a single, seemingly innocuous bad value rarely stays contained to one situation or use case and can have downstream consequences.</span></p><p><span>For example, a renewal date or expiration date may not be written explicitly in the document. It may be computed from other extracted fields: effective date plus initial term, potentially plus a renewal period minus a notice window.</span></p><p><span>If OCR silently corrupts the effective date, the computed expiration date may be wrong.</span></p><p><span>If OCR silently corrupts the term length, the renewal calculation may be wrong.</span></p><p><span>If OCR silently corrupts the notice period, the counterparty outreach or operational deadline may be wrong.</span></p><p><span>If OCR silently corrupts the contract value, the agreement may be routed to the wrong workflow, wrongfully excluded from a report, or treated as below a threshold that should have triggered additional review.</span></p><p><span>And the model may have no way to know, because every input it was handed looked valid on the surface.</span></p><p><span>That is how a small OCR problem becomes a contract intelligence problem.</span></p><p><span>This is also why scale changes the risk profile. In a small set of contracts, these errors may be annoying but manageable. In an enterprise repository with tens or hundreds of thousands of documents, even a low error rate can create meaningful downstream exposure.</span></p><p><span>OCR quality is not just a preprocessing issue. It is a contract data quality issue.</span></p><h3><strong><span>What can be done</span></strong></h3><p><span>There are various approaches to either reducing or mitigating issues that can come up from OCR or file quality issues. This discussion could easily be its own article.</span></p><p><span>At the OCR layer - some common approaches involve checking for existing OCR, re-running OCR software, testing different engines and capturing OCR confidence where available.</span></p><p><span>At the extraction layer - high-impact contracts or fields can be routed into more robust QC workflows.</span></p><p><span>At the QC or validation layer - rules can catch what is catchable: for example values outside plausible ranges.</span></p><p><span>In addition, evolving AI model capabilities may help with some of these as well.</span></p><h3><strong><span>Recap and the open question</span></strong></h3><p><span>To recap, there are three OCR failure modes worth separating:</span></p><p><span>1. </span><strong><span>No text layer at all</span></strong><span> &#8212; the document is readable to a human but may be blank or illegible to a text-based extraction system. This is usually relatively easy to detect by a well-designed system.</span></p><p><span>2. </span><strong><span>Visible OCR damage</span></strong><span> &#8212; the text layer exists, but contains obvious problems like non-words, junk dates, or corrupted phrases. This can also often be detected, and whether it matters depends on the use case.</span></p><p><span>3. </span><strong><span>Silent OCR damage</span></strong><span> &#8212; the text looks plausible, but is wrong. This is the most dangerous because it can flow through the pipeline undetected.</span></p><p><span>The open question is empirical: in a real population of post-executed enterprise contracts, how often does each failure mode occur? And how much does each one affect downstream contract AI accuracy?</span></p><p><span>That is a type of measurement I have not seen yet &#8212; and one that matters if contract AI is going to move from impressive demos to trusted enterprise systems.</span></p><p><span>Because the real issue is not simply whether an AI model is powerful enough.</span></p><p><span>It is whether the system has given the model the right contractual text to read in the first place.</span></p><p><span>If you found this useful, subscribe to The Contract Signal &#8212; and forward it to anyone who might find it helpful.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div class="poll-embed" data-attrs="{&quot;id&quot;:722721}" data-component-name="PollToDOM"></div><p></p>]]></content:encoded></item><item><title><![CDATA[What Is Post-Executed Enterprise Contract Intelligence?]]></title><description><![CDATA[The three ideas behind it &#8212; post-execution, enterprise scale, and contract intelligence &#8212; and why your already-signed agreements are full of untapped risk and value.]]></description><link>https://thecontractsignal.com/p/what-is-post-executed-enterprise-contract-intelligence</link><guid isPermaLink="false">https://thecontractsignal.com/p/what-is-post-executed-enterprise-contract-intelligence</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 28 Jun 2026 18:09:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b6182cd7-297a-4718-8c37-a97be61f4236_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>I spent 10 years working on </span><strong><span>post-executed enterprise contracts intelligence.</span></strong><span> This is a loaded concept, so let&#8217;s break it down into three parts:</span></p><ol><li><p><strong><span>Post-executed</span></strong><span> &#8211; these are contracts that have already been signed or executed, otherwise known as post-signature contracts, as compared to pre-execution or pre-signature contracts that have been drafted, negotiated or redlined but not signed by all parties.</span></p></li><li><p><strong><span>Enterprise </span></strong><span>&#8211; at least one party to the contract is a very large business, such as a F500 company.</span></p></li><li><p><strong><span>Contract Intelligence</span></strong><span> &#8211; enabling people to make legal and business decisions by turning contracts into structured data that can be searched, filtered, reported on, and analyzed at scale.</span></p></li></ol><p><span>Let&#8217;s dig into this in greater depth &#8211; including how these concepts influence one another.</span></p><p><span>Prefer to watch? Here's the video version: </span></p><div id="youtube2-3rpK6X4te38" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;3rpK6X4te38&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/3rpK6X4te38?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><h2><strong><span>Deeper Dive</span></strong></h2><h3><strong><span>Post-executed</span></strong></h3><p><span>First, we should dig into the implications of what it actually means for a contract to be post-executed</span><strong><span>.</span></strong></p><p><span>Contracts are written agreements containing the terms of an exchange, including legal obligations on both sides.</span><strong><span> </span></strong></p><p><span>Post-executed means the contract is already signed. This means we&#8217;re past the drafting stage, negotiation, and redlining &#8212; the language has been finalized and agreed to. In other words, pending an amendment to the original contract or another legal mechanism, the language is locked, and the rights and obligations in effect.</span></p><div><hr></div><h3><strong><span>Enterprise</span></strong></h3><p><span>Being a large enterprise has many significant implications for contracts as well.</span></p><p><strong><span>Range and volume of contract types</span></strong></p><p><span>One implication is that the volume of contracts is much, much higher for large enterprises than it is for smaller businesses.</span></p><p><span>This happens for a few reasons.</span></p><p><span>First, enterprises are large, diversified companies, often with multiple business units, many customers and vendors, many products and services, and a more complicated supply chain. For example, everything ranging from the need to manufacture things to the need to distribute products and services at scale means that there are many more significant business exchanges and many corresponding opportunities for contracts across a range of contract types. </span></p><p><span>These may include manufacturing agreements, reseller or distribution agreements, licensing agreements, partnership agreements, leases, and other types of customer and vendor contracts.</span></p><p><strong><span>Contract Families</span></strong></p><p><span>Another implication is that because enterprises have been around for a long time, they have business relationships and contracts from many different periods. </span></p><p><span>This includes older contracts with longer lifecycles, often including multiple rounds of contracting between entities. These relationships may include more than just an initial master agreement. They may also include NDAs, amendments, statements of work, order forms, addenda, and other related documents. </span></p><p><span>As a result, large enterprises can have very large families of related contracts.</span></p><p><strong><span>Complexity of individual contracts</span></strong></p><p><span>Relatedly, the complexity of the contracts themselves tends to be very high for enterprises. Large enterprises have all sorts of complicated requirements built into their contracting practices.</span></p><p><span>This is partly due to being subject to different regulations as a result of their broad reach of business activity, often spanning geographies and business lines. Also, because they have been in business for a long time, different laws, rules and conventions may have applied over different periods of time.</span></p><p><span>It&#8217;s also because they are at greater risk of litigation, leading to more robust risk allocation measures. This is driven in part by increased opportunities for something to go wrong, given their higher levels of business activity and transaction volume. It is also driven by the fact that large enterprises often have more money and assets to pursue, making them more attractive litigation targets relative to smaller entities. </span></p><p><span>As a result, they often have to be more aggressive about risk allocation. This can lead to lengthier, more detailed provisions that limit liability exposure in different circumstances.</span></p><p><span>Net of these considerations, the population of contracts at a given enterprise is highly mixed. It is not homogeneous or standardized.</span></p><p><span>Even the same type of contract may have many different variables and complexity drivers baked in.</span></p><p><span>That gets compounded further by other complexity drivers, such as:</span></p><ol><li><p><strong>File quality.</strong> Many of these documents include emails, scans, pages with coffee stains, handwriting in the margins, or image-only PDFs. Each of these can reduce the quality and reliability of OCR (optical character recognition). Bad OCR means the text may not be properly recognized by AI models. So before any AI touches the contract, the file often has to go through a chain of preprocessing steps just to become usable. Example here: <a href="https://www.industrydocuments.ucsf.edu/docs/yfdc0172">link</a></p></li><li><p><strong>Multi-document files.</strong> One file will often contain multiple discrete parts. For example, you may have the core contract, cover pages, exhibits, addenda, additional terms and conditions, and other related documents all in one file. Determining which document or clause takes precedence can be a very challenging exercise as a result. Example here: <a href="https://www.sec.gov/Archives/edgar/data/820609/000119312510160027/dex102.htm">link</a></p></li><li><p><strong>No system of record.</strong> Enterprises are diversified, with many business units, siloed repositories and multiple custodians. As a result, there is often no central place where all contracts live. Companies frequently can&#8217;t even find their contracts, let alone analyze them. And when they do find them, they are often dealing with legacy systems that contain partially executed agreements, duplicates or overlapping documents.</p></li></ol><p><strong><span>Compounding Complexity</span></strong></p><p><span>Each of these factors would create complexity on its own. Enterprises often deal with all of them at once, across hundreds of thousands of documents.</span></p><p><span>So there is a range of complexity drivers, and many of them are closely linked to large enterprises. That&#8217;s why the enterprise dimension is so important and why it has so many implications for contracts.</span></p><p><span>Net of these considerations, you do not have one big, centralized, clean dataset. You have many unstructured, complex, messy, difficult-to-find documents &#8212; subpopulations inside subpopulations, fragmented across many systems.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h3><strong><span>Contract Intelligence</span></strong></h3><p><span>Collectively, this is what creates the need and opportunity for contract intelligence.</span></p><p><span>Even finding contracts in an enterprise can be very challenging, let alone knowing </span><em><span>the substance of</span></em><span> what is in those contracts and what obligations they contain.</span></p><p><span>And this matters because hidden within these already-executed contracts are collectively huge opportunities, both in terms of risk allocation and mitigation, as well as increasing business value.</span></p><p><span>For example, it is not unheard of for a contract to lack a limitation of liability clause, which could introduce significant liability for the business, particularly if it is an older contract without other safeguards in place.</span></p><p><span>Payment terms are often suboptimal, leaving value behind or creating misalignment with the company&#8217;s standards, internal finance systems, and processes.</span></p><p><span>There may also be untapped opportunities to increase prices, or rights and restrictions to assignment, termination or exiting agreements.</span></p><p><span>The goal of contract intelligence is to take all of these contract documents and be able to search, filter, report on, and analyze them at scale.</span></p><p><span>The data points of interest vary based on the use cases and the business or legal problems people are trying to solve. </span></p><p><span>For example, in M&amp;A diligence and integration, you might be interested in assignment, change of control, governing law, and related provisions. </span></p><p><span>For a business-as-usual type of review where a question comes up from the commercial team, you may be interested in a specific clause, such as the indemnification language agreed to in previous similar situations, or how long a contract requires the company to support a product that is being sunset.</span></p><p><span>This is a very small sample of potential use cases. The relevant questions vary depending on the business line, industry, function, department, specific situation, and place within the larger supply chain.</span></p><p><strong><span>Contract Intelligence vs. Contract Summarization</span></strong></p><p><span>For sake of clarity, it is important to understand that the difference between contract intelligence and a contract summary. These are complementary use cases that serve different purposes.</span></p><p><span>Contract intelligence specifically relates to turning the natural language of contracts into a structured data asset at scale, so users can gain visibility into their contract data through search, filters, reports and analytics.</span></p><p><span>In contrast, with a contract summary, you are typically doing a deeper dive into one or more contracts, often at smaller volume, to better understand the underlying nuances.</span></p><p><span>The critical difference is that the output involves qualitative descriptions of text. These are not ideal for running searches, reports and analytics at scale. </span></p><p><span>However, summaries can be very helpful when creating a diligence memo or other report to support a decision across a contract population, such as whether to move forward with an acquisition or how to prioritize regulatory remediation.</span></p><div><hr></div><h2><strong><span>Summary</span></strong></h2><p><span>To summarize, post-executed enterprise contract intelligence involves three main components:</span></p><ol><li><p><strong>Post-executed</strong> &#8211; contracts that have already been signed or executed, otherwise known as post-signature contracts, as compared to pre-execution or pre-signature contracts that have been drafted, negotiated or redlined but not signed by all parties.</p></li><li><p><strong>Enterprise </strong>&#8211; at least one party to the contract is a very large business, such as a Fortune 500 company.</p></li><li><p><strong>Contract Intelligence</strong> &#8211; enabling people to make legal and business decisions by turning contracts into structured data that can be searched, filtered, reported on, and analyzed at scale.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p></li></ol><p><span>If you found this article helpful, subscribe to follow along and forward it to anyone you think might find it useful.</span></p>]]></content:encoded></item><item><title><![CDATA[Everyone in Contract AI Suddenly Has a “Benchmark.” Here’s How to Read One.]]></title><description><![CDATA[Four different things "benchmark" can mean in contracts AI &#8212; and how to tell real evidence from marketing.]]></description><link>https://thecontractsignal.com/p/how-to-read-contract-ai-benchmark</link><guid isPermaLink="false">https://thecontractsignal.com/p/how-to-read-contract-ai-benchmark</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 21 Jun 2026 22:16:07 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/249aac98-eff0-418c-a930-104fce3bcc30_2400x1256.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>A few weeks after I launched this newsletter &#8212; </span><em><span>The Contract Signal</span></em><span> &#8212; TermScout launched a &#8220;Contract Signals Report.&#8221;</span></p><p><span>I&#8217;ll take the overlap as validation. The framing works in part because contracts are becoming more than documents to store, negotiate, and search. Once you parse through the dense language, contracts start to look like signals: of risk, leverage, governance, market dynamics, opportunity, and operational behavior.</span></p><p><span>But the naming overlap is also a small symptom of a larger market pattern.</span></p><p><span>The volume of contract AI news keeps accelerating. Every week brings another product announcement, research release, workflow claim, or market report. And in that stream of news, one word keeps coming up: </span><strong><span>benchmark</span></strong><span>.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><p><span>Crosby published a contract-redlining benchmark. Harvey has published benchmarks for contract understanding and legal agents. TermScout, LexisNexis, and others benchmark contract terms against market norms. CLM vendors benchmark cycle time and operational performance.</span></p><p><span>One word &#8212; benchmark &#8212; is taking on multiple meanings.</span></p><p><span>Sometimes &#8220;benchmark&#8221; means a test. Sometimes it means a dataset. Sometimes it means the results from a market survey or a dashboard metric summarizing the same to compare against industry peers.</span></p><p><span>It has quietly become a word that sounds precise but is often focused on marketing.</span></p><p><span>So when (or before) the next benchmark lands in your inbox, the first move is simple: ask what kind of benchmark you are actually looking at.</span></p><p><span>Benchmark of what? Against what? Scored by whom? Using whose data? And for whose benefit?</span></p><div><hr></div><h3><strong><span>What the word actually means</span></strong></h3><p><span>Strip away the jargon and a benchmark is just a reference point. A fixed thing you measure against &#8212; a baseline, a starting line &#8212; so that any other number means something relative to that initial reference point.</span></p><p><span>You set the reference, then you compare: better than the benchmark, worse than the benchmark, by how much.</span></p><p><span>Contracts professionals already had a specific version of this long before AI showed up. In outsourcing and other long-term agreements, a benchmarking clause lets a customer bring in an independent third party to compare a supplier&#8217;s pricing and service levels against the wider market &#8212; and adjust the deal if the supplier has drifted off-market.</span></p><p><span>Note what made that mechanism trustworthy: an independent party, an agreed method, and a defined market to compare against.</span></p><p><span>In the AI era, the word has stretched to cover at least four genuinely different activities. They get announced with the same or similar vocabulary, which is exactly why they are hard to read.</span></p><p><span>The first evaluates models. The second evaluates contract positions (i.e. the values of given clauses or other data points). The third evaluates operational efficiency. The fourth evaluates AI or system improvement over time.</span></p><p><span>Separating these concepts can help provide clarity.</span></p><div><hr></div><h3><strong><span>The four benchmarks hiding under one word</span></strong></h3><p><strong><span>1. Performance benchmarks &#8212; &#8220;How good is the AI?&#8221;</span></strong></p><p><span>This is the one generating the loudest headlines: take a task, have AI models do it, and score them against a human expert or a model answer.</span></p><p><span>One recent example is Crosby&#8217;s Redline Bench, which scored frontier models on realistic SaaS-contract redlining against attorney-authored &#8220;golden&#8221; responses. The reported results put the top model around 50% and the rest clustered below it &#8212; a narrow spread, with humans still better at finding new routes to resolution while models tended to anchor on their opening positions.</span></p><p><span>Conceptually, there is nothing wrong with this. Measuring model performance against domain experts in realistic conditions is a reasonable thing to do. To their credit, the better recent benchmarks in this category disclose a lot: the use cases, the scoring dimensions, the rubric, and in some cases the whole thing is open for other labs to run.</span></p><p><span>That is valuable as it helps practitioners assess actual performance.</span></p><p><span>The trouble starts with how the numbers get used downstream.</span></p><p><span>&#8220;Model X scored 50.5%&#8221; gets repeated far from the methodology that produced it. Without that context, the number is almost impossible to interpret.</span></p><p><span>It may be directionally interesting. It may tell you something about frontier-model behavior in a simulated negotiation. It may help compare models under one defined test design.</span></p><p><span>But it does not automatically tell you which product to buy, which workflow to automate, or whether the same system will work inside your contract population.</span></p><p><span>That distinction matters.</span></p><p><span>A multi-turn negotiation benchmark is not the same thing as a data extraction benchmark. A redlining benchmark is not the same thing as a playbook-compliance benchmark. A test of SaaS agreement negotiation is not the same thing as a test of high-volume vendor onboarding, procurement triage, or post-signature obligation extraction.</span></p><p><span>And in many contract AI workflows, the advanced task depends on the foundational layer being right first.</span></p><p><span>Before a system can negotiate from a contract, it often has to identify the agreement, extract the relevant provisions, understand the parties, map the clause to the right playbook position, and preserve the context across the workflow. If those earlier steps are noisy, the later benchmarked task may inherit the errors. Sometimes the problem is not only model performance. It is error propagation &#8211; the snowballing or compounding of small mistakes at different stages in the process until the end result is a far cry from the original performance findings.</span></p><p><span>That is why the headline score is not enough.</span></p><p><span>Did the model miss obvious issues? Did it identify the right concern but propose a weaker redline? Did it produce something commercially reasonable but different from the attorney-authored answer? Was it penalized for not matching the exact path the human took? Would two senior attorneys have agreed on the same &#8220;perfect&#8221; answer?</span></p><p><span>A performance benchmark can be useful. But the headline score is not the benchmark. The benchmark is a combination of the task, the dataset, the scoring method, and the judgment calls underneath the score.</span></p><div><hr></div><p><strong><span>2. Norm benchmarks &#8212; &#8220;Is my contract typical?&#8221;</span></strong></p><p><span>A completely different activity wears the same label.</span></p><p><span>Here, you assemble a large body of contract data, derive what is standard for a given type of deal, and then compare a new agreement against that norm or standard.</span></p><p><span>LexisNexis&#8217;s Market Standards, </span>TermScout&#8217;s reports, <span>and various &#8220;state of contracts&#8221; reports live in this category. The output is not a model score. It is a reference set.</span></p><p><span>For example: for vendor agreements of this type, in this industry, at this deal size, mutual indemnification is standard; one-way indemnification is aggressive; this limitation of liability formulation is common; that AI-use restriction is becoming more frequent.</span></p><p><span>This is genuinely useful and very doable.</span></p><p><span>Most of it reduces to structured data. A clause like indemnification can be reduced to a small set of positions &#8212; both parties indemnified, one party indemnified, neither party indemnified &#8212; and once you have coded enough contracts and broken down each clause type sufficiently, you can say how common each position is. Control for industry, deal size, template source, whether the agreement is customer paper or vendor paper, and other criteria and you have a real reference point for negotiation and risk.</span></p><p><span>But norm benchmarks have a problem that rarely makes the announcement: representativeness.</span></p><p><span>Much of this analysis is built on public filings, and public contracts are a biased sample by construction. A contract gets filed with the SEC because it crosses a materiality threshold. That means the corpus skews toward large, heavily negotiated agreements that may look very different from the thousands of ordinary commercial contracts a business actually runs on.</span></p><p><span>Even the filed agreements are frequently redacted at the commercially sensitive points you would most want to benchmark.</span></p><p><span>I used SEC contracts myself in last week&#8217;s article on creating a contract data extraction benchmark, and the same caveat applied there: they are a convenient, legitimate starting point. They are not a stand-in for your private contract population.</span></p><p><span>A norm built on a skewed sample is a useful hint, not ground truth.</span></p><div><hr></div><p><strong><span>3. Process benchmarks &#8212; &#8220;How fast is my operation?&#8221;</span></strong></p><p><span>A third meaning has nothing to do directly with model quality or contract terms at all.</span></p><p><span>It is operational: cycle time, turnaround, throughput, percentage of contracts on standard templates, time-to-signature, number of negotiation rounds, fallback frequency, legal touch rate, approval bottlenecks.</span></p><p><span>CLM vendors increasingly sell &#8220;benchmarking&#8221; dashboards that compare your contracting operation against peers.</span></p><p><span>This is often the management-consulting meaning of benchmark, applied to legal ops. It may also be the most actionable version for many teams.</span></p><p><span>If your average NDA takes 12 days and comparable teams complete theirs in three, that is useful information to have. By the same token, if 80% of your low-risk agreements still require legal review while similar companies route most of these through self-service or your contracting cycle time spikes in one region or business unit, or by contract type, that is useful too.</span></p><p><span>But it is worth separating this cleanly from the other meanings.</span></p><p><span>A vendor saying it can &#8220;benchmark your contracts&#8221; might mean it can compare clause language to market norms. It might mean it can compare your contracting process to peer operations. It might mean it can measure model performance. Those are different use cases with different products, solutions, and datasets.</span></p><p><span>Process benchmarks are valuable when the metric maps to a real operational decision.</span></p><p><span>They are less valuable when the metric becomes theater: a dashboard showing you are slower than &#8220;peers&#8221; without explaining who the peers are, what work was included, or whether the comparison reflects your risk profile, company stage, industry, deal size, or contracting model.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong><span>4. Baseline benchmarks &#8212; &#8220;Is my system getting better?&#8221;</span></strong></p><p><span>The fourth meaning is the one practitioners building AI systems use most, and the one that almost never makes the press release.</span></p><p><span>Here, a benchmark is an internal baseline.</span></p><p><span>You fix a starting level of performance &#8212; an out-of-the-box model, a first prompt, a v1 pipeline, an existing review process &#8212; and then measure every later change against it.</span></p><p><span>Did the new prompt help? Did the more expensive model actually earn its cost? Did retrieval improve accuracy or just add latency? Did fine-tuning move the metric or just the vibes? Did the new clause taxonomy improve consistency? Did the latest model version break something that used to work?</span></p><p><span>This is how serious AI and data teams actually work.</span></p><p><span>And it carries two critical ideas the headline benchmarks tend to gloss over but often have a huge impact on assessing the benchmarks and making them actionable.</span></p><ol><li><p><span>The first is generalization. A model trained or tuned on one contract population can fall apart on a new population due to different drafting, file quality, document structures and conventions used.</span></p></li><li><p><span>The second is data drift. A system that scored well last quarter can quietly degrade as the documents flowing through it change.</span></p></li></ol><p><span>A single published score is a photograph: one snapshot in time, taken under a defined set of assumptions.</span></p><p><span>Truly effective benchmarking is more like a video &#8211; multiple points in time, multiple snapshots, multiple angles, and enough continuity to tell whether the system is improving, degrading or staying flat.</span></p><p><span>And the video is what tells you whether something works in production.</span></p><p><span>Four meanings. Four different questions. One word.</span></p><p><span>The moment you hear &#8220;benchmark,&#8221; your first move should be to figure out which one you are actually being sold.</span></p><div><hr></div><h3><strong><span>How to read any benchmark</span></strong></h3><p><span>Sorting the type gets you halfway there. The rest is actually performing the analysis.</span></p><p><span>There are three layers that rarely get discussed and almost always decide whether a number means anything: </span><strong><span>the use cases, the scoring and the incentives.</span></strong></p><p><strong><span>The use cases</span></strong></p><p><span>First, what was actually tested?</span></p><p><span>&#8220;Contract review&#8221; can mean redlining a standard NDA against a clear playbook. It can also mean a no-playbook judgment call on a bespoke, heavily negotiated agreement with limited commercial context.</span></p><p><span>Those are wildly different tasks.</span></p><p><span>Then, the specific scenarios corresponding to the use cases.</span></p><p><span>Before you trust a performance number, you want to know the distribution of difficulty behind it. Was the benchmark mostly easy cases? Edge cases? Common issues? Rare but high-risk provisions? First-pass issue spotting? Full redlining? Negotiation strategy? Escalation judgment?</span></p><p><span>That is the difference between &#8220;the AI is good&#8221; and &#8220;the AI is good at the easy version.&#8221;</span></p><p><strong><span>The scoring</span></strong></p><p><span>How was &#8220;right&#8221; decided?</span></p><p><span>This is where legal benchmarks are genuinely hard, because you are scoring dense, qualitative prose, not simple multiple-choice.</span></p><p><span>Did they grade it as a simple classification problem &#8212; risk flagged correctly or not, one point each over the total possible? Or did they use a rubric that accounts for subjectivity, ambiguity, commercial reasonableness, drafting quality, and the context the model was given?</span></p><p><span>And what does 100% mean?</span></p><p><span>Does it mean the model matched the attorney answer exactly? If so, would two senior attorneys have agreed 100% of the time?</span></p><p><span>Often they would not.</span></p><p><span>Law is full of gray areas. That is why it needs lawyers. A scoring method that pretends the gray areas are black and white will produce a confident number that does not survive contact with reality.</span></p><p><span>In addition, watch for false precision.</span></p><p><span>A score reported to a tenth of a percent &#8212; 50.5, 45.1, 44.4 &#8212; signals rigor. But the decimal is only as meaningful as the rubric underneath it.</span></p><p><span>When the rubric is scoring subjective legal judgment, three-significant-figure precision is often decoration. The confident number is the easiest part to produce and the least informative.</span></p><p><strong><span>The incentives</span></strong></p><p><span>Who ran the benchmark or experiment, and what do they sell?</span></p><p><span>This is the layer that matters most and gets discussed least.</span></p><p><span>Charlie Munger is often quoted as saying: &#8220;Show me the incentive and I will show you the outcome.&#8221;</span></p><p><span>A private sector benchmark is almost never published for philanthropy.</span></p><p><span>A law firm has an interest in a benchmark that shows legal judgment remains hard to automate. A vendor has an interest in a benchmark that makes its workflow look measurable and differentiated. A model lab has an interest in a benchmark that defines progress around the capabilities its model performs well on.</span></p><p><span>None of that means the benchmark is dishonest.</span></p><p><span>It means the benchmark has a point of view.</span></p><p><span>Wherever judgment enters &#8212; which tasks to include, how to score gray areas, what to call 100%, which dataset to use, which comparison group to define &#8212; the benchmark will reflect the author&#8217;s theory of the problem.</span></p><p><span>Even a scrupulously transparent benchmark is shaped by what its author chose to measure in the first place.</span></p><p><span>Transparency lets you check the math. It does not remove the incentive that selected the problem.</span></p><div><hr></div><h3><strong><span>How to spot a benchmark worth trusting</span></strong></h3><p><span>None of this is a reason to dismiss benchmarks.</span></p><p><span>They are one of the better things to happen to a field that ran on vibes and demos for years. But they need to be read like evidence, not like slogans.</span></p><p><span>To summarize, a benchmark you can actually lean on tends to disclose:</span></p><ul><li><p><strong><span>Scope</span></strong><span> &#8212; the tasks tested, how hard they were, and whether the headline average hides meaningful variation.</span></p></li><li><p><strong><span>Method</span></strong><span> &#8212; how &#8220;correct&#8221; was scored, who judged, and what a perfect score would mean.</span></p></li><li><p><strong><span>Data provenance</span></strong><span> &#8212; where the corpus came from and how representative it is of the contracts you care about.</span></p></li><li><p><strong><span>Authorship and incentive</span></strong><span> &#8212; who built it, who funded it, and what they sell.</span></p></li><li><p><strong><span>Reproducibility</span></strong><span> &#8212; whether anyone else can run it and get the same answer.</span></p></li><li><p><strong><span>Business relevance</span></strong><span> &#8212; whether the metric maps to a decision you actually need to make or problem you need to solve.</span></p></li></ul><p><span>Use that as a checklist.</span></p><p><span>A benchmark that discloses its scope, method, and data but is published by a party with an obvious stake in how it is used is not worthless. It just needs to be read with the stake in view.</span></p><p><span>A benchmark with a strong headline but no corpus, no rubric, no task design, no scoring explanation, and no discussion of limitations is not reliable. It is a claim dressed up as a benchmark, often for marketing purposes.</span></p><div><hr></div><h3><strong><span>The through-line</span></strong></h3><p><span>Last week I made the case that contract AI looks magical on one document and gets hard across a population &#8212; that the demo and the workflow are different problems.</span></p><p><span>A benchmark is the same story one level up.</span></p><p><span>The score is the demo. The methodology is the workflow.</span></p><p><span>A number that travels without its context is a promise without the receipts.</span></p><p><span>That is why independent reading of this market matters.</span></p><p><span>Independent does not mean neutral or detached. I have a point of view, shaped by a decade in enterprise contracts, taxonomy design, structured data and extraction workflows, AI, and legaltech product work.</span></p><p><span>It means something narrower and more important: I do not have a platform, model, law firm service, or benchmark to sell inside the analysis.</span></p><p><span>When the company publishing the benchmark also benefits from the way the benchmark is framed, the right response is not cynicism. It is professional skepticism.</span></p><p><span>What was tested? Against what? By whom? On whose data? And who benefits from the answer?</span></p><p><span>More signals soon.</span></p><p><span>If you found this useful, subscribe to follow along, forward it to someone who is about to make a tooling decision off a benchmark headline, and tell me which contract AI terms or claims you want me to unpack next.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[What Happens When You Ask AI to Extract Contract Metadata at Scale]]></title><description><![CDATA[A first benchmark on the gap between one-contract demos and contract-population workflows.]]></description><link>https://thecontractsignal.com/p/what-happens-when-you-ask-ai-to-extract</link><guid isPermaLink="false">https://thecontractsignal.com/p/what-happens-when-you-ask-ai-to-extract</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 15 Jun 2026 02:24:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d55d5138-4580-4ba2-83e1-1e3531c5db52_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>The gap nobody demos</strong></p><p>Ask a frontier model to pull the title, parties, and effective date from a single contract and it will probably get the answer right.</p><p>That is real progress.</p><p>It is also not the same problem as organizing a contract population.</p><p>So I ran a first benchmark: 100 real contracts, five metadata fields, three model runs, one fixed prompt, clean text-layer PDFs only.</p><p>The result was not a simple leaderboard. The more useful finding was where the workflow started to bend.</p><p>At one document, the experience feels almost magical. At 100 documents, the bottlenecks start to show: session limits, inconsistent handling of ambiguous fields, silent wrong answers, and model-specific trade-offs between accuracy, speed, and usage.</p><p>That matters because contract AI workflows are increasingly being built on top of specific model versions. And model access itself is now a planning variable. Last week&#8217;s export-control disruption affecting Anthropic&#8217;s Fable 5 and Mythos 5 models is an extreme example: access can change suddenly for reasons outside the workflow itself.</p><p>So this benchmark is partly about accuracy. But it is also about something more practical:</p><p><strong>What happens when a capability that works in a demo has to survive a realistic contract-population workflow?</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><p></p><div><hr></div><p><strong>What I tested</strong></p><p>Before you can do anything substantive with a contract population &#8212; M&amp;A diligence, renewals management, migration into a CLM, risk analysis &#8212; you have to know what you are looking at.</p><p>Which documents are contracts at all?</p><p>What types of contracts?</p><p>What does the document call itself in its own language?</p><p>Who are the parties?</p><p>When did the agreement take effect?</p><p>Organizing contracts is not glamorous, but it is the dependency for everything downstream. Almost every contract intelligence workflow starts here, which is why this is the first benchmark I ran.</p><p>The scope was deliberately narrow: extract five metadata fields from each document and return one CSV row per document.</p><p>The five fields:</p><ul><li><p><strong>Document title</strong> &#8212; as stated in the heading or preamble</p></li><li><p><strong>Party names</strong> &#8212; full legal names as stated in or around the preamble</p></li><li><p><strong>Effective date</strong> &#8212; stated effective date, falling back to other best date available in the preamble</p></li><li><p><strong>Preamble</strong> &#8212; verbatim introductory opening paragraph usually containing the party names and date</p></li><li><p><strong>Recitals</strong> &#8212; verbatim background or procedural history</p></li></ul><p>Nothing downstream. No clause extraction. No obligation tracking. No risk analysis. No product build.</p><p>That narrowness is the point. These are the fields many downstream workflows assume you already have, and they are mostly verifiable against the document itself. They also tend to appear near the beginning of the contract, which creates a more controlled extraction task than clause-level fields that may appear anywhere.</p><p>This benchmark is not trying to prove that AI can read contracts.</p><p>It is trying to test what happens when basic contract reading becomes part of a population-level workflow.</p><div><hr></div><p><strong>The documents</strong></p><p>The sample consisted of 100 material contracts sourced from SEC EDGAR filings.</p><p>These are real, executed agreements, not synthetic examples. They are not a perfect analog for enterprise contract populations. They skew toward negotiated, higher-value agreements and generally arrive in cleaner condition than the contracts sitting in a shared drive, inbox archive, or legacy CLM export.</p><p>But they are useful for a first benchmark because they are:</p><ul><li><p>real agreements,</p></li><li><p>publicly available at scale,</p></li><li><p>structurally similar to enterprise contracts,</p></li><li><p>and suitable for manual ground-truth annotation.</p></li></ul><p>For this first run, I used clean text-layer PDFs only.</p><p>That means this is not yet a test of scanned contracts, bad OCR, image-only files, or corrupted text layers. Those conditions matter &#8212; a lot &#8212; but I wanted to isolate model choice first before adding document-quality variation.</p><div><hr></div><p><strong>The models and prompt</strong></p><p>I tested three models against the same 100 documents using the same prompt semantics: Claude&#8217;s Sonnet 4.6 and Opus 4.8, and OpenAI&#8217;s GPT-5.5. I used Claude Cowork and OpenAI&#8217;s Codex to run these.</p><p>Model choice was the only variable in this run.</p><p>The prompt was fixed across all runs. It was approximately one page of context, written to strike a practical balance: specific enough to define the extraction task and output format, but not so optimized that it became a bespoke engineering exercise.</p><p>In other words, this was not a zero-shot &#8220;please extract metadata&#8221; prompt. But it was also not a heavily tuned production workflow.</p><p>It was the kind of first-version prompt a knowledgeable contracts professional could plausibly draft in an hour with a clear understanding of the desired output.</p><p>I&#8217;m not publishing the full prompt or workflow notes here because the point of this piece is the benchmark result, not a replication guide. But the methodology matters, so I&#8217;ve described the corpus, fields, scoring approach, and limitations clearly enough for the findings to be contextualized and evaluated.</p><div><hr></div><p><strong>How I scored the results</strong></p><p>I scored the outputs across three dimensions.</p><p>First: <strong>field return rate</strong>. Did the model return a value for the field at all? I tracked this separately from accuracy because a blank output and a wrong output create different risks. In this run, it was less of a differentiator between models.</p><p>Second: <strong>accuracy</strong>. For free text fields &#8212; title, parties, preamble, recitals &#8212; I assessed substantive match against the manually annotated ground truth. For effective date &#8211; I used exact match after date conversion to YYYY-MM-DD.</p><p>Third: <strong>time and usage</strong>. I tracked wall-clock runtime and session usage consumed per model, per run.</p><p>A note on ground truth: contract metadata is often less objective than it looks.</p><p>Take document title. Sometimes the cover page says one thing and the opening paragraph says another. Sometimes the heading includes &#8220;Executed Version,&#8221; &#8220;Confidential,&#8221; &#8220;Exhibit 10.3,&#8221; or party names that are not really part of the title. Sometimes the preamble characterizes the document differently from the header.</p><p>Effective dates are similar. Across this sample, effective dates appeared in several forms: stated explicitly in the preamble, defined by reference elsewhere in the document, implied by an &#8220;as of&#8221; convention, or left to the signature block.</p><p>So the benchmark uses documented annotation rules applied consistently. Ambiguous cases are flagged rather than forced.</p><p>The goal is not a perfect golden set. For contract populations, the term &#8220;golden set&#8221; often oversells the certainty available. The goal is a reasonably standardized reference set &#8212; and the difficulty of building even that is part of the finding.</p><div><hr></div><p><strong>The headline result</strong></p><p>The headline result is not simply that one model performed better than another.</p><p>The more useful findings are that 1-the errors were patterned and 2-there are trade-offs between speed, usage and accuracy.</p><p>The models did not fail randomly. They struggled on the same kinds of contract complexity:</p><ul><li><p>ambiguous titles,</p></li><li><p>party-name exactness,</p></li><li><p>effective dates defined by reference,</p></li><li><p>and multi-party agreements.</p></li></ul><p>That matters because it changes how to interpret the benchmark.</p><p>If errors were random, the answer would be &#8220;use the most accurate model.&#8221; But if errors cluster around specific document structures, then knowing your contract population becomes key prior to, or in parallel with model selection.</p><p>A better model helps. But a better understanding of the document population may help more.</p><div><hr></div><p><strong>Speed and usage</strong></p><p>In addition, the three models diverged sharply on both accuracy as well as speed and session usage.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BYIA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BYIA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 424w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 848w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 1272w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BYIA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png" width="1006" height="392" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:392,&quot;width&quot;:1006,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55511,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/202056707?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BYIA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 424w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 848w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 1272w, https://substackcdn.com/image/fetch/$s_!BYIA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7ab9cd97-6ded-4b20-8da9-3d99bbd32e44_1006x392.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A few caveats.</p><p>These runs were performed using paid consumer tiers (Anthropic Pro, ChatGPT Plus), not enterprise licenses. That matters because most company use cases would likely require enterprise controls for data security, confidentiality, admin oversight, and related needs. Usage limits and economics may differ materially in those environments.</p><p>The GPT-5.5 usage figure also deserves a specific note. Current GPT-5.5 pricing appears unusually favorable and may be subsidized as part of a competitive user-acquisition strategy. That does not mean the figure is meaningless, but it should not be treated as a durable long-term cost benchmark. API usage at scale may tell a different story.</p><p>Still, the operational difference was real.</p><p>A single Claude Opus 4.8 run on 100 documents consumed roughly 71% of a paid consumer-tier session. That is manageable for one experiment. It is far less likely to be a comfortable foundation for routine extraction work at scale.</p><p>At 1,000 documents, consumer-tier session limits would likely become a workflow design constraint well before the task itself is complete.</p><div><hr></div><p><strong>Accuracy by field</strong></p><p>Claude Opus 4.8 performed best in terms of accuracy both overall and for each specific field. GPT-5.5 and Claude Sonnet 4.6 performed similarly overall but varied significantly in terms of field-level performance.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gZvI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gZvI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 424w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 848w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 1272w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gZvI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png" width="1006" height="516" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:516,&quot;width&quot;:1006,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53894,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/202056707?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gZvI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 424w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 848w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 1272w, https://substackcdn.com/image/fetch/$s_!gZvI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7f2d6f-35fd-4f15-8075-95cf9004dc06_1006x516.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>n = 100 documents. Effective date scored on normalized exact match; free-text fields scored on substantive match.</p><p>The accuracy gap between models was real, but the spread was not the whole story. The more important question is <em><strong>where</strong></em> each model was wrong, and whether those errors would matter for the workflow.</p><p>For example, if one model is five percentage points more accurate but consumes nearly three times the session capacity, that may or may not be worth it. The answer depends on the use case, QC tolerance, document volume, and whether the incremental errors are concentrated in fields that actually matter.</p><p>That is why this benchmark is less useful as a leaderboard than as a workflow diagnostic.</p><p>As another example, for party names: more lenient criteria (including partial matches &#8211; such as where the model found one but not all of the parties) yielded the following results, with substantial improvements for both GPT-5.5 and Sonnet 4.6:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hslY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hslY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 424w, https://substackcdn.com/image/fetch/$s_!hslY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 848w, https://substackcdn.com/image/fetch/$s_!hslY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 1272w, https://substackcdn.com/image/fetch/$s_!hslY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hslY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png" width="944" height="196" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:196,&quot;width&quot;:944,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:21210,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://thecontractsignal.com/i/202056707?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hslY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 424w, https://substackcdn.com/image/fetch/$s_!hslY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 848w, https://substackcdn.com/image/fetch/$s_!hslY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 1272w, https://substackcdn.com/image/fetch/$s_!hslY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb059b75-3cc2-4333-b749-3f9578c1f5a9_944x196.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><div><hr></div><p><strong>Where the models failed</strong></p><p>The failure modes were not random.</p><p>Here are some common ones:</p><p><strong>1. Document title ambiguity</strong></p><p>Title was more ambiguous than it initially appears.</p><p>Some documents had a clean title in the heading. Others had a heading that included filing labels, stamps, party names, or version indicators. In some cases, the heading and preamble characterized the document differently.</p><p>For example: &#8220;FOURTH AMENDMENT TO COMMON SECURITY AND ACCOUNT AGREEMENT&#8221; (doc header) vs. &#8220;Fourth Amendment&#8221; (preamble).</p><p>This produced two kinds of issues.</p><p>First, models sometimes selected different title sources: header versus preamble. Second, even within the same model run, title-selection behavior was not always consistent.</p><p>That inconsistency matters more than a one-off wrong answer. A single bad title can be corrected. An inconsistent title-selection rule creates downstream cleanup work.</p><p><strong>2. Party-name exactness</strong></p><p>There were some significant party name mistakes that included missing relevant entity names, particularly in multi-party agreements.</p><p>However, more often models returned party names that were substantively correct but not exact.</p><p>Examples include:</p><ul><li><p>Including extraneous text in addition to a relevant party name,</p></li><li><p>omitting legal suffixes,</p></li><li><p>omitting punctuation</p></li></ul><p>In normal reading, these second class of errors are not serious mistakes. A person understands the answer.</p><p>But in a structured data workflow, small variations matter. Party names often become merge keys, deduplication inputs, entity-resolution candidates, or search filters. &#8220;Close enough&#8221; may be fine for a summary but problematic for a migration or analytics workflow.</p><p>This is one of the recurring lessons of contract data work: an answer can be substantively right and still operationally messy.</p><p><strong>3. Effective-date variation</strong></p><p>Effective dates were one of the hardest structured fields.</p><p>The reason is not that models cannot read dates. The reason is that contracts state effective dates in various ways.</p><p>Some are explicit:</p><p>&#8220;This Agreement is effective as of January 1, 2024.&#8221;</p><p>Others are indirect:</p><p>&#8220;This Agreement is entered into as of the date first written above.&#8221;</p><p>Others define the date elsewhere, refer to a closing, rely on signature timing, or contain several dates that look plausible but serve different functions.</p><p>When models relied too heavily on a standard preamble pattern, they missed dates defined elsewhere. When models searched more broadly, they sometimes surfaced the wrong date: a prior agreement date, notice date, recital date, or signature date.</p><p>This is exactly the kind of field where a model can look confident while being wrong.</p><div><hr></div><p><strong>What did not fail</strong></p><p>It is worth saying what worked.</p><p>On clean text-layer PDFs with standard preamble structure, all three models were generally consistent. They identified ordinary titles, ordinary party names, and explicitly stated effective dates reasonably well.</p><p>That is important. This is not an argument that model-based extraction does not work.</p><p>It does work.</p><p>The question is where it stops working cleanly, how visibly it fails, and how much workflow you need around it before the output can be trusted.</p><p>The failures above are patterned, which means they are addressable. Better prompts can help. Post-processing rules can help. Entity normalization can help. QC sampling can help. A human-in-the-loop review step can help.</p><p>But none of that is free.</p><p>The model is only one part of the workflow.</p><div><hr></div><p><strong>The silent failure problem</strong></p><p>The most important operational finding was not that the models made mistakes.</p><p>It was that the wrong answers looked exactly like the right answers.</p><p>Across all three models, every field in every document received a value. No uncertainty flags. No &#8220;I could not determine this.&#8221; No visible indication in the CSV that a given value required review.</p><p>That is convenient until it is dangerous.</p><p>At 100 documents, a human QC pass can catch many of these issues.</p><p>At 1,000 documents, that is no longer true unless sampling is built into the workflow explicitly. At 10,000 documents, even a small systematic error becomes a data-quality problem. At 100,000 documents, a 2% intervention rate means 2,000 human interventions.</p><p>The operational risk is not merely that the model gets something wrong.</p><p>It is that the model gets something wrong confidently, and confidence is the only signal available downstream.</p><div><hr></div><p><strong>What this means</strong></p><p><strong>The first lesson:</strong> model choice mattered, but document structure mattered more.</p><p>Across all three models, the lowest-accuracy documents were not random. They were the documents missing preambles, where effective dates were not clearly defined, where there were more than two named parties, or title/preamble mismatches.</p><p>That means knowing your contract population is key prior to model selection. If your population has a high concentration of those structures, prompt design, field definitions, and QC process may matter more than switching models.</p><p><strong>The second lesson:</strong> speed and usage tell a different story than accuracy.</p><p>Claude Opus 4.8 was directionally strongest on accuracy, but it was also the slowest and consumed the most session capacity. Claude Sonnet 4.6 was materially faster and used less session capacity. GPT-5.5 was fastest and cheapest by apparent session usage, but with pricing and durability caveats.</p><p>That creates a practical trade-off.</p><p>If Claude Opus 4.8&#8217;s overall accuracy advantage is 22 percentage points, the decision is not simply &#8220;use Claude Opus 4.8.&#8221; The decision is whether those extra correct answers are worth the added time, usage constraint, and workflow friction for the specific task.</p><p><strong>The third lesson:</strong> consumer-tier tools can benchmark the workflow, but they are not the workflow at scale.</p><p>At 100 documents, they are usable. At 1,000, session limits and manual handling become serious constraints. At higher volumes, API access, retry logic, QC sampling, structured storage, and review workflows become the relevant comparison.</p><p><strong>The fourth lesson:</strong> silent failures require explicit QC.</p><p>Do not treat a complete CSV as a correct CSV.</p><p>A model returning values for every field is not the same thing as a validated extraction workflow. If the output will be used for diligence, migration, analytics, obligation tracking, or system-of-record updates, some QC process has to be designed into the workflow.</p><div><hr></div><p><strong>Takeaways</strong></p><p>These thresholds are directional, based on this first run. I will refine them as the benchmark develops.</p><p><strong>10 documents or fewer</strong></p><p>Use the chat interface.</p><p>At this volume, setup overhead matters more than workflow sophistication. Drag-and-drop, ask for the fields, copy the output into a spreadsheet, and manually check the results.</p><p>Model choice is probably less important than clear instructions and human review.</p><p><strong>Around 100 documents</strong></p><p>The chat interface still works, but the constraints become visible.</p><p>You need a repeatable prompt, a consistent output schema, a place to store the results, and a QC pass before using the output downstream. Watch especially for ambiguous titles, effective dates, and multi-party agreements.</p><p>This is the tier where many ad hoc legal ops and in-house requests live. It is also where teams can fool themselves into thinking they have a process when they really have a clever manual workaround.</p><p><strong>100 to 1,000+ documents</strong></p><p>The workflow around the model starts to matter more than the model.</p><p>You need consistent prompting, structured output, QC sampling, failure logging, and a storage layer that is not a manually assembled spreadsheet. Consumer-tier session limits become impractical. API access becomes the relevant comparison.</p><p>The per-document model cost may be low. The human time cost of managing an ad hoc process at this scale is likely much higher.</p><p><strong>10,000+ documents</strong></p><p>The question changes.</p><p>This is no longer &#8220;which model should I use?&#8221; It becomes &#8220;what operating model do we need?&#8221;</p><p>Build, buy, outsource, or run a formal internal data operation &#8212; but do not mistake model access for workflow readiness. At this scale, reliability, review design, exception handling, and data governance become the work.</p><div><hr></div><p><strong>The useful distinction</strong></p><p>The demo version of contract extraction asks:</p><p>Can the model read this contract?</p><p>The operational version asks:</p><p>Can the workflow produce structured data across this population, with known accuracy, acceptable review burden, manageable failure modes, and outputs that downstream users can trust?</p><p>Those are different questions.</p><p>The first is increasingly easy to answer.</p><p>The second is where most of the real work still lives.</p><div><hr></div><p><strong>What comes next</strong></p><p>This is a first data point, not a conclusion.</p><p>I will publish additional cuts as the benchmark produces useful findings: document-quality conditions, additional models, larger scale tiers, and the operational differences between chat-based and API-based workflows.</p><p>If you found this useful, the best signal you can send is forwarding it to someone who is about to make a decision about contract extraction tooling based on a demo.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The contracts AI announcements keep coming. They’re not solving the same problem.]]></title><description><![CDATA[The market is converging on contracts AI. But 'contract intelligence,' 'agentic workflows,' and 'practice-area plugins' aren't the same thing &#8212; or where most of the time goes.]]></description><link>https://thecontractsignal.com/p/the-contracts-ai-announcements-keep</link><guid isPermaLink="false">https://thecontractsignal.com/p/the-contracts-ai-announcements-keep</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 07 Jun 2026 23:06:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/08921a2f-2f3c-494d-b6b4-f70e941b8442_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the last few months, Harvey announced Contract Intelligence. Anthropic launched Claude for Legal with practice-area plugins. DocuSign unveiled agentic contract workflows.</p><p>Last year brought more of the same: LexisNexis announced Prot&#233;g&#233;&#8482; General AI, Thomson Reuters announced agentic capabilities in CoCounsel Legal, OpenAI published a detailed account of building an internal contract data agent.</p><p>That&#8217;s not a coincidence. </p><p>The market is converging around the contracts space using ever-improving AI capabilities &#8212; from creating &#8220;contract intelligence&#8221; to building end-to-end &#8220;workflows.&#8221;</p><p>But there&#8217;s a lot missing in the conversation around these announcements. So let&#8217;s unpack what it means and why it matters.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong>To start &#8212; contract intelligence</strong></p><p>I&#8217;ve spent about a decade working on this problem. </p><p>I started in 2015, inside a small contracts business at Axiom that eventually became Knowable, a LexisNexis company. Back then, the deliverable was a spreadsheet. The engine was people: project managers and analysts running a quality-controlled review process on enterprise contracts, turning dense legal language into structured data someone could actually use. </p><p>There was no software to buy. The concept we now call &#8220;contract intelligence&#8221; predated the term <strong>&#8212; </strong>and predated most of the tools now claiming it.</p><p>That is why the current wave is so interesting. The concept is old. The technology is genuinely new and continues to evolve. And the gap between those two things is where the market is hardest to understand and is where the useful analysis lives.</p><div><hr></div><p><strong>Same space, different layers</strong></p><p>Here&#8217;s what&#8217;s worth noticing: these announcements are solving for different things in different ways, but they&#8217;re all filed under similar labels.</p><p>Harvey&#8217;s Contract Intelligence appears to be a portfolio-level playbook management system &#8212; surfacing negotiation patterns, fallback positions, and clause language from executed agreements to make future review faster. The promise is that every signed contract feeds back into and updates the playbook automatically. That&#8217;s a workflow-compounding play aimed at in-house teams doing high-volume, BAU contract reviews.</p><p>Claude for Legal is a horizontal capability layer: practice-area plugins, first-pass review and redlining through Cowork, and MCP integrations that connect Claude to other systems in your stack. It&#8217;s positioned as the intelligence underneath other tools, not a standalone contracts product.</p><p>OpenAI&#8217;s contract data agent is an internal build for their own finance team &#8212; ingesting PDFs and scans, extracting structured data with retrieval-augmented prompting, and serving it up for human-in-the-loop review. The goal is to meet the scale of a growing business with limited headcount.</p><p>DocuSign&#8217;s agents operate inside its Intelligent Agreement Management platform: checking agreements against company standards, flagging risks, tracking obligations, and allowing teams to build custom agents for deal and renewal workflows through Agent Studio. The use cases overlap with what others serve, but the emphasis is on a platform-ecosystem anchored in their e-signature capabilities.</p><p>These are not competitors in the simple way a feature comparison might suggest. </p><p>They&#8217;re addressing different layers of a workflow that most announcements don&#8217;t bother to decompose. </p><p>And if you&#8217;re an end user &#8212; a contracts professional, in-house counsel, legal ops leader, or someone running a deal process &#8212; the useful question is not simply &#8220;which one is best?&#8221;</p><p>It is: </p><p><strong>Which layer of my actual problem does each one touch, what&#8217;s still missing, and what trade-offs come with choosing one approach over another?</strong></p><div><hr></div><p><strong>The part the announcements skip</strong></p><p>That question is harder to answer than it sounds because the announcements are built as much to impress as to inform. </p><p>They rarely locate themselves precisely in the workflow. They rarely explain what must already be true for the product to work well. And they rarely dwell on the gap between a staged demo and a system that survives real documents, real users, and real organizational complexity.</p><p>Every one of these announcements assumes a set of prerequisites that practitioners know are the actual hard part: </p><ul><li><p>That the contracts have been collected. </p></li><li><p>That you know which contracts are in scope. </p></li><li><p>That the entity names in your CRM match the party names in the agreements. </p></li><li><p>That users have access to the right documents. </p></li><li><p>That someone has decided what data points matter.</p></li><li><p>That the output can be reviewed, trusted and used in the company&#8217;s legal and business processes.</p></li></ul><p>Very little of that is solved by the announcement itself. </p><p>All of it is where most of the time goes.</p><p>The decade I spent building in this space taught me one pattern above everything else: each wave of technology solves the layer everyone is staring at, and the value quietly relocates to the layer nobody was watching. </p><p>Services gave way to software. Traditional machine learning improved extraction. Generative models changed what could be attempted. Each time, the extraction layer got better &#8212; genuinely, materially better.</p><p>And each time, the constraint moved somewhere upstream or downstream:</p><ul><li><p>Which contracts matter?</p></li><li><p>Where are they?</p></li><li><p>Who has access?</p></li><li><p>What should we extract?</p></li><li><p>How do we evaluate quality?</p></li><li><p>How do we surface the answer clearly?</p></li><li><p>How does the output change the way the business actually makes decisions?</p></li></ul><p>That pattern is playing out again now, across every one of these announcements.</p><p>It is the thing most worth tracking.</p><div><hr></div><p><strong>Why I&#8217;m writing this</strong></p><p>I come at this from a decade of working across ML and LLM product development, contract analysis, solution sales, and end-user research, with cross-functional visibility into how enterprise contracts work actually gets delivered.</p><p>I have no vendor to sell and no employer to promote. </p><p>The goal of <em>The Contract Signal</em> is to give end users &#8212; the people who actually work with contracts &#8212; an independent, practitioner-level read on what the capabilities are, what they are not, what is coming, and how to use them to solve real problems.</p><p>The next pieces will get hands-on: testing tools against real contract workflows, looking at what they can actually do, where they break, and what they quietly assume you have already solved.</p><p>If you work with contracts, buy or evaluate AI tools for legal workflows, or build in this space and see something I&#8217;m missing, I&#8217;d like to hear from you.</p><p>Subscribe to follow along.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Announced M&A Deal is the Worst Lead You'll Get All Year]]></title><description><![CDATA[Why M&A announcements are a misleading signal for anyone trying to originate contract work &#8212; and what actually predicts demand.]]></description><link>https://thecontractsignal.com/p/the-announced-deal-is-the-worst-lead</link><guid isPermaLink="false">https://thecontractsignal.com/p/the-announced-deal-is-the-worst-lead</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 31 May 2026 21:53:23 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b5b75e90-97a4-4b8b-9c4f-c1cc0a084b3a_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>My first assignment out of law school was monitoring M&amp;A deals.</p><p>I sat in front of a subscription tool that aggregated announced mergers, acquisitions, financings, and the rest. You could sort by deal size, by type, by buyer. You could set alerts. And the strategy my team ran on top of it was simple enough to explain in one sentence: a big deal gets announced, big deals mean lots of contracts to review, lots of contracts mean someone needs help, so we call them and win the work.</p><p>It sounds reasonable.</p><p>It&#8217;s also wrong in two structural ways that took me years to fully appreciate. If you&#8217;re building an origination motion for contract-review work today &#8212; whether you&#8217;re a vendor, a services firm, or a legal-ops team trying to staff ahead of demand &#8212; those two failures are worth as much as any feature comparison.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong>Failure one: by the time it&#8217;s news, the work is gone</strong></p><p>The first thing that breaks is timing.</p><p>A deal that&#8217;s been announced is a deal that&#8217;s already been planned. The diligence scope is set. The advisors are picked. Often the contract-review approach &#8212; internal team, outside counsel, a vendor, some combination &#8212; was decided weeks or months before the press release went out. You are reading about a decision, not catching one in flight.</p><p>So the announcement, which feels like the <em>start</em> of an opportunity, is usually the end of one. The companies that show up in your alert feed are disproportionately the ones you can no longer help on this deal. The genuinely addressable work was upstream, in the quiet period when nobody outside the principals knew it was happening.</p><p>This is the part that impacts everything else. If announced deals are mostly too late, then reactive monitoring isn&#8217;t an origination strategy &#8212; it&#8217;s a lagging indicator dressed up as a lead source. The work that actually converts comes from being known <em>before</em> the deal exists: proactive relationships, so the company comes to you when the deal is still a secret. You can complement this &#8212; but not replace &#8212; with news monitoring.</p><div><hr></div><p><strong>Failure two: volume doesn&#8217;t mean need</strong></p><p>The second failure is subtler and, I think, the more useful one, because it survives even if you fix the timing problem.</p><p>The whole premise &#8212; more deal volume means more contract work &#8212; assumes a correlation that just isn&#8217;t reliable. Deals are structured, and the structure determines whether contracts matter at all. A few examples that look similar in a volume column and behave nothing alike:</p><p><strong>The acqui-hire.</strong> A handful of people command an enormous package &#8212; you can get to several hundred million, even half a billion, on a few names in the current AI talent market. It will light up any deal-size screen you build. The contract review attached to it is close to zero. Nobody&#8217;s reviewing a customer book; there barely is one. Huge number, no work.</p><p><strong>The asset purchase.</strong> A buyer takes a discrete set of assets &#8212; say, a fleet of trucks. The assets are owned outright; there&#8217;s not much contractual tail to diligence on the buy side, and any related agreements may be getting retired or terminated as part of the structure anyway. So buy-side need is low. But flip it: the <em>seller</em> may have real work understanding how hard those contracts are to exit. The same deal generates demand on one side and almost none on the other, and a volume metric can&#8217;t see the difference.</p><p><strong>Business-as-usual masquerading as nothing.</strong> Meanwhile, the company quietly enters into fifty thousand new agreements a year in the ordinary course of business. That&#8217;s a completely separate pool of potential work from anything M&amp;A-driven &#8212; larger and steadier &#8212; and it never shows up in a deal feed at all, because it never gets announced.</p><p>Put those together and the lesson is uncomfortable for anyone who loves a clean dashboard: high deal volume is weakly correlated with contract-review demand, and the cases where it&#8217;s <em>most</em> wrong are often the flashy ones that dominate the news.</p><p><strong>What actually predicts demand</strong></p><p>None of this means prediction is hopeless. It means the naive screen is the floor, not the ceiling.</p><p>The floor &#8212; and I mean &#8220;naive&#8221; as a term of art, not an insult &#8212; is heuristic. Big, diversified companies with the cash flow or borrowing capacity to keep doing deals are likely serial transactors; their historical deal cadence tells you something. Working with physical goods helps: a business with a manufacturing or supply-chain footprint tends to carry more contractual surface area than a lighter software-first business, though large enough software companies blow that rule up with sheer customer and vendor count. Fortune 500, deal history, industry density &#8212; that&#8217;s a defensible level-one list, and it&#8217;s cheap to build.</p><p>The ceiling is where you stop treating those as the answer and start treating them as inputs. Layer in the things a screen can&#8217;t see: your relationship history, your own pipeline, deal-structure patterns, the buy-side/sell-side asymmetry above. Feed that into a model and you&#8217;re no longer asking &#8220;who&#8217;s big?&#8221; &#8212; you&#8217;re asking &#8220;who&#8217;s likely to transact <em>in a way that actually generates contract work</em>, and where are we well-positioned<strong> </strong>to win it?&#8221; That&#8217;s a prediction worth making.</p><p>One honest caveat, because it&#8217;s the thing most product pitches skip: the highest-value version of this is driven by <em>your</em> internal context &#8212; your relationships, your data, your read on structure. An outside-in tool can get you the level one list. It can&#8217;t replicate the part that makes the prediction good, which is the proprietary context only the firm doing the work has. Proceed accordingly and be wary of anyone selling you the ceiling as if it came in a box.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://thecontractsignal.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><em>If you&#8217;re working on origination or capacity planning around contract work and any of this maps to (or contradicts) what you&#8217;re seeing, I&#8217;d love to discuss.</em></p>]]></content:encoded></item><item><title><![CDATA[Replaying an Old Assignment Through an Agents Lens]]></title><description><![CDATA[A few years into my career at Axiom, I was handed an assignment that seemed simple on the surface: put together a newsletter.]]></description><link>https://thecontractsignal.com/p/replaying-an-old-assignment-through</link><guid isPermaLink="false">https://thecontractsignal.com/p/replaying-an-old-assignment-through</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 06 Apr 2026 01:31:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A few years into my career at Axiom, I was handed an assignment that seemed simple on the surface: put together a newsletter.</p><p>The context mattered. We were a nascent contracts data business operating inside a larger alternative legal services firm. Our business &#8212; turning contracts into structured data for M&amp;A due diligence and integration &#8212; was maybe ten percent of the company&#8217;s revenue. The other ninety percent ran a different model entirely: placing attorneys with in-house legal departments on a consulting basis. That larger business was our best source of referrals. When one of their senior account executives recognized that a client&#8217;s problem was better suited to what we did, they&#8217;d send it our way &#8212; and they&#8217;d still get credit toward their goals, so incentives were aligned.</p><p>The head of our business wanted to double down on that referral engine. His solution was internal marketing &#8212; get the other teams excited about what we were building, make them feel connected to it, make them more likely to think of us when the right client situation came up. He tasked me, roughly a year and a half into the job, with making that happen through a newsletter compiling the latest cross-functional updates.</p><p>I&#8217;ve been thinking about that assignment lately because it&#8217;s a useful test case for something that comes up constantly: where do agents actually fit into knowledge work, and where does the person still have to drive?</p><div><hr></div><h4><strong>What the Assignment Actually Involved</strong></h4><p>Before getting into the AI question, it&#8217;s worth being specific about what the work required, step by step.</p><p>The first thing I had to do was understand the business problem behind the assignment. The newsletter wasn&#8217;t the goal &#8212; increasing sales revenue was, through increased origination from the referral channel. That distinction mattered because it shaped every subsequent decision about what to include and how to frame it.</p><p>The second thing was figuring out what to compile. I started with our CRM to pull sales history, stakeholder context, deal status, and anything useful I could surface without taking up anyone&#8217;s time. Then I reached out across sales, product, marketing, and client service for anything the CRM didn&#8217;t capture &#8212; their read on what was worth highlighting, recent wins, delivery milestones, collateral. This produced a pile of good raw material: deals closed, new clients landed, product updates shipped, branding work completed.</p><p>The third thing &#8212; and this is where most of the judgment lived &#8212; was deciding what actually mattered and why, and how to communicate that to our audience.</p><p>Winning a deal with a new enterprise client was significant not just for its revenue, but for what it signaled about our ability to win that class of business and what it meant for the senior person on the other side of the house who originated it. A major renewal and upsell was worth highlighting not just as a financial outcome but as a recognition of the delivery team whose work made it possible. The connections between how collaboration across different parts of the business enabled our sales motions &#8212; and giving specific credit to the people involved &#8212; that&#8217;s what made the newsletter land, not the facts themselves.</p><p>Finally, the fourth thing was packaging it into a format that salespeople, executives, and other business stakeholders would actually read, let alone enjoy and be inspired by.</p><div><hr></div><h4><strong>Running It Back With Agents</strong></h4><p>The interesting question isn&#8217;t whether an agent could do this today. Parts of it clearly could be addressed. The interesting question is which parts, and what that leaves for the person.</p><p><strong>Compilation is largely addressable.</strong> The mechanical work of gathering updates across departments &#8212; if that information lives in a shared system with appropriate access &#8212; is exactly the kind of multi-step retrieval task agents handle well. Connect to the CRM, pull recent updates by department, surface them in a structured format. That&#8217;s real and useful. In my case I was doing much of this manually through CRM data research, email and conversations. Some of that friction goes away.</p><p>That said, the compilation step has its own constraints worth naming. The relevant information isn&#8217;t always well-structured or correctly tagged. Ambiguity in underlying data produces ambiguous outputs. And there are legitimate access and privacy boundaries &#8212; not every relevant conversation is logged, not every system is connected, and not every organization should want them to be. These aren&#8217;t temporary engineering problems. They&#8217;re features of how organizations actually function.</p><p><strong>Summarization is addressable, and honestly it was already addressable before agents.</strong> Given a set of inputs, an LLM produces coherent summaries. This isn&#8217;t a new capability agents unlock &#8212; it&#8217;s baseline LLM behavior. Useful to have in the workflow, not a step-change.</p><p><strong>Formatting and packaging is partially addressable.</strong> A first draft of the newsletter structure, given clear inputs and a well-defined template, is something an LLM handles comfortably. The blank page problem gets easier. The judgment in the final version is still a human job.</p><p><strong>What doesn&#8217;t move, at least not yet: the judgment calls at the beginning and middle of the process.</strong></p><p>Understanding why the assignment existed &#8212; that the real goal was referral activation, not newsletter production &#8212; required knowing the business, the person making the ask, and the organizational dynamics well enough to read between the lines. That context wasn&#8217;t written down anywhere. It came from being embedded in the organization.</p><p>Deciding what actually mattered in a pile of raw updates required knowing which outcomes were genuinely significant versus routine, which relationships between teams were worth reinforcing, and what would actually resonate with a specific audience of attorneys and senior salespeople being asked to care about something adjacent to their primary work. An agent surfacing everything equally would have produced a longer document, not a better one.</p><div><hr></div><h4><strong>A More Recent Illustration</strong></h4><p>There&#8217;s a more current version of the same dynamic that I&#8217;ve been thinking about while using Claude Code for my own projects.</p><p>When Claude Code runs, it periodically pauses and asks whether you want to auto-accept its next actions or review them manually. It&#8217;s a small UX detail, but it captures something important: the model itself doesn&#8217;t know when human review is warranted. It can execute well across a wide range of tasks. What it can&#8217;t do is determine, in your specific context with your specific constraints, which outputs are safe to accept without review and which ones need your eyes on them first.</p><p>That decision &#8212; when to trust the output and when to verify &#8212; is a judgment call. And it&#8217;s not incidental that the tool surfaces it explicitly as a user choice rather than resolving it internally. The people who built it understand that this is a structurally human decision, not a capability gap to be closed in the next model release.</p><p>The newsletter assignment had the same structure. An agent could have helped me compile, draft, and format. What it couldn&#8217;t have done is tell me which deal was worth leading with, or how to frame a delivery team&#8217;s contribution in a way that would actually motivate a sales floor to send us more referrals. Those calls required context that lived in relationships and organizational dynamics that no system was tracking.</p><div><hr></div><h4><strong>What This Suggests Practically</strong></h4><p>The newsletter assignment was never going to be fully automated, and that&#8217;s not a failure condition. The right way to think about it is as a workflow with a natural division of labor: the person drives the judgment-heavy front end, the tools handle the execution-heavy middle, and the person reviews and refines the back end.</p><p>In practice that looks something like this. You identify the business problem clearly enough that it can guide downstream decisions &#8212; human work. You use the tools to pull and structure relevant inputs from the systems that have them &#8212; agent territory. You evaluate what came back against the judgment you brought in at the start &#8212; human again. You use the tools to draft and format &#8212; mixed, depending on how clearly you&#8217;ve defined what good looks like. You review and refine with your audience in mind &#8212; human, and not optional.</p><p>The thing I&#8217;d push back on is the framing that treats this division as temporary &#8212; that eventually the agent handles the judgment calls too as models improve. Maybe. But the judgment calls in the newsletter assignment depended on context that is structurally difficult to provide to any system, not just ones with current capability limits. The organizational dynamics, relationship histories, and political considerations that informed my decisions weren&#8217;t in any system. Getting them there fully would require a level of organizational surveillance that most environments wouldn&#8217;t and shouldn&#8217;t accept.</p><p>Until that changes, the person isn&#8217;t a bottleneck to be engineered around. They&#8217;re the thing that makes the output worth producing in the first place.</p>]]></content:encoded></item><item><title><![CDATA[AI Agents: Getting More Done Isn’t the Same as Creating Value]]></title><description><![CDATA[Article 7 in a series on building with and using AI tools as a non-native technologist]]></description><link>https://thecontractsignal.com/p/ai-agents-getting-more-done-isnt</link><guid isPermaLink="false">https://thecontractsignal.com/p/ai-agents-getting-more-done-isnt</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 29 Mar 2026 17:08:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Article 7 in a series on building with and using AI tools as a non-native technologist</em></p><div><hr></div><h3>Intro</h3><p>There&#8217;s been a lot of noise lately about AI agents &#8212; most recently OpenClaw, before that OpenAI&#8217;s agents, Anthropic&#8217;s, Google&#8217;s, and the growing ecosystem of tools that let you automate workflows, string tasks together, and run complex operations with minimal intervention. The pitch is compelling: more getting done, less effort spent doing it.</p><p>I&#8217;ve found most of that pitch to be true in a narrow but important sense. These tools genuinely do help you get more done. I use them regularly for exactly that purpose.</p><p>But I keep noticing a gap between the claim and what&#8217;s actually happening &#8212; between <em>getting more done</em> and <em>creating value</em>. They&#8217;re not the same thing. The confusion between them, I think, is one of the more consequential misunderstandings in how these tools are being talked about right now.</p><p><strong>Accompanying audio note:</strong> </p><div class="native-audio-embed" data-component-name="AudioPlaceholder" data-attrs="{&quot;label&quot;:null,&quot;mediaUploadId&quot;:&quot;9d4578d7-4440-4c60-9b48-fc75c67c1fe5&quot;,&quot;duration&quot;:632.92084,&quot;downloadable&quot;:false,&quot;isEditorNode&quot;:true}"></div><div><hr></div><h3>The Productivity Claim</h3><p>When people talk about AI productivity gains, they&#8217;re usually describing something real: tasks that previously took an hour now take ten minutes. Analyses that required manual effort can be automated. Research that involved dozens of searches can be compressed into a few prompts.</p><p>This is genuinely useful. I don&#8217;t want to dismiss it. But productivity, strictly speaking, just means output per unit of input. Getting more things done in less time. And productivity is only valuable to the degree that the things being done are worth doing.</p><p>That sounds obvious&#8230;but watch what happens in practice. Someone deploys an agent that automates a multi-step process. They count the hours saved. They extrapolate to the team, multiply by average hourly rate, and arrive at a dollar figure. The number sounds large. The announcement goes out: <em>$X million in value created.</em></p><p>In most cases, this conflates two very different things: the cost of the activity and the value of the outcome.</p><p>Getting more done faster is only as valuable as what&#8217;s getting done. If an AI agent helps you accelerate a process that was already misaligned with what your customers needed, you&#8217;ve become more efficient at something that shouldn&#8217;t have existed in the first place. That&#8217;s not value creation &#8212; it&#8217;s waste, just at a higher throughput.</p><div><hr></div><h3>The More Interesting Question</h3><p>None of this means productivity gains are worthless. It means the question worth asking is: <em>what do you do with the time you get back?</em></p><p>There are two very different answers. One is that the freed capacity gets absorbed by more tasks of the same kind &#8212; a never-ending queue of lower-priority work that expands to fill available bandwidth. More things get done. The ratio of important things to unimportant things stays roughly the same. The volume increases.</p><p>The other answer is that the freed capacity gets redirected toward genuinely higher-leverage work &#8212; the things that actually move outcomes, that create something worth having, that compound over time. If that happens, then the productivity gain is real in the deeper sense: it bought you time to work on the things that matter.</p><p>Which answer you get depends almost entirely on whether you have the judgment and the agency to make that call. Judgment to identify what the high-leverage work actually is. Agency to actually redirect toward it &#8212; which, in an organizational setting, is far from guaranteed.</p><p>In a previous article I wrote about the difference between where you build in the stack and where your time actually compounds. The same logic applies here. Productivity tools don&#8217;t answer the allocation question. They just widen the gap between people who are asking it well and people who aren&#8217;t.</p><div><hr></div><h3>A Third Category</h3><p>There&#8217;s one more thing I&#8217;ve been sitting with, and it&#8217;s harder to articulate cleanly.</p><p>Most discussions of AI productivity focus on one of two categories: either you were doing the task before and now do it faster, or you weren&#8217;t doing it and now you can. What I&#8217;ve started to notice is a third category &#8212; things you technically could have done before, but that were just annoying and friction-laden enough that you didn&#8217;t, in practice, ever do them.</p><p>A specific example: I&#8217;ve had FSA accounts for years, with funds that expire if unused &#8212; the classic use-it-or-lose-it situation. Submitting claims, especially for borderline or unfamiliar expense types, has always involved navigating opaque requirements, decoding denial reasons, and figuring out exactly what documentation was needed. Annoying enough that I&#8217;d let some money expire rather than waste time fighting the process.</p><p>Earlier this year I used ChatGPT to work through a claim denial. Not just as a search engine &#8212; as a conversational partner. I described the situation, asked about common denial reasons for that category, worked through what documentation would address each one, and drafted a resubmission. The claim was approved. Several hundred dollars I would have left on the table.</p><p>The productivity framing undersells what happened there. I didn&#8217;t do that task faster &#8212; I did a task I wasn&#8217;t going to do at all. The LLM didn&#8217;t just reduce time cost. It reduced friction below the threshold where I&#8217;d actually act.</p><p>There&#8217;s probably a psychological dimension to this as well. Something about the conversational format &#8212; the ability to ask follow-ups, to get a response that acknowledges your specific situation rather than returning generic search results &#8212; lowers the activation energy in a way that a list of web links doesn&#8217;t. I&#8217;m not sure whether that&#8217;s because the medium maps to how people have evolved to interact with and process information, their mental model or something else. But the effect is real.</p><p>This matters because it suggests a category of value that doesn&#8217;t show up in time-savings calculations. Not hours recaptured. Not existing tasks accelerated. But things that now happen that otherwise wouldn&#8217;t &#8212; claims filed, questions answered, decisions made with better information. That&#8217;s not productivity. It&#8217;s enablement.</p><div><hr></div><h3>What This Means in Practice</h3><p>The honest version of the AI agent story is more nuanced than the headlines suggest. These tools genuinely do help you get more done. In the right conditions &#8212; when there&#8217;s judgment about what&#8217;s worth doing and agency to act on it &#8212; that translates to real value. When those conditions are absent, you get faster hamster wheels.</p><p>For anyone building with or using these tools, I think the most important questions aren&#8217;t about the tools themselves. They&#8217;re about the human side of the equation: What are you actually trying to accomplish? Where does the allocation of your time currently mismatch what you&#8217;d choose if you thought clearly about it? And what category of things &#8212; the ones you&#8217;re not doing because the friction cost is too high &#8212; might be worth reconsidering?</p><p>Getting more done is the easy part. These tools are genuinely good at it. The harder part &#8212; which the tools don&#8217;t touch &#8212; is knowing what to do with the time.</p><div><hr></div><p><em>This is Article 7 in a series on building with and using AI tools as a non-native technologist. Previous articles covered financial planning with AI tools, a framework for city selection, evaluating LLM tools, building the Seattle Nature Access Map, why judgment was always the bottleneck, and where in the stack to build.</em></p>]]></content:encoded></item><item><title><![CDATA[Where in the Stack Should You Build?]]></title><description><![CDATA[At one point I had created a library of over 50 SQL and Opensearch template queries I used each week to wrangle up data for ML training.]]></description><link>https://thecontractsignal.com/p/where-in-the-stack-should-you-build</link><guid isPermaLink="false">https://thecontractsignal.com/p/where-in-the-stack-should-you-build</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 22 Mar 2026 18:07:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At one point I had created a library of over 50 SQL and Opensearch template queries I used each week to wrangle up data for ML training.</p><p>This was back before LLM tools became mainstream.</p><p>The queries weren&#8217;t sophisticated, but they handled my use cases well. They were also easy to understand and troubleshoot if something didn&#8217;t seem to be working.</p><p>I learned to write these &#8211; passably. Enough to pull data, join tables, answer a question without waiting for an engineer.</p><p>Once I reached a good enough level - I deliberately never got much better at it than that.</p><p>Every hour I could have spent becoming a stronger SQL writer was an hour I wasn&#8217;t doing the thing I was actually better at: figuring out which query mattered, what the data meant, and what it implied for the product. Where the models seemed to be doing well, where they needed to improve, and where we had no idea.</p><p>I could always find someone to write better queries. I couldn&#8217;t always find someone who knew what to ask.</p><p>That wasn&#8217;t laziness. It was a judgment about where my time compounded.</p><p>I&#8217;ve been thinking about that decision a lot since starting to build with AI tools in earnest &#8212; first the financial planning work I wrote about in Article 1, then the nature access map in Article 4. The more I build, the more I notice that the most consequential decisions aren&#8217;t about <em>how</em> to build something. They&#8217;re about <em>where</em> in the process your leverage actually lives.</p><div><hr></div><p><strong>The Stack Has Layers &#8212; And They&#8217;re Not Equal</strong></p><p>When people talk about &#8220;building with AI,&#8221; they&#8217;re usually collapsing several genuinely different things into one.</p><p>At the foundation, there&#8217;s training foundational models &#8212; the work Anthropic, OpenAI, and Google do. Enormous compute, proprietary data, specialized research teams. Almost no one reading this should be operating here.</p><p>One level up: fine-tuning and domain adaptation. Taking an existing model and adjusting it for a specific task or dataset. Still technical, but accessible to a much wider group &#8212; though the relevant question increasingly is whether the base model already does what you need, because often it does.</p><p>Above that: building AI-powered applications. Connecting LLMs to data, interfaces, and logic to make something useful. This is where most of the interesting work is happening right now, and where the barrier to entry has dropped most dramatically.</p><p>And above that: using AI-powered tools to create things that aren&#8217;t primarily engineering at all &#8212; writing, analysis, structured research, products where the judgment and domain knowledge are the core value, and the code is just the vehicle.</p><p>These are genuinely different layers. Which one you should work at isn&#8217;t obvious, and &#8220;I can build this now&#8221; doesn&#8217;t answer it.</p><div><hr></div><p><strong>The Right Level of Abstraction Is the Same Question</strong></p><p>In a previous article, I wrote about how judgment was always the bottleneck &#8212; that vibe coding made it visible, not new. This piece is about a related but distinct question: given that you have judgment, and given that you can now build things you couldn&#8217;t before, <em>where should you be building?</em></p><p>This is a question I&#8217;ve been trying to answer for myself.</p><p>There&#8217;s a useful frame from computer science education that I found clarifying here. When I first learned to code, I started with a site that abstracted away almost everything &#8212; you typed commands, things ran, but you had no idea why. Then I took a more rigorous introductory class that went deeper. Then, eventually, a course that started just above the transistor level &#8212; actual hardware, zeros and ones, how you get from electricity to software. Each level of abstraction revealed something the previous one had hidden.</p><p>The lesson wasn&#8217;t &#8220;always go to the lowest level.&#8221; Going down to the physics of transistors was helpful for understanding how computers work. It wasn&#8217;t necessary for building a web application. The insight was that the <em>right</em> level of abstraction depends on what you&#8217;re trying to do &#8212; and choosing wrong in either direction creates problems.</p><p>Too high: you can&#8217;t diagnose what breaks. This is what happened when I hit the rendering bug during the nature access map build. Claude kept proposing fixes that introduced new problems, and we went in circles until I suggested using browser dev tools to surface actual error information. Abstracting away all understanding of the underlying system meant I was stuck the moment something went wrong.</p><p>Too low: you&#8217;re spending your time on infrastructure that isn&#8217;t your actual problem. If I&#8217;d tried to build a more capable machine learning pipeline from scratch for the nature access scoring, instead of using public data and simpler distance calculations, I&#8217;d have spent months on something that wasn&#8217;t what the app was focused on.</p><p>Choosing where in the stack to build is the same decision as choosing the right level of abstraction. The mistake is treating it as fixed.</p><div><hr></div><p><strong>The Chef Doesn&#8217;t Make the Spatula</strong></p><p>Here&#8217;s the underlying principle: being capable of doing something and that being a good use of your time are different things.</p><p>A skilled chef could, in principle, forge their own knives, mill their own flour, build their own oven. They might have enough mechanical intuition to figure it out. But that&#8217;s not where their differentiation lives, and spending time there means spending less time on the cooking &#8212; which is where they&#8217;re actually hard to replace.</p><p>The question for anyone building now is the same one I was asking about SQL years ago: where does <em>my</em> time compound? Where is the gap between what I can do and what most people can do actually widest?</p><p>During the ten years I&#8217;ve spent in legaltech, the honest answer was never &#8220;write better ML code.&#8221; It was understanding what types of contract data matter to attorneys, salespeople, and finance teams &#8212; and how to structure and surface it in a way that was most useful to them.</p><p>That meant understanding what clause types mattered to an M&amp;A attorney under what conditions, how to translate a 91% accuracy rate into something meaningful for a lawyer managing risk, which edge cases would surface in exactly the situations that mattered most. Those were judgment calls that required domain knowledge I&#8217;d spent years building. Better tooling wouldn&#8217;t have made them easier &#8212; it would have made the engineering around them cheaper, which is a different thing.</p><p>That calculus hasn&#8217;t changed. What&#8217;s changed is that the floor of &#8220;good enough building&#8221; is now much higher, which means more people can reach it &#8212; and that makes the question of where you&#8217;re actually differentiated more urgent to answer explicitly. Before LLMs, if you couldn&#8217;t write the code, the question was moot. Now it isn&#8217;t, and that&#8217;s mostly a good thing. It just means the &#8220;where should I build?&#8221; question is no longer optional.</p><div><hr></div><p><strong>So Where Does Your Time Compound?</strong></p><p>Building toward something of my own has sharpened this question more than anything else. If I can build more easily than before, what should I actually build? At which layer?</p><p>The honest version of the question isn&#8217;t &#8220;can I do this?&#8221; It&#8217;s: where is my specific combination of domain knowledge, judgment, and taste producing something that isn&#8217;t easily replicated by someone working adjacent to me?</p><p>For some people, that points toward a lower layer &#8212; they have genuine ML depth, or systems architecture intuition, or infrastructure knowledge that&#8217;s still genuinely hard to replicate. For most of us building on top of these models, the answer points up: toward the quality of the problem being solved, the specificity of the insight, the understanding of the user, the design of the experience.</p><p>This doesn&#8217;t mean the engineering disappears. As I wrote in <em>Judgment Was Always the Bottleneck</em>, engineering judgment &#8212; knowing when a latency issue will create a UX crisis under load, knowing when a context window limitation will cap what&#8217;s possible, knowing when something technically works but creates maintenance debt that isn&#8217;t worth it &#8212; that&#8217;s still judgment, just expressed through technical decisions. The stack doesn&#8217;t have a clean separation between &#8220;engineering&#8221; and &#8220;everything else.&#8221;</p><p>What it means is that for most people who came to technology laterally &#8212; through law, finance, a domain specialty, a problem they needed to solve &#8212; the differentiation probably isn&#8217;t at the foundational layer. It&#8217;s in knowing what to build that&#8217;s actually worth building.</p><p>The spatula is cheaper now. The chef is still the chef.</p><div><hr></div><p><em>This is Article 6 in a series on building with AI tools as a non-native technologist &#8212; covering what actually changes when someone without an engineering background starts building products. Previous articles covered financial planning with Claude Code, a framework for city selection, evaluating LLM tools, building the Seattle Nature Access Map, and why judgment was always the bottleneck.</em></p>]]></content:encoded></item><item><title><![CDATA[There Are No Laws, Only Principles]]></title><description><![CDATA[Highlights:]]></description><link>https://thecontractsignal.com/p/there-are-no-laws-only-principles</link><guid isPermaLink="false">https://thecontractsignal.com/p/there-are-no-laws-only-principles</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 15 Mar 2026 22:48:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="native-audio-embed" data-component-name="AudioPlaceholder" data-attrs="{&quot;label&quot;:null,&quot;mediaUploadId&quot;:&quot;06e4f788-8b8c-4615-8342-2fb25a5ca970&quot;,&quot;duration&quot;:1496.0587,&quot;downloadable&quot;:false,&quot;isEditorNode&quot;:true}"></div><p><strong>Highlights:</strong></p><ul><li><p>All laws are are actually principles applied under certain conditions or assumptions</p></li><li><p>The strength of a principle is a function of how tightly you define the conditions under which it applies &#8212; narrow the scope, strengthen the law</p></li><li><p>Complex systems are hard to predict because multiple counteracting principles operate simultaneously</p></li><li><p>This is why macroeconomics gets predictions wrong repeatedly: too many counteracting principles operating simultaneously</p></li><li><p>The same pattern shows up in machine learning: a model trained on all contracts performs worse than one scoped to a specific contract type</p></li><li><p>And in UX: a design principle that worked for one user group often fails when applied to a different one without re-examining the underlying assumptions</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Judgment Was Always the Bottleneck]]></title><description><![CDATA[This is Article 5 in a series on building with AI tools as a non-native technologist.]]></description><link>https://thecontractsignal.com/p/judgment-was-always-the-bottleneck</link><guid isPermaLink="false">https://thecontractsignal.com/p/judgment-was-always-the-bottleneck</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 08 Mar 2026 20:01:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is Article 5 in a series on building with AI tools as a non-native technologist.</em></p><div><hr></div><p>There&#8217;s a popular narrative being told right now about vibe coding, and it goes roughly like this: coding used to be the bottleneck. Now it isn&#8217;t. The new bottleneck is judgment and taste. If you can think clearly about what to build, the tools will handle the rest.</p><p>It&#8217;s a compelling narrative. But it only tells half the story.</p><p>Not because judgment doesn&#8217;t matter &#8212; it does, enormously, and I&#8217;ll spend most of this article on why. But because the framing implies that judgment is <em>new</em> to the equation. That before vibe coding, the separating factor between great products and mediocre ones was mainly technical execution.</p><p>My re-frame is simpler: judgment was always the bottleneck. Vibe coding just made it visible.</p><div><hr></div><p><strong>What Actually Changed</strong></p><p>Things did change, substantially. Claude Code and similar tools collapsed the barrier to entry for certain categories of development work in a way that prior no-code tools never quite managed. The prototype-to-production gap is narrower than it was. The floor of what a non-engineer can ship went up substantially.</p><p>But here&#8217;s the disconnect: we&#8217;re confusing a change in what we can <em><strong>see</strong></em> as a change in what was <em><strong>present</strong></em>.</p><p>Before widespread vibe coding, poor product judgment was largely contained inside institutions. It existed &#8212; and produced plenty of mediocre features and failed products &#8212; but it was wrapped inside teams at companies, hidden behind proprietary systems, filtered through organizational process. You didn&#8217;t see it at scale in the open.</p><p>Now you do. Thousands of people are shipping publicly, often free, often building in the open. So a much higher volume of things to look at means a much higher volume of poorly crafted things to look at. The ratio of strong products to weak ones was probably always similar. What changed is that you now have a massive lens on all of it, versus a microscope pointed mostly at what large companies chose to release.</p><p>There&#8217;s a marketing layer on top of this too. &#8220;I&#8217;m a builder&#8221; has become the new &#8220;I&#8217;m a disruptor&#8221; &#8212; a buzzword with similar dynamics to disruption a decade ago. When everyone&#8217;s shipping, people pay attention to what gets shipped. The scrutiny goes up alongside the volume. Both effects make judgment failures more visible without changing the underlying rate.</p><p>The story isn&#8217;t &#8220;judgment matters now.&#8221; It&#8217;s &#8220;judgment failures are harder to hide now.&#8221;</p><div><hr></div><p><strong>Judgment Isn&#8217;t New &#8212; It Was Always Load-Bearing</strong></p><p>To understand why, it helps to break down what building something actually requires. At the highest level, three things: knowing <em>what</em> to build, knowing <em>how</em> to build it, and getting it to market.</p><p>The first category &#8212; what to build &#8212; means understanding the problem you&#8217;re solving, for whom, and conceptually how. When I built the nature access map, I was my own first user. I knew the problem in depth because I&#8217;d lived it: I wanted to understand, by neighborhood, which areas of Seattle had parks and trails within walking distance before committing to a move. That gave me a clear problem, a clear user, and a clear mental model of what a useful solution would look like.</p><p>The second category &#8212; how to build &#8212; means going from that concept to something tangible: mock-ups, prototypes, and eventually working code. This is where the LLM tools have changed things most dramatically. Historically, you couldn&#8217;t get to a working prototype at all without substantial technical fluency. Now, for many categories of work, you can describe what you want in plain English and get something functional back.</p><p>The third category &#8212; distribution and commercialization &#8212; means actually getting the product in front of users and, if you&#8217;re building something commercial, making money from it.</p><p>What changed is the second category. What didn&#8217;t change is the first and the third. And the first category &#8212; judgment about what to build, for whom, and why &#8212; was always the harder problem. It&#8217;s just that when the second category was blocking most people entirely, the first category&#8217;s difficulty was invisible. You can&#8217;t expose a judgment problem in something that never gets built.</p><p>I&#8217;ve spent a decade in legaltech, turning contracts into structured data for enterprise customers. The engineering was real &#8212; machine learning models, complex document processing pipelines, large-scale data extraction and classification, and increasingly GenAI-powered tools. But the decisions that actually determined whether the product worked were almost never purely engineering decisions. Was the model accuracy on this clause type sufficient for the risk appetite of an in-house counsel or M&amp;A attorney? Which edge cases would appear infrequently enough that we could accept them, versus which would surface in exactly the situations that mattered most? How did you explain a 92% accuracy rate to a lawyer in a way that mapped to how they actually processed risk?</p><p>Those were judgment calls. They required domain knowledge, pattern recognition, and the ability to translate between technical realities and human contexts. No amount of better tooling would have made them easier. The people who made them well were valuable long before vibe coding existed.</p><div><hr></div><p><strong>The Complexity Matrix No One Is Talking About</strong></p><p>Part of why the popular narrative feels limiting is that it collapses a genuinely complex set of distinctions into a single claim.</p><p>Whether vibe coding meaningfully changes the engineering bottleneck depends entirely on where you are across multiple dimensions at once:</p><p><strong>Internal tool vs. external product.</strong> An internal tool built for yourself has radically different tolerances &#8212; for rough edges, for edge cases, for performance under load. When I built the nature access map, it ran locally, served one user, and if it broke I fixed it. That&#8217;s a completely different problem than a SaaS product with uptime requirements, paying customers and users who aren&#8217;t you.</p><p><strong>Prototype vs. production.</strong> The gap between something that works in a demo and something that works under real conditions &#8212; concurrent users, unexpected inputs, evolving requirements &#8212; remains enormous, and LLMs haven&#8217;t closed it. Prior no-code tools produced a high ratio of prototypes to production-ready products for real reasons. There&#8217;s an honest open question about whether this generation finally crosses that threshold. I think the answer is &#8220;maybe, in limited contexts, but we don&#8217;t fully know yet.&#8221;</p><p><strong>Simple vs. complex use case.</strong> Some design patterns have been solved thoroughly enough that you can implement them at a raised level of abstraction and they&#8217;ll work reliably. Others haven&#8217;t. New interfaces &#8212; AR, voice, novel interaction paradigms &#8212; haven&#8217;t accumulated the solved patterns that web and mobile have built up over decades. What worked on desktop didn&#8217;t always translate to mobile. What works now won&#8217;t automatically translate to whatever comes next. Every new interface introduces substantial new variables and trade-offs.</p><p><strong>Regulated vs. unregulated.</strong> Financial services, legal, healthcare &#8212; these domains introduce constraints that compound complexity in ways that aren&#8217;t just about technical difficulty. The cost of certain failure modes is qualitatively different. The judgment required to navigate those trade-offs is domain-specific in ways that can&#8217;t be substituted by general technical fluency.</p><p><strong>Scale.</strong> A few users is categorically different from a few thousand or a few million. Engineering judgments that produce acceptable outcomes at small scale often produce catastrophic ones at large scale. Latency barely noticeable to one user becomes a UX crisis under load. Data pipelines that work fine on small datasets break in ways that are hard to predict without depth.</p><p>The claim that &#8220;coding is no longer the bottleneck&#8221; may be true in the simplest corner of this matrix &#8212; internal tools, prototype-grade, unregulated, small scale. It becomes progressively less true as you move in any direction.</p><div><hr></div><p><strong>Engineering Judgment Is Still Judgment</strong></p><p>There&#8217;s a subtler problem buried in how &#8220;judgment&#8221; is being used right now: it&#8217;s being treated as a single thing when it&#8217;s actually several.</p><p>Product judgment &#8212; understanding user problems, prioritizing correctly, knowing when good enough is good enough &#8212; that&#8217;s one kind. Design judgment &#8212; taste in visual hierarchy, interaction patterns, information density &#8212; that&#8217;s another. Engineering judgment is another &#8212; and it doesn&#8217;t disappear because LLMs can write code.</p><p>Engineering judgment is knowing that a given API&#8217;s rate limits will create a bad user experience under realistic load, before you find out by watching it fail. It&#8217;s recognizing that a context window limitation will cap what&#8217;s possible for your use case, and structuring your approach accordingly. It&#8217;s knowing when the cost of using an LLM makes it the wrong tool for a particular subtask, even if it would technically work.</p><p>When I was building the nature access map and hit a frontend rendering bug that sent Claude into circles &#8212; each fix creating a new problem &#8212; what broke the loop wasn&#8217;t better prompting. It was recognizing that we needed browser developer tools to surface actual error information, rather than continuing to diagnose from assumptions. That was a judgment call, made by a human, that the LLM hadn&#8217;t arrived at on its own. And this was a simple, single-user, locally-run application. The engineering complexity scales from here fast.</p><p>You can&#8217;t decouple engineering decisions from user experience decisions, or from business decisions. The people who understood this &#8212; who made proactive choices about downstream needs before being asked, who proposed building test infrastructure before it became obviously necessary, who recognized when a technically feasible solution would create maintenance debt that wasn&#8217;t worth it &#8212; were always valuable because of those judgment behaviors. They just happened to express them through engineering work.</p><p>Vibe coding doesn&#8217;t make that less relevant. It reduces the coordination costs of separating judgment from execution across different people. That&#8217;s genuinely useful. It doesn&#8217;t synthesize the judgment.</p><div><hr></div><p><strong>The Economics of It</strong></p><p>When the cost of development decreases, the value of complementary capabilities increases. This is basic economics &#8212; cheaper peanut butter raises demand for jelly &#8212; and it explains why judgment, taste, and domain expertise are getting more attention right now.</p><p>But the implied corollary &#8212; that the previous bottleneck disappears &#8212; doesn&#8217;t follow. It becomes less constraining for simpler cases. It remains fully present for complex ones. And new bottlenecks emerge in places previously masked by the old one.</p><p>When I built the nature access map, part of what made it worth building was that the search cost of finding a tool that answered my specific question exceeded the development cost of building it myself. That math used to reliably go the other direction. Now it sometimes flips. That&#8217;s a real shift &#8212; but one that made domain knowledge more load-bearing, not less. The development cost dropped. The judgment requirements didn&#8217;t.</p><p>There&#8217;s a second economic dynamic worth naming. When anyone can build, the supply of apps proliferates rapidly. At some point supply exceeds demand. And in a buyer&#8217;s market, buyers have two levers: pay less for the same thing, or demand higher quality. These are in tension &#8212; as quality requirements increase, fewer people can meet them, which changes pricing again. But the net effect is that &#8220;good enough&#8221; stops being sufficient in the same way it was when supply was scarce. When you were one of a few people who could ship anything at all, the bar was low by necessity. When thousands of people are shipping similar things, the baseline shifts. The judgment required to clear it doesn&#8217;t go away. It goes up.</p><div><hr></div><p><strong>The Open Question</strong></p><p>We&#8217;ve had &#8220;low-code&#8221; and &#8220;no-code&#8221; tools for years. Their history is consistent: useful for prototypes, insufficient for production, high ratio of demos to shipped products. They lowered the floor without raising the ceiling. People built more things, most of which didn&#8217;t make it.</p><p>The open question is whether this generation of tools &#8212; Claude Code and its successors &#8212; finally crosses that threshold. Whether the production gap will actually be closed for a meaningful range of use cases in a way it wasn&#8217;t before.</p><p>My honest answer is: maybe. Given enough time, probably, for some categories of use cases and some levels of complexity. But we don&#8217;t fully know yet, and the people making the strongest claims &#8212; in either direction &#8212; are probably overstating their case.</p><p>What I do know is that the judgment question isn&#8217;t new. It was always the hard part. The fact that more people can now see it &#8212; because more people can now build things and see the results &#8212; doesn&#8217;t mean it recently emerged from new tools.</p><p>It was always there, hidden in the gap between shipping and shipping something worth using.</p><div><hr></div><p><em>This is Article 5 in a series on building with AI tools as a non-native technologist &#8212; covering my real experiences and observations building products as someone without a formal engineering background. Previous articles covered financial planning with Claude Code, a framework for city selection, using LLM tools to help with a cross-country move, and building the Seattle Nature Access Map.</em></p>]]></content:encoded></item><item><title><![CDATA[I Couldn’t Find the App I Needed, So I Built It. Here’s What I Learned.]]></title><description><![CDATA[I spent weeks asking ChatGPT, Google, and every tool I could find the same question: if I move to this neighborhood, how close is the nearest park I&#8217;d actually use?]]></description><link>https://thecontractsignal.com/p/i-couldnt-find-the-app-i-needed-so</link><guid isPermaLink="false">https://thecontractsignal.com/p/i-couldnt-find-the-app-i-needed-so</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sun, 01 Mar 2026 20:01:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HWql!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I spent weeks asking ChatGPT, Google, and every tool I could find the same question: if I move to this neighborhood, how close is the nearest park I&#8217;d actually use?</p><p>No one could answer it well.</p><p>Google Maps can route you between two points. LLMs can tell you that Seattle &#8220;has great parks.&#8221; Walk Score can give you a general walkability number. But none of them answered my actual question: for a given block within a given neighborhood, what&#8217;s the walking distance to a meaningful green space?</p><p>So I built it myself.</p><p>This is the story of how a personal frustration during my move from New York to Seattle became my first public app&#8212;built with Claude as a coding partner, shipped as an MVP over a handful of nights and weekends, and designed to answer one specific question that no existing tool answered well.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HWql!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HWql!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 424w, https://substackcdn.com/image/fetch/$s_!HWql!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 848w, https://substackcdn.com/image/fetch/$s_!HWql!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 1272w, https://substackcdn.com/image/fetch/$s_!HWql!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HWql!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png" width="1456" height="861" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:861,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A map of the united states\n\nAI-generated content may be incorrect.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A map of the united states

AI-generated content may be incorrect." title="A map of the united states

AI-generated content may be incorrect." srcset="https://substackcdn.com/image/fetch/$s_!HWql!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 424w, https://substackcdn.com/image/fetch/$s_!HWql!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 848w, https://substackcdn.com/image/fetch/$s_!HWql!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 1272w, https://substackcdn.com/image/fetch/$s_!HWql!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F13a712e9-b781-402d-b51b-fb72c7fa761b_2510x1484.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h4><strong>Why This Existed as a Problem</strong></h4><p>In my previous articles, I wrote about the framework I used to pick a city and how I used LLMs during the planning process. One of my core criteria was convenient walking access to nature&#8212;not &#8220;nature exists somewhere in the metro area,&#8221; but &#8220;I can walk to a real park in under 15 minutes from my front door.&#8221;</p><p>This mattered to me because of lived experience, not aspiration. Living in New York, I&#8217;d discovered that regular morning walks&#8212;first through Carl Schurz Park, then Central Park&#8212;were one of the highlights of my daily life. I felt measurably better on days I walked. Over time, I realized that a city with great restaurants and nightlife but mediocre park access would be a worse fit for my actual life than a city with fewer amenities but a large park within walking distance.</p><p>The problem was that once I had this criterion, I couldn&#8217;t evaluate it efficiently.</p><p>I did rounds of Google searches. I asked LLMs. I checked existing nature and hiking apps. Nothing quite fit. There are plenty of apps for finding trails or planning hikes&#8212;but I wasn&#8217;t looking for a weekend adventure. I wanted to know: if I live on this block, can I walk to a good park easily? And how does that compare to the next block over, or the next neighborhood?</p><p>That question turns out to be surprisingly hard to answer without building something.</p><div><hr></div><h4><strong>When Finding Is Harder Than Building</strong></h4><p>This experience led me to a realization I hadn&#8217;t expected: for a sufficiently specific personal need, the search costs of finding the right tool can actually exceed the development costs of building one yourself.</p><p>Historically, that wouldn&#8217;t have been true. Building a map-based web app required professional software engineering skills and weeks or months of development time. The rational move was always to search harder, compromise on fit, or just do the research manually.</p><p>But something has shifted. Between vibe coding with LLM tools like Claude, publicly available geographic data, and open-source mapping libraries, the cost of building a simple, personalized tool has dropped dramatically. Meanwhile, the cost of searching&#8212;scrolling through apps that almost-but-don&#8217;t-quite solve your problem, configuring tools not designed for your use case, cobbling together manual research across multiple sources&#8212;hasn&#8217;t changed much.</p><p>For me, the math flipped. I&#8217;d already spent hours on manual research and still didn&#8217;t have what I wanted. Building the thing I actually needed took less additional time than continuing to search for it would have.</p><p>This has broader implications for anyone building with AI tools. As development costs decrease, the value of complementary capabilities increases&#8212;things like judgment about what to build, taste in how it should work and look, and access to the right data. The bottleneck isn&#8217;t &#8220;can I code this?&#8221; anymore. It&#8217;s &#8220;do I understand the problem well enough to build the right thing?&#8221;</p><div><hr></div><h4><strong>What I Built</strong></h4><p>The concept was deliberately simple: a color-coded map of Seattle showing walking distance to parks, broken down by neighborhood and then by census block.</p><p>That&#8217;s it. One question, answered visually, at a glance.</p><p><strong>Why a map and not a report?</strong> I actually tried the text-based approach first. I ran LLM queries and compiled natural language descriptions of each neighborhood&#8217;s park access. It wasn&#8217;t very useful. Geographic and spatial relationships are fundamentally hard to convey in text. &#8220;Wallingford is 0.8 miles from Woodland Park&#8221; is less informative than seeing Wallingford on a map, colored green, with a park icon nearby. The visual format communicates distance, relative position, and comparison across neighborhoods simultaneously&#8212;something that would take paragraphs to express in words and still wouldn&#8217;t land as intuitively.</p><p>This connects to a broader principle I&#8217;ve been thinking about: the right interface depends on the type of information. Natural language works well for conceptual questions and subjective reasoning. But for spatial data, structured comparisons, or anything where relationships between data points matter as much as the data points themselves&#8212;a visual interface is almost always better.</p><p><strong>The flow works like this:</strong></p><p>Start with the city overview. Every neighborhood is color-coded from green (excellent park access) to red (poor access), scored 0&#8211;100. You can see at a glance which areas of Seattle have the best walking access to nature. Hover over a neighborhood and you see its score.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2hWq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2hWq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 424w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 848w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 1272w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2hWq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png" width="1456" height="862" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:862,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A map of different colored areas\n\nAI-generated content may be incorrect.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A map of different colored areas

AI-generated content may be incorrect." title="A map of different colored areas

AI-generated content may be incorrect." srcset="https://substackcdn.com/image/fetch/$s_!2hWq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 424w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 848w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 1272w, https://substackcdn.com/image/fetch/$s_!2hWq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F585740e0-9aa0-488f-bbd4-d82bc95c2504_2036x1206.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Click into a neighborhood. Now you see the individual census blocks within it, each with its own color and score. This matters because neighborhood averages can be misleading&#8212;a neighborhood might score 65 overall, but the blocks closest to a park could be 85+ while blocks on the far edge might be 35.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xzRA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xzRA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 424w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 848w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 1272w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xzRA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png" width="1456" height="756" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:756,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A map of a city\n\nAI-generated content may be incorrect.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A map of a city

AI-generated content may be incorrect." title="A map of a city

AI-generated content may be incorrect." srcset="https://substackcdn.com/image/fetch/$s_!xzRA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 424w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 848w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 1272w, https://substackcdn.com/image/fetch/$s_!xzRA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e00322a-7542-42c1-8a1c-57814100de58_2704x1404.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Click a specific block. You see the address, the nature access score, the nearest park name, the distance, the block area, and a link to view it on Google Maps.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WjFi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WjFi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 424w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 848w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 1272w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WjFi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png" width="1456" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A screenshot of a map\n\nAI-generated content may be incorrect.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A screenshot of a map

AI-generated content may be incorrect." title="A screenshot of a map

AI-generated content may be incorrect." srcset="https://substackcdn.com/image/fetch/$s_!WjFi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 424w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 848w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 1272w, https://substackcdn.com/image/fetch/$s_!WjFi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e071aac-e948-415d-b816-ef0004727c2d_2682x1452.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The whole thing follows a principle called progressive disclosure: show the minimum useful information first, then let people drill deeper based on what they care about. For someone just exploring neighborhoods, the city overview is enough. For someone evaluating a specific apartment, the block-level detail is what matters.</p><div><hr></div><h4><strong>How I Built It (And What Was Harder Than Expected)</strong></h4><p>I used Claude as my coding partner for the entire build. The overall process followed what most people describe as vibe coding: provide a high-level plan, iterate step by step, and make trade-off decisions along the way.</p><p><strong>The plan was conceptually simple.</strong> I needed three things: data (parks, neighborhoods, geographic boundaries), a frontend (the map interface), and backend logic (calculating distances, generating scores, connecting everything). This is essentially the model-view-controller pattern&#8212;data, interface, and logic working together.</p><p><strong>Where it got interesting was the trade-offs.</strong></p><p>The data layer was the biggest source of real decisions&#8212;not because getting data was impossible, but because every data choice had downstream consequences.</p><p>First, how do you define a &#8220;park&#8221;? Is a 2-acre playground sufficient? What about a botanical garden? A greenway? For MVP, I kept it simple: parks above a minimum size threshold, focused on the kind of green space you&#8217;d actually walk through for 20+ minutes. That meant leaving out some edge cases, but it kept the scope manageable.</p><p>Second, where does the data come from? Publicly available geographic datasets exist, but they come with trade-offs: API rate limits, licensing restrictions, varying data quality, and inconsistent naming conventions across sources. Park data uses one taxonomy. Neighborhood boundary data uses another. Census block data uses yet another. Making these work together required judgment calls&#8212;how to define neighborhood boundaries, how to handle edge cases where a block spans two neighborhoods, what to do when data sources disagree.</p><p>Third, how do you serve the data? Making live API requests for every user interaction would be slow and fragile&#8212;and could hit rate limits fast if multiple people use the app simultaneously. So I downloaded the data upfront and serve it locally. For an MVP with a small user base, this works well and keeps the experience fast. Scaling would require a different approach, but that&#8217;s a future problem.</p><p><strong>The frontend was where taste mattered most.</strong> The color scheme, the scoring scale, the level of detail in the popup, the decision to show census blocks rather than arbitrary grid squares&#8212;these are all judgment calls that determine whether the app feels useful or confusing. No amount of coding skill replaces knowing what the user (in this case, initially me) actually wants to see.</p><p><strong>The backend logic surfaced subtle complexity.</strong> Distance calculations between geographic points aren&#8217;t as straightforward as they appear&#8212;coordinate systems, projection methods, and the difference between straight-line distance and walking distance all matter. For MVP, I used a reasonable approximation that&#8217;s good enough for &#8220;is this a 5-minute walk or a 20-minute walk?&#8221; If I need higher precision later, I can refine it.</p><div><hr></div><h4><strong>Where Vibe Coding Helped&#8212;and Where It Didn&#8217;t</strong></h4><p>Building with Claude as a coding partner was significantly faster than building alone would have been. The LLM was particularly good at generating boilerplate code for map rendering and data processing, suggesting library choices and architectural patterns, helping debug issues when I could describe the problem clearly, and handling repetitive tasks like data formatting and conversion.</p><p>Where it struggled was when problems required context that&#8217;s hard to express in a prompt. For example, I hit a rendering bug where the frontend wasn&#8217;t displaying blocks correctly. Claude tried multiple fixes, each of which introduced a new problem&#8212;we went in circles. What eventually broke the loop was me suggesting we use browser developer tools to inspect the actual errors, which gave us concrete diagnostic information to work with. In another case, a more capable model release solved a problem the previous model couldn&#8217;t crack.</p><p>This reinforced something I&#8217;ve come to believe through this process: understanding computer science fundamentals matters, even in a vibe coding world. Not because you need to write every line yourself, but because when something breaks&#8212;and it will&#8212;you need enough understanding to diagnose the problem, ask the right questions, and evaluate whether the LLM&#8217;s proposed fix makes sense. If you abstract away all understanding of how the system works, you&#8217;re stuck the moment something goes wrong.</p><p>That said, I want to be honest about the tradeoff: for a simple, single-user application like this, the fundamentals matter less than they would for a production system serving thousands of users. The stakes of a bug are low. The complexity is manageable. If you&#8217;re building something for yourself and learning as you go, vibe coding is a remarkably effective way to do both simultaneously. Just know that the further you scale&#8212;more users, more data, more features&#8212;the more that underlying understanding pays off.</p><div><hr></div><h4><strong>Lessons Learned</strong></h4><p><strong>Build for yourself first, then share.</strong> I built this because I genuinely wanted it. That meant I understood the problem deeply, had strong opinions about what the solution should look like, and could evaluate quality without user testing. Solving your own problem is the fastest path to a useful product.</p><p><strong>Start with one question.</strong> The app answers one question: how walkable is the nearest park from this location? Every feature decision flowed from that. When I was tempted to add hiking trail data, transit overlays, or restaurant proximity, I asked: does this help answer the core question? If not, it&#8217;s post-MVP.</p><p><strong>Data is the real bottleneck, not code.</strong> The most time-consuming part wasn&#8217;t writing code&#8212;it was finding reliable data sources, understanding their limitations, reconciling different taxonomies, and making judgment calls about quality. As vibe coding makes development easier, data access and data quality become key differentiators for applications.</p><p><strong>The right interface matters more than the right model.</strong> I tried answering this question with LLMs (natural language), with manual research (spreadsheets and notes), and with a visual map. The map won&#8212;not because it has more information, but because it presents spatial information in the format humans process most naturally. Choosing the right interface for your information type is a design decision that no amount of better AI can substitute for.</p><p><strong>Ship the MVP, then iterate.</strong> The current version has rough edges. The scoring could be more sophisticated. The data could be more comprehensive. The styling could be more polished. None of that matters as much as having a working tool that answers the core question. Everything else is refinement.</p><div><hr></div><h4><strong>What&#8217;s Next</strong></h4><p>The immediate plan is to continue refining the app&#8212;improving the interface, exploring additional cities, and seeing whether other people find it as useful as I do.</p><p>But the bigger takeaway from this experience is about a shift that&#8217;s happening in who can build useful tools. A few years ago, building a map-based web application required a team of engineers. Today, a product manager with some CS background and an LLM coding partner can ship a functional MVP in weeks.</p><p>That doesn&#8217;t mean the engineering is trivial. It means the barrier to entry has moved. The scarce resource is no longer &#8220;can you code it&#8221; but &#8220;do you understand the problem well enough, and do you have the judgment to make the right trade-offs?&#8221;</p><p>For me, this started as solving my own problem. It turned into my first public app. And it taught me more about building products than years of managing them from the other side of the table.</p><p>The app is available here: <a href="http://natureaccessmap.com">nature-access-map</a>.</p><p><em>This is the fourth article in a series. <a href="https://leonidprilutskiy.substack.com/p/claude-code-for-financial-planning?r=1iyonb">Article 1</a> covered using Claude Code for financial planning. <a href="https://leonidprilutskiy.substack.com/p/i-spent-10-years-in-the-wrong-city?r=1iyonb">Article 2</a> covered the framework I used to evaluate potential cities to move to. <a href="https://leonidprilutskiy.substack.com/p/i-used-chatgpt-and-claude-to-help?r=1iyonb">Article 3</a> covered where LLMs helped and failed during the planning process.</em></p>]]></content:encoded></item><item><title><![CDATA[I Used ChatGPT and Claude to Help Pick a New City. Here’s Where They Helped—and Where They Didn’t.]]></title><description><![CDATA[This is the third article in a series on using AI tools for real-world decisions.]]></description><link>https://thecontractsignal.com/p/i-used-chatgpt-and-claude-to-help</link><guid isPermaLink="false">https://thecontractsignal.com/p/i-used-chatgpt-and-claude-to-help</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 23 Feb 2026 02:26:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is the third article in a series on using AI tools for real-world decisions. <a href="https://open.substack.com/pub/leonidprilutskiy/p/claude-code-for-financial-planning?r=1iyonb&amp;utm_campaign=post&amp;utm_medium=web">Article 1</a> covered using Claude Code for financial planning. <a href="https://leonidprilutskiy.substack.com/p/i-spent-10-years-in-the-wrong-city?r=1iyonb">Article 2</a> covered the framework I used to evaluate cities based on my actual day-to-day life.</em></p><div><hr></div><p>I asked ChatGPT for the best cities to live in matching my criteria.</p><p>It gave me Portugal.</p><p>I asked for mild weather.</p><p>It gave me South Carolina.</p><p>That&#8217;s when I realized: LLMs are powerful research tools, but they&#8217;re solving for the average person&#8212;not for you.</p><p>Over three months of planning a cross-country move, I learned exactly where AI tools help and where they fail. The key wasn&#8217;t finding the &#8220;right prompt.&#8221; It was understanding which parts of the problem are objective&#8212;and which are subjective&#8212;and treating them completely differently.</p><p>Here&#8217;s what I found.</p><div><hr></div><h3><strong>The Real Problem: Your Criteria Aren&#8217;t as Clear as You Think</strong></h3><p>When I started this process, I had what I thought were well-defined criteria: mild weather, walkability, access to nature, reasonable cost of living.</p><p>But here&#8217;s the thing&#8212;most of those words don&#8217;t actually mean anything specific.</p><p>&#8220;Mild weather&#8221; means something different if you walk everywhere carrying groceries versus if you drive. I&#8217;ve done both. I&#8217;ve walked through Atlanta summers carrying grocery bags, and I can tell you that 85&#176;F with humidity when you&#8217;re on foot for 20 minutes is a fundamentally different experience than 85&#176;F when you&#8217;re stepping from an air-conditioned car into an air-conditioned store.</p><p>&#8220;Walkability&#8221; means something different if you work remotely and mainly care about parks and errands versus if you commute downtown daily.</p><p>&#8220;Access to nature&#8221; means something different if you want weekend alpine hikes versus a daily walk through a 50+ acre park within 15 minutes of your front door.</p><p>The hard part isn&#8217;t that these criteria are wrong. It&#8217;s that they contain hidden assumptions&#8212;and you often don&#8217;t realize it until you get results that feel off.</p><p>Part of the reason is that many preferences are internalized. You don&#8217;t actively think about them. They&#8217;re visceral. You know the Atlanta summer grocery run feels bad, but you might not think to tell an LLM, &#8220;I walk to the grocery store three times a week and I carry bags, so your temperature recommendation needs to account for extended pedestrian exposure in humidity.&#8221; That level of context doesn&#8217;t come naturally.</p><p>It&#8217;s similar to how a UX researcher has to coax out a user&#8217;s real workflow&#8212;people don&#8217;t volunteer the details that matter most because those details feel obvious to them.</p><div><hr></div><h3><strong>The Insight That Changed Everything: Objective vs. Subjective</strong></h3><p>After weeks of getting well-intentioned but unhelpful results, I noticed a pattern.</p><p>LLMs performed very differently depending on whether my question was <strong>objective</strong> or <strong>subjective</strong>&#8212;and much of the frustration came from treating subjective questions as if they were objective.</p><p><strong>Objective questions</strong> have verifiable, well-structured answers. Historical temperature data. Cost-of-living indices. Transit system maps. Average precipitation by month. These are previously solved, commonly asked, publicly available, and quantifiable. LLMs are excellent at these&#8212;often as good as or better than manual research because they can synthesize across sources quickly.</p><p><strong>Subjective questions</strong> involve personal interpretation, context-dependent meaning, or ambiguous language. &#8220;Is it mild?&#8221; &#8220;Is this neighborhood walkable?&#8221; &#8220;Is there good nature access?&#8221; These feel like they should have clear answers, but they don&#8217;t&#8212;because the answer depends on who&#8217;s asking and why.</p><p>The mistake I kept making was asking subjective questions and expecting objective answers.</p><p>When I asked for cities with &#8220;mild weather,&#8221; the LLM had to interpret &#8220;mild.&#8221; Its interpretation was reasonable&#8212;but it wasn&#8217;t mine. South Carolina has mild winters. It also has summers that would make my grocery walks miserable. The LLM didn&#8217;t know that because I hadn&#8217;t connected the dots between &#8220;mild&#8221; and &#8220;I walk carrying bags in the heat.&#8221;</p><p>This distinction&#8212;objective vs. subjective&#8212;became the foundation for everything else I did.</p><div><hr></div><h3><strong>Working With Objective Criteria: Let the LLM Do the Heavy Lifting</strong></h3><p>For objective criteria, LLMs were genuinely excellent.</p><p>Once I learned to ask for data instead of adjectives, the quality of answers jumped dramatically.</p><p>Instead of: <em>&#8220;Which of these cities has mild weather?&#8221;</em></p><p>I asked: <em>&#8220;For each of these 15 cities, what is the average, median, and range of monthly high temperatures over the last 5 years? Include humidity and precipitation.&#8221;</em></p><p>Instead of: <em>&#8220;Is Raleigh walkable?&#8221;</em></p><p>I asked: <em>&#8220;What is Raleigh&#8217;s Walk Score? What percentage of residents commute without a car?&#8221;</em></p><p>The principle is simple: if a question can be answered with publicly available, well-structured data, ask for the data directly. Don&#8217;t ask the LLM to interpret it for you.</p><p>This worked well for criteria like temperature ranges, cost-of-living comparisons, transit infrastructure, general park and green space data, and airport access and logistics.</p><p>The results weren&#8217;t always perfect&#8212;LLMs can be confidently wrong on specifics, and data can be outdated&#8212;but they were directionally reliable. Good enough to build a shortlist. Good enough to eliminate obvious misfits.</p><div><hr></div><h3><strong>Working With Subjective Criteria: That&#8217;s Your Job (But LLMs Can Help You Think)</strong></h3><p>The subjective criteria were where things got interesting&#8212;and where I had to change my approach entirely.</p><p>For these, the LLM&#8217;s role shifted from &#8220;researcher&#8221; to &#8220;thought partner.&#8221; Instead of asking it to give me answers, I used it to help me figure out what my questions actually meant.</p><p><strong>Technique 1: Benchmark against cities you know.</strong></p><p>Rather than trying to define &#8220;mild&#8221; in the abstract, I used cities I&#8217;d actually lived in as reference points.</p><p><em>&#8220;I&#8217;ve lived in New York and Atlanta. Compare the summer pedestrian experience in Raleigh, Seattle, and Pittsburgh to those two cities&#8212;specifically for someone who walks daily and carries groceries.&#8221;</em></p><p>This forced specificity. The LLM couldn&#8217;t hide behind &#8220;mild climate.&#8221; It had to compare against my actual lived experience.</p><p><strong>Technique 2: Ask yourself why&#8212;and tell the LLM.</strong></p><p>When I caught myself using a vague word, I&#8217;d ask: why do I actually care about this?</p><p>&#8220;Mild weather&#8221;&#8212;why? Because I walk for exercise and errands daily. Because I don&#8217;t want to dread going outside for four months of the year. Because I don&#8217;t want to rely on AC or heating as a coping mechanism for a climate I fundamentally don&#8217;t enjoy.</p><p>Each of those &#8220;becauses&#8221; translates into a more specific, testable question. And once I fed that context to the LLM, the recommendations improved noticeably.</p><p><strong>Technique 3: Pressure-test with specific examples.</strong></p><p>When an LLM told me a city had &#8220;good nature access,&#8221; I&#8217;d follow up: <em>&#8220;What specific parks are within 20 minutes walking distance of downtown? How large are they? Do they have trails or is it a playground?&#8221;</em></p><p>The specifics often revealed that &#8220;good nature access&#8221; meant something very different from what I needed. A city park with a playground and a walking loop is not the same as a 300-acre urban forest with trail networks. Both are technically &#8220;nature access.&#8221;</p><p>This is actually a pattern I recognized from a previous career chapter. Years ago, I managed machine learning annotation projects&#8212;labeling training data for AI models. Even with seemingly well-defined categories, annotators would apply labels subjectively. The fix was always the same: go through specific examples, compare them, explain your reasoning, surface the disagreements, and refine the criteria. The same principle applies when you&#8217;re &#8220;annotating&#8221; your own preferences for an LLM.</p><div><hr></div><h3><strong>The Division of Labor</strong></h3><p>After a few weeks, a clear division of labor emerged:</p><p><strong>I used LLMs for</strong> thought partnership to pressure-test my criteria, researching and summarizing candidates against those criteria, helping me narrow down on specific dimensions, drafting checklists (apartment tours, moving logistics, donation steps), and estimating logistics and surfacing things I might be missing.</p><p><strong>I kept for myself</strong> defining what I actually want and how to prioritize it, calibrating subjective labels (&#8221;mild,&#8221; &#8220;walkable,&#8221; &#8220;quiet&#8221;), verifying anything that could cost money or time, and running real-world tests (walking neighborhoods, simulating daily life, trying errands).</p><p>The frame that made it click: LLMs are great at research, organization, and estimation. They&#8217;re not built for judgment&#8212;and personal decisions are mostly judgment.</p><p>Once I internalized this, the tool stopped being a frustrating &#8220;decision maker&#8221; and became a genuine force multiplier.</p><div><hr></div><h3><strong>Top-Down First, Then Bottom-Up</strong></h3><p>A mistake I made early was trying to get the perfect answer in one shot.</p><p>What worked was a two-phase approach:</p><p><strong>Top-down (cast a wide net):</strong> I asked for 20&#8211;40 candidate cities given my criteria. I deliberately wanted a long list with false positives, because false negatives are more costly&#8212;if you never consider a city, you never discover it&#8217;s a fit.</p><p>At this stage, I accepted messy, imperfect results. The goal wasn&#8217;t precision. It was coverage&#8212;and using the results to refine my own thinking about what mattered.</p><p><strong>Bottom-up (go deep on the shortlist):</strong> Once I&#8217;d narrowed to a handful of cities, I shifted to targeted, specific queries. Neighborhood-level research. Specific parks, transit routes, grocery stores. Temperature data by month. This is where the objective/subjective framework paid off most&#8212;I knew which questions to ask the LLM (data) and which to answer myself (judgment).</p><p>The key insight: city-level averages hide neighborhood-level reality.</p><p>A city can be &#8220;unwalkable&#8221; overall and still have pockets that are perfect for a car-free lifestyle. Or the reverse. So I treated each promising city&#8217;s neighborhoods as a separate research problem&#8212;and that&#8217;s actually where the decision happened.</p><p><strong>Where LLMs Consistently Fell Short</strong></p><p>A few patterns emerged:</p><p><strong>Ambiguous terms without operationalization.</strong> Every time I used a word like &#8220;mild,&#8221; &#8220;safe,&#8221; &#8220;quiet,&#8221; or &#8220;diverse&#8221; without defining it, I got generic results. The fix was always the same: replace adjectives with numbers, ranges, or specific examples.</p><p><strong>Hidden assumptions.</strong> The LLM would assume I had a car, or that &#8220;affordable&#8221; meant a certain range, or that &#8220;city&#8221; meant metro area rather than a specific neighborhood. The fix: ask the model to list its assumptions, or preempt by stating constraints (&#8221;assume no car,&#8221; &#8220;budget is Y&#8221;).</p><p><strong>Trade-off weighting.</strong> If you ask an LLM to optimize across 10 criteria simultaneously, you&#8217;re asking it to make value judgments about your life. It can&#8217;t do that. The fix: ask for a longer list with trade-offs noted, and do the weighting yourself.</p><p><strong>Freshness and specificity.</strong> LLMs can be confidently wrong on details that change&#8212;pricing, availability, neighborhood dynamics. The fix: treat every output as a lead to verify, not a fact to trust.</p><div><hr></div><h3><strong>The Limits of Research (And Why Real-World Testing Matters)</strong></h3><p>At some point, no amount of prompting eliminates uncertainty. You have to test.</p><p>I found that the best way to reduce uncertainty was simply visiting&#8212;or simulating a visit as closely as possible. Walking the actual routes I&#8217;d walk. Checking the actual grocery stores I&#8217;d use. Riding the actual transit I&#8217;d depend on.</p><p>This is also where I learned that not everything should be an LLM conversation. Some questions are better answered by a map. Some by a transit app. Some by walking around for an hour.</p><p>I actually started building a tool for one of these gaps&#8212;a neighborhood park access map that shows walking distance to parks at the block level&#8212;because it was a question I kept asking that no existing tool (including LLMs) answered well. I&#8217;ll share more about that in the next article.</p><p>The principle: LLMs generate hypotheses. Targeted tools and real-world experience verify them.</p><div><hr></div><h3><strong>What I&#8217;d Do Differently</strong></h3><p>If I were starting this process from scratch, I&#8217;d do three things earlier:</p><p><strong>First, separate objective from subjective criteria on day one.</strong> This alone would have saved me weeks of frustration. Objective criteria get delegated to the LLM immediately. Subjective criteria get a &#8220;define what I actually mean&#8221; session before any research happens.</p><p><strong>Second, use benchmark cities from the start.</strong> Instead of trying to define preferences in the abstract, ground everything in places you&#8217;ve actually experienced. It&#8217;s faster and more reliable.</p><p><strong>Third, go neighborhood-level sooner.</strong> City-level research is useful for elimination, but the decision lives at the neighborhood level. I spent too long comparing cities when I should have been comparing neighborhoods within the top 3&#8211;5.</p><div><hr></div><h3><strong>Conclusion</strong></h3><p>Using LLMs for this move didn&#8217;t eliminate uncertainty. Moving is still a leap.</p><p>What it did was compress the process&#8212;from fuzzy to structured, from overwhelming to manageable, from &#8220;I don&#8217;t even know where to start&#8221; to &#8220;I have a clear shortlist and I know what to test.&#8221;</p><p>The trick was learning that the most important skill isn&#8217;t prompting. It&#8217;s knowing which parts of the problem are yours to solve&#8212;and using the tool to handle everything else.</p><p>LLMs are great at helping you think. They&#8217;re great at helping you research and organize.</p><p>They&#8217;re not great at being you.</p><p><em>Next article: I built a neighborhood park access map because no existing tool answered my question. Here&#8217;s what I learned about going from an internal tool to a public product.</em></p>]]></content:encoded></item><item><title><![CDATA[I Spent 10 Years in the Wrong City. Here’s The Framework I Wish I Had.]]></title><description><![CDATA[Conventional wisdom on selecting cities tends to be overly generic.]]></description><link>https://thecontractsignal.com/p/i-spent-10-years-in-the-wrong-city</link><guid isPermaLink="false">https://thecontractsignal.com/p/i-spent-10-years-in-the-wrong-city</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 16 Feb 2026 03:22:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Conventional wisdom on selecting cities tends to be overly generic.</p><p>I know this because I spent ten years living in a New York City &#8211; a city that looks great on paper.</p><p>It&#8217;s globally recognized, full of opportunity, and has every possible convenience. And yet, over time, it just didn&#8217;t feel right for me.</p><p>It wasn&#8217;t a bad experience. It just felt like I was in the wrong place for reasons I couldn&#8217;t quite articulate.</p><p>I&#8217;d had this feeling for a long time.</p><p>But given the range of possible cities to choose from, the uncertainty and ambiguity around them, and a wide range of experiences being reported&#8230;</p><p>Information overload and decision paralysis followed.</p><p>Combine that with psychological tendencies towards pain avoidance (the delightful moving process) and doubt avoidance (how do I even know the new city will be better?) &#8211; and slight variations of the status quo becomes the norm.</p><p>I.e. stay in the same city.</p><p>When I started planning a possible move late last year, I tried to address this by getting advice from both people who knew me (and of course, LLMs).</p><p>The advice was well-intentioned but mostly unhelpful.</p><p>Not that I felt like anyone was trying to mislead me&#8212;but because the advice wasn&#8217;t personalized to my life. People were giving me recommendations filtered through <em>their</em> definitions of what matters.</p><p>Someone says, &#8220;Don&#8217;t move to Seattle, it rains a lot.&#8221;</p><p>Someone else says, &#8220;You should move to Nashville or Austin.&#8221; When pressed, it&#8217;s because they took a two-day business trip, went to great restaurants, and enjoyed the vibe.</p><p>Even well-vetted advice can be wrong in context. When I lived in Atlanta for a few years, the default line was: <em>you need a car.</em> Maybe that&#8217;s true for many lifestyles. It wasn&#8217;t true for mine. What mattered more was how I actually spent my day: school, studying, errands, gym, basic routines.</p><p>I lived near where I needed to be, walked more than most people assumed, and used taxis and rideshares strategically. In Atlanta, monthly parking was $~150. Ubers averaged less than $20/week.</p><p>The city&#8217;s reputation didn&#8217;t match my lived experience&#8212;because my routine didn&#8217;t match the &#8220;average&#8221; routine implied by the reputation.</p><p>That&#8217;s the fundamental problem:</p><p><strong>Most advice about choosing a city fails because it isn&#8217;t personalized.</strong></p><p>The good news is: this is solvable.</p><p>It just requires a framework that starts with your real day-to-day life.</p><div><hr></div><p><strong>Step 1: Audit your day-to-day (the foundational move)</strong></p><p>Before you research cities, ask a simpler question:</p><p><strong>What does my life actually look like right now?</strong></p><p>For a couple weeks, track your routine at a practical level. Not &#8220;I value health.&#8221; Not &#8220;I love nature.&#8221; The actual sequence:</p><ul><li><p>Do you work remotely or commute? If you do commute, how?</p></li><li><p>How often do you leave the house on weekdays?</p></li><li><p>Do you cook, buy pre-cooked meals, or mostly eat out?</p></li><li><p>What errands happen weekly (groceries, pharmacy, gym, laundry)?</p></li><li><p>What do you do on weekends when you have free time?</p></li><li><p>What do you do for fun / to blow off steam?</p></li></ul><p>This is where people get tripped up. They choose cities based on identity (&#8220;I&#8217;m a nature person&#8221;), aspiration (&#8220;I&#8217;m going to hike every weekend&#8221;), or aesthetics (&#8220;I want charm&#8221;).</p><p>None of those are inherently wrong.</p><p>But they&#8217;re unreliable inputs if they aren&#8217;t grounded in practical behavior and habits.</p><p>One of the biggest insights I had was realizing that some &#8220;preferences&#8221; were actually <em>requirements</em> for my mood, health, and productivity&#8212;and some things I thought mattered were mostly nice-to-haves.</p><p>Example: nature access. It didn&#8217;t become real for me because I wrote it down as a value. It became real when I started walking in the park regularly&#8212;early, even when it was cold, even when the weather wasn&#8217;t ideal. That behavior was evidence. It told me: this isn&#8217;t a nice-to-have or a fantasy. It&#8217;s a proven requirement for my wellbeing.</p><div><hr></div><p><strong>Step 2: Turn your routine into criteria (make it testable)</strong></p><p>Once you understand your day-to-day, convert it into criteria that are specific enough to evaluate.</p><p>Start with broad categories (most people share these):</p><ul><li><p>Cost of living (true cost, not just rent)</p></li><li><p>Housing (size, age, noise, amenities)</p></li><li><p>Employment</p></li><li><p>Food access (diet, convenience)</p></li><li><p>Walkability and transportation</p></li><li><p>Healthcare access (if relevant)</p></li><li><p>Weather (what you enjoy, tolerate or dislike)</p></li><li><p>Access to nature</p></li><li><p>Safety (as you define it)</p></li><li><p>Logistics (airports, shipping, family visits, errands)</p></li><li><p>Density / stimulation (do you want energy or calm?)</p></li></ul><p>Then make them <strong>concrete</strong>. &#8220;Access to nature&#8221; is not a criterion. It&#8217;s a label.</p><p>A criterion is something you can test, like:</p><ul><li><p>&#8220;Within a 15-minute walk: a park I&#8217;d actually use.&#8221;</p></li><li><p>&#8220;I can do 80% of my weekly errands without a car.&#8221;</p></li><li><p>&#8220;My grocery options support how I actually eat (not how I wish I ate).&#8221;</p></li><li><p>&#8220;Rent + utilities + predictable &#8216;extras&#8217; fit within my comfort range.&#8221;</p></li><li><p>&#8220;Noise and stress in the home environment stay below a threshold.&#8221;</p></li></ul><p>The point is not to create a perfect scoring system. The point is to stop making decisions based on vibes you can&#8217;t operationalize.</p><div><hr></div><p><strong>Step 3: Prioritize (must-haves vs nice-to-haves, reality vs aspiration)</strong></p><p>Now the real work: prioritization.</p><p>People frequently confuse a preference with a requirement.</p><p>A <strong>must-have</strong> is a constraint you can&#8217;t realistically violate without breaking the move:</p><ul><li><p>budget reality</p></li><li><p>commute requirements (if you have them)</p></li><li><p>core health needs</p></li><li><p>housing necessities (space, noise insulation, safety, etc.)</p></li><li><p>logistics you can&#8217;t wish away</p></li></ul><p>A <strong>nice-to-have</strong> is something you want, but could sacrifice if other wins are strong.</p><p>A second consideration matters just as much: <strong>Daily needs vs episodic needs.</strong></p><p>Daily needs dominate your happiness because they repeat. A great restaurant scene matters less if you eat at home 90% of the time. A world-class airport matters less if you fly twice a year. On the flip side, if you travel constantly, the airport becomes a daily-life consideration.</p><p>Finally: <strong>Current reality vs future aspiration.</strong></p><p>If you aren&#8217;t exercising now, don&#8217;t overweight &#8220;strong fitness culture&#8221; as the deciding factor. That doesn&#8217;t mean you can&#8217;t change&#8212;but it means you should be honest about the likelihood that a move alone will change you.</p><p>A useful test here:</p><ul><li><p>If this criterion disappeared, would I be mildly disappointed&#8212;or miserable?</p></li><li><p>If I got &#8220;good enough&#8221; here, would that free me to prioritize other things?</p></li></ul><div><hr></div><p><strong>Step 4: Apply the framework to your current city first</strong></p><p>As with any experiment &#8211; having a benchmark or starting point is best.</p><p>Run the framework on where you currently live:</p><ul><li><p>What works well?</p></li><li><p>What&#8217;s constantly annoying?</p></li><li><p>What are you compensating for with money, time, or stress?</p></li><li><p>Which friction points keep showing up?</p></li></ul><p>This is how you learn what you&#8217;re <em>actually trying to change</em>.</p><p>For me, this is also helped flesh out interdependencies. The living environment affects stress. Stress affects sleep. Sleep affects diet. Diet affects health. Health affects productivity. Productivity affects earning power and confidence. Suddenly the city choice isn&#8217;t just lifestyle&#8212;it&#8217;s downstream of everything.</p><div><hr></div><p><strong>Step 5: Research cities&#8212;but evaluate at the neighborhood level</strong></p><p>Here&#8217;s a simple truth that collapses a lot of noise:</p><p><strong>Most cities aren&#8217;t one experience. They&#8217;re a bundle of neighborhoods.</strong></p><p>Even in cities with strong overall characteristics, your day-to-day is determined by:</p><ul><li><p>where you live</p></li><li><p>where you work</p></li><li><p>what&#8217;s within walking distance</p></li><li><p>how you move around</p></li><li><p>what your immediate environment feels like</p></li></ul><p>Atlanta has walkable areas. It has areas where you can live without a car. It also has areas where that would be miserable. &#8220;Atlanta&#8221; as a single recommendation is meaningless without the neighborhood.</p><p>Same for Seattle. Same for NYC. Same for almost anywhere that isn&#8217;t tiny.</p><p>A practical research method that helped me:</p><ul><li><p>List cities you think you&#8217;d like</p></li><li><p>Also list cities you&#8217;re confident you wouldn&#8217;t</p></li><li><p>Use the contrast to refine your constraints (e.g., &#8220;I don&#8217;t tolerate extreme heat&#8221; or &#8220;I don&#8217;t want months of sub-freezing temperatures&#8221;)</p></li></ul><div><hr></div><p><strong>Step 6: Down-select to 3&#8211;10 options (then stop browsing)</strong></p><p>At some point, &#8220;more research&#8221; becomes procrastination.</p><p>Build a shortlist. For me, it included places like Boulder, Denver, Nashville, Seattle, Santa Clara, San Francisco, Oakland, Chicago, Tampa, Raleigh, Bellevue.</p><p>Your list will differ. That&#8217;s the whole point.</p><p>Use a simple rule:</p><ul><li><p>Must-haves = pass/fail</p></li><li><p>Nice-to-haves = 1-5 score</p></li><li><p>Keep notes on trade-offs, not just ratings</p></li></ul><div><hr></div><p><strong>Step 7: Experiment (visits prevent expensive mistakes)</strong></p><p>If there&#8217;s one cheat code here, it&#8217;s this:</p><p><strong>Visit. Simulate a normal week or at least a few days.</strong></p><p>Don&#8217;t just do tourist stuff. Do your routine:</p><ul><li><p>walk to a grocery store you&#8217;d actually use</p></li><li><p>try the transit route you&#8217;d actually take</p></li><li><p>go to a gym you&#8217;d actually join</p></li><li><p>walk in the neighborhood at the time you&#8217;d normally walk</p></li><li><p>sit in a caf&#233; and work for an hour if you work remotely</p></li><li><p>pay attention to friction: noise, safety vibe, stress level, convenience</p></li></ul><p>This is also where you learn things you can&#8217;t learn online.</p><p>I almost chose a building/neighborhood combination that looked perfect on paper&#8212;until I visited and noticed sketchy behavior from the building manager. The neighborhood and apartments were great. But the &#8220;day to day reality&#8221; felt off &#8211; ranging from spikes in traffic noise, hidden fees and a lot of &#8220;it depends&#8221; answers to my questions. That&#8217;s not a minute detail. That&#8217;s the kind of thing that turns a move into a chronic stress generator.</p><div><hr></div><p><strong>Step 8: Decide (don&#8217;t let perfect be the enemy of good)</strong></p><p>There will always be trade-offs.</p><p>Some people try to optimize to a mythical &#8220;perfect city&#8221; and end up stuck, endlessly comparing. Instead, aim for:</p><p><strong>Significantly better than your current state.</strong></p><p>Then commit.</p><p>You can refine over time&#8212;sometimes within the same region, sometimes within the same city, often just by switching neighborhoods.</p><p>This is where diminishing marginal returns matters. Getting to 8/10 on your must-haves usually beats sacrificing multiple criteria to chase 10/10 on one. &#8220;Good enough&#8221; is not settling; it&#8217;s calculated decision.</p><div><hr></div><p><strong>Lessons learned (what I wish I could tell my past self)</strong></p><p><strong>1) Test aspirations vs reality.</strong><br>If you don&#8217;t do it now, don&#8217;t overweight it. Validate needs through behavior.</p><p><strong>2) Social life is portable.</strong><br>&#8220;Seattle is less friendly&#8221; is too generic to be actionable. You can meet people anywhere and be lonely anywhere. Your habits and communities matter more than city stereotypes.</p><p><strong>3) Experimentation prevents expensive mistakes.</strong><br>Visits, neighborhood testing, and realism checks beat internet certainty.</p><p><strong>4) Hidden costs matter.</strong><br>Not just money&#8212;time and friction. Filters, locks, insurance, parking, delivery fees, and all the little things that add up. A rent number without lifestyle math is incomplete.</p><p><strong>5) Diminishing marginal returns are everywhere.</strong><br>You don&#8217;t need perfection. You need a strong baseline across key variables. Then you can make a decision regarding what to optimize and when/where to fit in nice to haves.</p><p><strong>6) &#8220;Access to nature&#8221; is not a single preference.</strong><br>A 5am hiker and a casual park walker are describing different lives. Same label, different requirement.</p><div><hr></div><p><strong>Closing Thoughts</strong></p><p>I&#8217;ve walked in the Seattle rain. I&#8217;ve seen gray days. And I&#8217;m still glad I moved&#8212;because the decision wasn&#8217;t based on a city&#8217;s reputation. It was based on whether the city fit my real day-to-day life.</p><p>If you&#8217;re considering a move, don&#8217;t start with &#8220;What&#8217;s the best city?&#8221;</p><p>Start with:</p><p>1. What does my life actually look like?</p><p>2. What would make it meaningfully better?</p><p>3. Which places can realistically support that?</p><p>Then build your shortlist&#8212;and go test it.</p>]]></content:encoded></item><item><title><![CDATA[Claude Code for Financial Planning – Top-Down vs. Bottom-Up Approach]]></title><description><![CDATA[Claude Code]]></description><link>https://thecontractsignal.com/p/claude-code-for-financial-planning</link><guid isPermaLink="false">https://thecontractsignal.com/p/claude-code-for-financial-planning</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Mon, 09 Feb 2026 04:02:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Late last year, I tried using Claude Code to assist with some financial planning. This was driven in part by a planned move to a new city and understanding the corresponding cost of living.</p><p>I explored a few different approaches in parallel &#8211; a top-down type of workstream and a bottom-up workstream.</p><p>In this context &#8211; a top-down approach means larger scope: effectively I compiled all of the financial statements and transactions from a given period of time, say the last 3-5 years, and then used Claude Code and other tools to actually parse out the transactions, compile them, categorize them, and analyze them or enable my own analysis of them.</p><p>I tested a few different approaches that uses this sort of top-down workflow.</p><p>The original overall approach or pipeline was:</p><p>- Create an overall plan or outline</p><p>- Compile / download financial statements (PDFs)</p><p>- Drop into a local folder and have claude code work with them end to end:</p><blockquote><ul><li><p>Provide initial context and refine plan using Claude Code&#8217;s plan mode</p></li><li><p>Convert each PDF into Excel or csv</p></li><li><p>Parse out the transactions</p></li><li><p>Aggregate the transactions into a single Excel or csv</p></li><li><p>Categorize the transactions</p></li><li><p>Perform follow-on analysis</p></li><li><p>Iterate</p></li></ul></blockquote><p>After a few hours it became clear that parsing out transactions from PDFs to Excel / csv is deceptively challenging (even for Claude), particularly when working with different data sources. So I replaced that specific step with using a third-party PDF to Excel conversion tool that I already had access to and was both much faster and more reliable.</p><p>Otherwise, the overall top-down approaches were the same.</p><p>My findings from the top-down approaches were as follows:</p><ul><li><p>First of all, handling large scope to start can be overwhelming in terms of the amount of the amount of necessary code (and of course corresponding context windows)</p></li><li><p>Second of all, because of the volume of code and data corresponding to the scope, the human testing costs become much higher</p><ul><li><p>In other words, the realm of possibilities of things that can go wrong are much broader, require more rigorous testing and the testing is more time consuming to perform</p></li><li><p>Unless an issue is glaring or there is some clear prioritization that can be done &#8211; for example identify what the largest transactions (such as rent payments), the transactions require a fair amount of individual review</p></li><li><p>However even when prioritizing these things there are all sorts of nuances that can get introduced &#8211; for example transactions being combined together, or amounts being off, or occasional exceptions that are difficult to spot (e.g. imagine a paycheck and a bonus being combined together or a rent payment being combined with a security deposit or a move from one apartment to another leading to different rent costs, etc.)</p></li><li><p>Oftentimes the context is situation-specific as well which makes more automated tests or handling this via the planning phase more difficult</p></li></ul></li><li><p>Third of all, in terms of the reliability and accuracy of the code &#8211; the top-down approach is decent at creating estimates but if exactness or precision is necessary, it is of limited utility as it would require a fair amount of manual quality control and testing. This may be okay as a starting point but usually requires complementing with other approaches to get more precise data</p></li></ul><h3>The bottom up approach</h3><p>Given the limitations of the top-down approach, the question came up: is there a complementary or perhaps alternative approach to the top-down process that focuses on a smaller scope?</p><p>For example, starting with a specific category of spend, mastering that and then repeating the process for other categories.</p><p>The business analogy I would use here is how Amazon focused on a specific product category to start (books). There are similar best practices when starting a business &#8211; specifically to focus on an underserved niche rather than trying to explore too large of a market and risk more competition, expensive customer / user acquisition, etc.</p><p>So the bottom up approach has the benefit of targeting specific categories one at a time (e.g. rent, grocery spending). This in turn involves smaller volumes, similarly structured transactions and less corresponding complexity so it involves less code, simpler testing, less cognitive load and greater transparency into both the data and the code.</p><p>This has been my experience so far.</p><p><strong>To summarize: the key benefit of the bottom-up approach is simplicity.</strong> Less categories, smaller transaction volume, less things that can go wrong, and less things to consider at any given time.</p><p>Now compare that with the top-down case where you have many statements across multiple categories. So you have a much larger volume of transactions of different types. You have different types of data structures for each. You have different types of ways in which things are expressed and net of those considerations there are many things that can go wrong potentially at each stage of the process, particularly when you&#8217;re parsing, aggregating, and analyzing the transactions.</p><h3>An Open Question &#8211; Building Up to Complex Use Cases?</h3><p>In addition, there is an additional potential benefit for more complex systems with a large volume of categories. My hypothesis is that if you build it up from the bottom step-by-step, category by category and each category is sufficiently constrained, over time you may be able to create an orchestration agent to actually coordinate each of these as if they were separate services.</p><p>For example you can have a grocery purchase analyzer as one service.</p><p>You can have a rent payment analyzer as another service, a medical appointment analyzer, a transportation payment analyzer, and so on.</p><p>You can do something similar for basically every category of items that is sufficiently standardized or constrained. Then, as you (and the system) develop a strong understanding of these you can then potentially automate the system entirely.</p><p>I need to pressure test the feasibility of this by running some experiments, but early indications are promising.</p><h2>Conclusion</h2><p>Overall, Claude Code has been very helpful in assisting me with financial planning and other use cases. So far &#8211; the top-down approach has been preferable to get a decent estimate of overall spend while the bottom-up approach has shown the most promise when exactness and detailed results are necessary.</p><p></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Shopify’s Fulfillment Experiment: Lessons on Focus and Knowing When to Quit]]></title><description><![CDATA[Balancing focus with expansion opportunities is one of the oldest strategic trade-offs in business.]]></description><link>https://thecontractsignal.com/p/shopifys-fulfillment-experiment-lessons</link><guid isPermaLink="false">https://thecontractsignal.com/p/shopifys-fulfillment-experiment-lessons</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sat, 18 Oct 2025 21:50:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Balancing focus with expansion opportunities is one of the oldest strategic trade-offs in business.</p><p>We can see this tension everywhere. Google must balance its dominance in search with new bets on AI and conversational interfaces like Gemini. Amazon went from selling books to becoming an &#8220;Everything Store&#8221;, then turned the internal logistics and computing system that powered that store into an entirely new external business &#8212; AWS &#8212; one of the most successful divisions in corporate history.</p><p>But expansion carries risks. Venture too far from your strengths, and you risk competing where you are weakest.</p><p>Capital, time, and attention &#8212; all scarce resources &#8212; get spread thin and diverted from the highest impact areas, leading to underperformance and, in the extreme, existential risk.</p><p>Yet focus alone carries its own danger. When customer preferences shift or technology evolves, companies that stay too narrowly focused can be blindsided by simpler, cheaper, or more relevant competitors. Disruption thrives where incumbents refuse to experiment.</p><h3>Shopify&#8217;s Fulfillment Bet</h3><p>Shopify&#8217;s short-lived fulfillment business captures this tension perfectly.</p><p>The business &#8211; known as the Shopify Fulfillment Network &#8211; launched in 2019.</p><p>For years, Shopify&#8217;s value proposition was clear: provide small and medium-sized merchants with software tools to compete online. The platform handled websites, payments, and inventory management, but left physical logistics to merchants themselves or third-party providers.</p><p>This created a gap. Small merchants lacked the scale to negotiate favorable rates with logistics providers or build their own warehouses. Meanwhile, Amazon&#8217;s fulfillment network gave its sellers a powerful competitive advantage. If Shopify could offer similar infrastructure, its merchants could compete more effectively.</p><p>The idea was bold: extend Shopify&#8217;s software platform into the physical world by building fulfillment and logistics infrastructure for small merchants. If successful, it would give small sellers the same capabilities as giants like Amazon &#8212; affordable, reliable, fast delivery &#8212; while letting Shopify capture a larger share of the merchant ecosystem. It would enable a more seamless, end-to-end experience for merchants and their customers.</p><p>On paper, the strategy made sense. Merchants already used Shopify to run their online storefronts. Why not also help them manage inventory and shipping? It was a logical extension of the &#8220;all-in-one commerce platform&#8221; narrative.</p><p>In practice, though, it pulled Shopify into a fundamentally different business.<br><br>Running fulfillment centers, managing warehouses, and coordinating last-mile delivery is operationally complex and capital-intensive &#8212; the opposite of Shopify&#8217;s asset-light software model. Suddenly, Shopify was competing in Amazon&#8217;s backyard, a domain where scale and logistics expertise matter far more than software elegance.</p><h3>Knowing When to Walk Away</h3><p>With hindsight, it&#8217;s easy to label Shopify&#8217;s fulfillment expansion a misstep.</p><p>Shopify sold the business in 2023, ultimately deciding to partner with existing providers instead.</p><p>But that decision shouldn&#8217;t be viewed as simple failure &#8212; it&#8217;s better understood as an <em>experiment that ran its course.</em></p><p>The willingness to run large experiments is a sign of an innovative culture. Shopify bet big on improving its merchants&#8217; experience, mirroring Amazon&#8217;s customer-obsessed ethos. It saw an unmet need and tried to fill it.</p><p>The more impressive part, though, is the willingness to walk away.<br>When the economics didn&#8217;t justify continued investment &#8212; and when it became clear that the effort was drawing focus away from Shopify&#8217;s software and payments businesses &#8212; leadership made the call to exit. That takes discipline. Many companies hold onto sunk-cost projects far too long.</p><h3>Changing Context, Changing Calculus</h3><p>Part of the decision was driven by evolving company and industry dynamics. As Shopify&#8217;s merchant base expanded to include larger enterprises, the value proposition for an in-house fulfillment network weakened; big merchants already had established logistics partners, reducing the need for Shopify to build its own.</p><p>The context had shifted &#8211; the same strategic move that once seemed visionary now risked becoming a distraction.</p><h3>Lessons Learned</h3><p>There are three broader lessons from Shopify&#8217;s entrance into and exit from the fulfillment space:</p><p>1. <strong>Experiment boldly, but deliberately.</strong><br>Growth often requires venturing beyond your comfort zone and running experiments with high-risk and high-reward. Shopify&#8217;s decision to make a large bet on a new business model was a testament to its experimental, high-agency builder culture.</p><p>2. <strong>Know when to quit.</strong><br>Success isn&#8217;t about avoiding mistakes &#8212; it&#8217;s about reallocating time and capital quickly when experiments don&#8217;t pan out. Strategic focus and adaptability are competitive advantages.</p><p>3. <strong>Adapt as the world changes.</strong><br>A good strategy at one moment can become a poor fit later. Shopify&#8217;s willingness to re-evaluate assumptions as context evolves is part of its disciplined execution and has helped sustain its momentum since.</p><h3>The Bigger Picture</h3><p>Whether at the level of a company, a team, or an individual career, the same logic applies. Experimentation is essential for learning and innovation &#8212; but so is focus. The goal isn&#8217;t to avoid failed experiments; it&#8217;s to make smart bets, learn fast, and redirect effort when the evidence points elsewhere.</p><p>Shopify&#8217;s fulfillment story isn&#8217;t just about logistics. It&#8217;s about the art of balancing curiosity with clarity: knowing when to build, when to explore, and when to let go.</p><p>From 2023 to 2024, Shopify&#8217;s revenues have increased 26% to over $8.8 billion while net income climbed to over $2 billion.</p>]]></content:encoded></item><item><title><![CDATA[Pricing Strategy - Lessons from Shopify’s 2024 10-K (Pt 2)]]></title><description><![CDATA[Shopify&#8217;s business model is a master class in pricing strategy.]]></description><link>https://thecontractsignal.com/p/pricing-strategy-lessons-from-shopifys</link><guid isPermaLink="false">https://thecontractsignal.com/p/pricing-strategy-lessons-from-shopifys</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sat, 11 Oct 2025 20:31:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Shopify&#8217;s business model is a master class in pricing strategy.</p><p>They have two main revenue components:</p><ul><li><p><strong>Subscription Solutions (~26%)</strong> &#8211; recurring revenue from monthly subscriptions to access the Shopify platform, plus related activities like app sales and domain registrations.</p></li><li><p><strong>Merchant Solutions (~74%)</strong> &#8211; transaction-based revenue from payment processing, currency conversion, and other services linked to merchant success.</p></li></ul><div><hr></div><h2>Subscription Solutions</h2><p>Subscription solutions are mostly comprised of platform access, and are priced accordingly (mix of fixed costs and variable platform fees). There are different tiers of platform access with different corresponding pricing. The reasoning being that when entrepreneurs are starting out, simplified / entry-level platform access makes sense but as a business grows and has more complex needs it will likely need access to a broader range of features to navigate the complexity involved (for example managing multiple sales/marketing channels or managing more customer accounts). Shopify Plus is mentioned as their premium tier with significantly higher pricing than the more basic out of box offerings.</p><p>It&#8217;s noteworthy that for any given reporting period, no single merchant has accounted for more than 5% of their revenue (which speaks to how diversified the business is).</p><p>Also, there are some other recurring revenue sources such as, sale of apps, registration of domain names and sale of themes (online store/website templates). These do not appear to be significant contributors to revenue but help provide a one-stop-shop in some respects, enabling customers (merchants) to get their needs met without having to explore other alternatives, which provides convenience and an overall better user experience.</p><div><hr></div><h2>Merchant Solutions</h2><p>Merchant solutions are priced on a transaction basis &#8211; as more transaction volume comes through the Shopify platform, Shopify collects payment processing and currency conversion fees, referral fees from partners, Shopify Capital (lending to / financing of qualified merchants), and Transaction Fees.</p><p>In addition, they generate revenue from services and products like sale of shipping labels, sale of point-of-sale (&#8220;POS&#8221;) hardware, advertising on the Shopify App Store, and Shop Campaigns (buyer acquisition offering).</p><p>While individually small, these involve similar dynamics as the subscription solutions segment of their business. Specifically, these offerings reduce friction and keep merchants operating within Shopify&#8217;s ecosystem &#8211; strengthening the user experience and retention.</p><div><hr></div><h2>Incentives and the Flywheel Effect</h2><p>From an incentives standpoint &#8211; when merchants succeed, Shopify succeeds.</p><p>This happens in a few different ways:</p><ul><li><p><strong>Subscription solutions</strong></p><ul><li><p>When a business is starting out &#8211; they don&#8217;t need access to premium platform features</p></li><li><p>Free trials (particularly during times of market slowdown) help incentivize entrepreneurs and up and comers to start using the platform. If they start gaining traction as a business and view the platform as helpful in doing so, they&#8217;ll start paying</p></li><li><p>If the business succeeds over time, it&#8217;s likely growing and will be a prime target to upsell/upgrade to the next tier</p></li><li><p>In addition, while most businesses pay month to month, some elect to pay annually or multi-year subscriptions (with longer terms likely incentivized by discounts, akin to volume-based pricing)</p></li></ul></li><li><p><strong>Merchant solutions</strong></p><ul><li><p>The more business a customer (merchant) does through the platform, the more Shopify collects in transaction-based fees. This is even better aligned in terms of incentives as Shopify immediately benefits while the merchant collects most of the profit associated with the sale, creating a win-win situation</p></li><li><p>Referral fees from partner fees are more likely to be generated as Shopify captures more of the market in its core offerings and develops increased brand recognition, and allow all parties in the ecosystem - including customers (merchants), partners and Shopify itself - to benefit accordingly from the network effect</p></li><li><p>Through Shopify Capital, lending to qualified merchants with strong business prospects generates revenue in the short term (loan interest) while functioning as a form of customer acquisition for the core offerings. As these businesses grow, they are more likely to upgrade to premium subscription tiers and process higher transaction volumes.</p></li><li><p>The other offerings &#8211; while not significant revenue drivers help customers (merchants) focus on their core business by reducing search/switching costs as well as stay in the platform (possibly reducing risk of customer churn over time)</p></li></ul></li></ul><div><hr></div><h2>Conclusion</h2><p>The genius of Shopify&#8217;s model lies in its dynamic, stage-appropriate pricing: it meets merchants where they are and grows with them. Free trials convert to paid plans, paid plans upgrade to premium tiers, and transaction fees scale with success&#8212;all while ancillary services reduce switching costs. It&#8217;s a masterclass in aligning company incentives with customer outcomes, and the 28% YoY revenue growth in 2024 provides strong evidence that it works.</p><p>Next time &#8211; I&#8217;ll explore Shopify&#8217;s profitability and cost structure.</p>]]></content:encoded></item><item><title><![CDATA[Some Things I Learned Reading Shopify’s Latest Annual and Quarterly Filings]]></title><description><![CDATA[Pt 1]]></description><link>https://thecontractsignal.com/p/some-things-i-learned-reading-shopifys</link><guid isPermaLink="false">https://thecontractsignal.com/p/some-things-i-learned-reading-shopifys</guid><dc:creator><![CDATA[Leonid Prilutskiy]]></dc:creator><pubDate>Sat, 04 Oct 2025 20:35:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aR1x!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd35163fd-b60f-4c97-94ed-f30c0d7bf908_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Shopify is one of the world&#8217;s leading e-commerce infrastructure platforms. It helps merchants of all sizes start, grow, market, and run their businesses.</p><p>Here are some things I learned from reading Shopify&#8217;s Recent Quarterly and Annual Filings:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Leonid&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>1. <strong>Two Main Revenue Streams</strong></p><p>Shopify&#8217;s revenue is divided into two broad categories:</p><ul><li><p><strong>Subscription Solutions (~26%)</strong> &#8211; recurring revenue from monthly subscriptions to access the Shopify platform, plus related activities like app sales and domain registrations.</p></li><li><p><strong>Merchant Solutions (~74%)</strong> &#8211; transaction-based revenue from payment processing, currency conversion, and other services linked to merchant success.</p></li></ul><p>Together, these two segments reflect Shopify&#8217;s hybrid model: predictable recurring revenue plus upside from merchants&#8217; transaction volume.</p><div><hr></div><p><strong>2. Customers</strong></p><p>Shopify&#8217;s customers are <strong>merchants</strong>&#8212;ranging from solo entrepreneurs to large enterprises. </p><div><hr></div><p>3. <strong>How Shopify Measures Performance</strong></p><p>Shopify highlights two key <strong>operating metrics</strong> in its filings:</p><ul><li><p><strong>Monthly Recurring Revenue (MRR):</strong> the following month&#8217;s expected subscription revenue based on currently active subscriptions</p><ul><li><p>up from <strong>$144 million in 2023</strong> to <strong>$178 million in 2024</strong>, a <strong>~24% increase</strong>.</p></li></ul></li><li><p><strong>Gross Merchandise Volume (GMV):</strong> this is a &#8220;merchant success&#8221; metric corresponding to the total merchant transactions processed over a period of time, including shipping + taxes, net of all cancellations / refunds</p><ul><li><p>up from <strong>$235.9 billion in 2023</strong> to <strong>$292.3 billion in 2024</strong>, a <strong>~24% increase</strong>.</p></li></ul></li></ul><p>It&#8217;s interesting to note that the growth rates are the same despite corresponding to separate parts of the business. This could be a coincidence but is worth noting when reviewing the results from different periods.</p><div><hr></div><p>4. <strong>Growth and Pricing Strategy</strong></p><ul><li><p>In 2022, Shopify introduced <strong>Paid Trial Incentives</strong>&#8212;low-cost access periods designed to attract new merchants after a post-COVID slowdown.</p></li><li><p>Many of those merchants upgraded in 2023, boosting growth, followed by a more normalized growth rate in 2024.</p></li><li><p>This illustrates how Shopify invests in the long term - providing increased value up front at a time when customers are more price-sensitive and then upselling when the time is right </p></li></ul><div><hr></div><p>5. <strong>Partnerships and the Payment Stack</strong></p><ul><li><p>Shopify relies on third-party partners such as <strong>PayPal</strong> and <strong>Stripe</strong> to process some payments.</p></li><li><p>However, most transactions run through <strong>Shopify Payments</strong>, the company&#8217;s own payment processing system.</p></li></ul><div><hr></div><p>6. <strong>Deferred Revenue Explained</strong></p><ul><li><p>Because merchants often pay upfront for monthly or annual access, Shopify records those payments as <strong>deferred revenue (a liability)</strong> and recognizes them gradually over the service / delivery period.</p></li><li><p>It&#8217;s a classic SaaS accounting principle that provides visibility into future recognized revenue.</p></li></ul><div><hr></div><p>7. <strong>Other Operational Notes</strong></p><ul><li><p><strong>App Store:</strong> Shopify operates its own App Store, where developers can build and sell plug-ins that extend store functionality.</p></li><li><p><strong>Sale of Logistics Business (2023):</strong> Simplified operations and improved margins, but makes year-over-year comparisons more complex.</p></li><li><p><strong>Seasonality:</strong> Q4 (holiday season) is historically the busiest quarter for Shopify merchants, boosting both GMV and Merchant Solutions revenue.</p></li></ul><div><hr></div><p>8. <strong>Open Questions Worth Exploring</strong></p><ol><li><p><strong>How large is the Shopify App Store</strong> in terms of merchant adoption or developer revenue?</p></li><li><p><strong>What impact could tariffs or trade policy</strong> shifts have on cross-border merchant sales?</p></li><li><p><strong>How does Shopify Payments add value</strong> on top of partnerships with PayPal and Stripe?</p></li><li><p><strong>Is Shopify Payments lower-margin per transaction</strong> but higher in operating margin as part of the overall business?</p></li><li><p><strong>Why did Shopify switch from filing Form 40-F to Form 10-K</strong>? Does this mean it transferred its legal presence or main operating presence to the United States?</p></li><li><p><strong>How does the logistics sale affect year-over-year comparability</strong> in Merchant Solutions revenue?</p></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://thecontractsignal.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Leonid&#8217;s Substack! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>