Engineering does not have an AI problem. It has an operations problem.
The instinct, when a technologist looks at an engineering firm, is to look at the calculations. That is the visible, technical, impressive part — and it is the wrong target. It is the part of the business with the least slack in it, the highest liability, and the smallest share of the payroll.
The rest of the firm is where the money leaks. Proposals. Estimating. Contract review. Scheduling. Submittal logs. RFIs. Meeting minutes. Specification cross-checks. Document control. Invoicing. Collections. Closeout. None of this requires a professional seal. All of it consumes licensed professionals' hours. And it is exactly the work that a coordinated set of AI agents is now good at.
The industry-level evidence that something is structurally wrong here is not subtle, and it is not new. In McKinsey's words: "Construction productivity improved by only 10 percent, or 0.4 percent annually, from 2000 to 2022, compared with 90 percent, or 3.0 percent annually, in manufacturing."2 In the United States the picture is worse than stagnation: Goolsbee and Syverson found that value added per construction worker was about 40% lower in 2020 than in 1970, a decline large enough to measurably drag on aggregate US productivity growth.3
Two decades of CAD, BIM, cloud document management, and project-management SaaS were sold into this industry over exactly the window in the chart above. The line did not move. That is the single most important fact in this paper, and I will come back to it twice — once as the argument for why the opportunity is real, and once as the strongest objection to Oxvelo entering at all.
Two out of five paid hours are already not billable.
Zoom from the sector to the firm. Deltek's Clarity study is the industry's longest-running financial benchmark of architecture and engineering firms; the 47th edition surveyed roughly 900 of them.4 Its numbers describe a business model under compression.
Read those together and the arithmetic is stark. Roughly 41% of paid professional hours are already non-billable before a single overhead-department salary is counted — and then every remaining billable dollar of labor drags $1.61 of overhead behind it. Meanwhile operating profit fell nearly five points in a single year off a ten-year high.
A number this paper deliberately does not use
The most-quoted statistic in AEC-technology marketing is that "construction professionals spend 35% of their time — over 14 hours a week — on non-productive activities," at a cost of $177 billion in US labor annually.18 It is real and well-sourced, but the surveyed population was 49% general contractors and 36% specialty trades — not design professionals. Applying it to engineers and architects would be the kind of small dishonesty that makes a reader stop trusting everything else in the document. The 58.9% utilization figure above is the defensible one.
The firm, not the task, is the unit of transformation.
Engineering companies are staffed almost entirely by engineers, and engineers think like engineers. That is a strength in the technical work and a blind spot everywhere else. It produces firms that will spend two years optimizing a structural analysis workflow and zero years asking why it takes eleven days to turn a signed contract into a project workspace.
The alternative framing is to organize the opportunity by business function rather than by engineering discipline. Every engineering business contains business development, estimating, contracts, project management, technical production, document control, construction administration, finance, people operations, compliance, and knowledge management. Each of these has an AI counterpart. Together they are not a feature set — they are a workforce.
Underneath the workforce sit two more layers, and they are where the durable advantage actually lives:
Institutional intelligence. Every project, document, standard, decision, lesson learned, estimate, and historical outcome becomes searchable and reusable. Today this asset exists in every firm and is worth almost nothing, because it is scattered across file servers, inboxes, and the memory of a principal who retires in four years. In 2022 an estimated 184,175 engineers retired or left the field, 85,175 of them from civil, mechanical, and electrical disciplines.15 Every one of them took an unindexed corpus with them.
Workflow orchestration. Once a project is won, the system initiates project setup, schedules, deliverable registers, budget tracking, compliance logs, document structures, and draft work products — escalating to humans at defined control points rather than waiting to be asked.
Start where no one has to seal anything.
Some engineering work is now so repeatable that the intelligence has already been embedded — in standards, templates, prior designs, and established procedures. What remains is execution, checking, documentation, and sign-off. That is a very different problem from "design a bridge," and it is the correct place to begin.
The right way to sequence entry is to plot every function on two axes at once: how automatable it is, and how much professional liability it carries. Those two axes are not independent — and the safe, valuable territory is the top-left, not the top-right.
Note what this rules out. It rules out an AI that independently approves submittals. It rules out an AI in responsible charge of a design. It rules out selling "the model sealed it." What it rules in is a firm where the proposal was assembled in an afternoon instead of two weeks, the contract deviations were flagged against a company standard before a principal ever opened the file, the deliverable register built itself at kickoff, and the RFI log answered 60% of its own questions by citing the spec section — all of it auditable, all of it escalating to a named licensed human at the points where the law requires one.
We no longer need engineers to produce the design. We need them to QA it.
This is the sharpest form of the thesis, and it deserves to be stated without hedging: a design team of a hundred engineers is an artifact of a production constraint that no longer binds. Two engineers, reviewing and taking responsibility for AI-produced output, can carry the production load that a large team once required. The financial, throughput, accuracy, and cycle-time consequences of that are not incremental. They are a different business.
The reason this is not reckless is that the profession already has a legal category for exactly this arrangement, and has had one for decades. It is called review.
Responsible charge never required authorship
Read the rules again with this claim in mind — and read the current ones, because NCEES rewrote the key definition in August 2025 in a way that has gone almost entirely unremarked. The old Model Law defined responsible charge as "direct control and personal supervision of engineering or surveying work." The August 2025 Model Law replaced that with a functional test: responsible charge now means "to exercise full professional knowledge of and control over work," elaborated as four duties — the authority to review, change, reject or approve the work; personal awareness of its scope; the capability to answer questions about it; and acceptance of full responsibility for it.20
That change matters more than any commentary I have seen acknowledges. The old test asked whether you supervised the person doing the work. The new one asks whether you know, control, can defend, and own the work itself. A licensed engineer reviewing AI output can satisfy every one of those four duties. A licensed engineer cannot "personally supervise" a model in any meaningful sense — under the superseded language, the argument was awkward. Under the current language, it is clean.
The rest of the framework points the same way. NCEES Model Rules §240.20 D requires that the seal be placed on work "only when it was under the licensee's responsible charge."20b New York's rule defines unprofessional conduct as sealing documents for which the services "have not been performed by, or thoroughly reviewed by, the licensee" — the disjunction is in the regulation.59 NSPE's Code expressly contemplates an engineer accepting responsibility for coordination of an entire project and sealing its documents, provided each technical segment is signed and sealed by the qualified engineer who prepared it.21b
Every large engineering firm on earth already runs on this. A principal seals work produced by junior staff they did not personally draft. The senior engineer's value was never the drafting — it was the judgment applied to the drafting. AI does not change the legal structure of that relationship at all. It changes the marginal cost of the thing being reviewed from an engineer's salary to a compute bill.
NSPE's Board of Ethical Review confirmed the boundary precisely in Case 24-2. An engineer used AI to draft a groundwater report — which he verified thoroughly, cross-checking against journal articles — and to draft design documents, which he reviewed only cursorily; the client found misaligned dimensions and omitted safety features. The Board's finding was not that using AI was wrong: "The use of AI-assisted drafting tools by Engineer A was not unethical per se. However, Engineer A's misuse of the tool, by failing to maintain Responsible Charge over the AI tool and its output before sealing the document… was unethical."21 The violation was insufficient review, not delegation. That is a green light with a condition attached, and the condition is the entire product.
Two details of that case are worth carrying forward rather than glossing. First, the Board judged the report use "partly ethical, and partly unethical" — thorough verification satisfied the competence obligation, but the engineer separately failed to obtain client permission before disclosing private information to the tool and failed to document required technical citations. Verification alone was not sufficient. Second, the Board held there was no obligation to disclose AI use to the client, while noting that "ethical principles favor transparency when AI plays a substantial role."21 Data handling and citation discipline are therefore product requirements alongside review — which is exactly what the provenance layer in Figure 6 exists to satisfy.
The arithmetic of leverage
The economics follow mechanically. US architecture firms generate average net billings of $143,000 per employee — up from $86,000 in 20117 — against a total labor cost of $119,511 per employee at A&E firms.45 That is the entire industry's problem in two numbers: revenue per head and cost per head are almost the same number, and the 161.3% overhead rate sits between them. (The $143,000 is a mean; AIA does not publish a median, and in a distribution this skewed the median is likely lower — which makes the gap worse, not better.)
Now introduce a leverage ratio — call it L, the number of engineer-equivalents of production output that one licensed professional can review and take responsibility for. Today, structurally, L ≈ 1. Every engineer produces roughly their own work. Move L and the revenue-per-professional line moves with it, linearly, while headcount cost does not.
The "two engineers instead of a hundred" claim sits at the far right of that curve — L ≈ 50 for the production function specifically. Two clarifications make it defensible rather than merely bold:
- It applies to production, not to the whole firm. A hundred-person engineering firm is not a hundred people drafting. It contains project managers, construction administration staff, principals, finance, and business development. The claim is that the design production function collapses to a small review team — while the AI workforce in Figure 3 simultaneously compresses the other functions on their own curves. Those are two different arguments and both are in this paper.
- It applies first to templatized, repeatable work. The blue band in Figure 4 — where the intelligence is already embedded in standards, prior designs, and established procedures, and what remains is execution, checking, and documentation. On a first-of-its-kind structure, L stays near 1 for a long time, and it should.
So why is Oxvelo hiring engineers at all?
The obvious objection to everything above is that it appears to contradict the plan. If engineers are not needed to produce the design, why assemble a team of engineers?
Because the ceiling on what an AI system produces is set by the quality of the subject-matter expertise embedded in it — and that expertise has to come from somewhere. A general-purpose model dropped into an engineering firm produces generic output. The same model, working against a curated corpus of that firm's standards, prior designs, acceptance criteria, checking procedures, failure history, and hard-won conventions, produces the firm's work. The difference between those two systems is not the model. It is the engineers who built the second one.
This is exactly what the capability research implies, read forward instead of backward. AECBench found LLMs weakest at interpreting building-code tables and at complex reasoning38 — which means the fix is retrieval against curated, SME-validated sources and deterministic checks with SME-defined acceptance criteria. Curating those sources and defining those criteria is senior engineering work. Hallucination is a structural product of training objectives that reward guessing;39 the counter is a provenance layer whose ground truth an SME had to establish. Every layer in Figure 6 that makes review cheap is itself a deposit of expert judgment made once and spent many times.
Eudia's leadership said the same thing when they bought a 300-person legal services firm rather than hiring engineers to build around a model: "building truly transformative systems requires deep human expertise embedded within the technology itself."42 cove.tool's AI architecture practice was built on a claimed decade of R&D by practicing architects, not by an AI team that read a code book.45
So the headcount does not go to zero. It changes shape. The engineer's job splits into three, and every one of them is a higher-value job than drafting:
Notice what this does to the hiring profile. It does not call for a large team of production drafters. It calls for a small number of unusually senior people — the ones with the judgment worth encoding. And that is precisely the resource the market is shortest of. ACEC's 2025 workforce survey found that only 34% of engineering students feel well-prepared with the skills most important to succeed in their chosen profession, while mid-career and older professionals estimate that just 14% of their young colleagues are well-prepared on those same skills.15 Set that against 184,175 engineers leaving the workforce in a single year.15 The scarce asset in this industry is not production capacity. It is senior judgment — and senior judgment is the one input AI cannot supply and the one an AI-native firm can leverage furthest.
Two caveats on those figures, since they carry weight here. The 34% comes from a 195-student convenience sample recruited via three universities and ACEC scholarship applicants, not a representative national panel. And the 14% is the estimate of mid-career and senior practitioners — it is a perception of preparedness, not a measurement of it. Both point the same direction; neither is precise.
The magnitude claims, and what actually supports them
The strongest available quantitative anchor is McKinsey's statement that "AI has the potential to automate 50 percent of nonphysical work in the architecture and engineering sectors."2 Note what that means in leverage terms: it is L ≈ 2 across the whole enterprise — a doubling, not a fiftyfold. The high-L claim is a claim about a narrow, deliberately chosen slice, and it should always be stated that way. (A caution on a figure you will see quoted alongside it: MGI's $228 billion is the projected annual US value for the entire AEC industry by 2030 from "AI and other automation technologies" — it is not the dollar value of the 50%, and it is not AI alone.2)
The most direct operating evidence comes from cove.tool, which launched Cove Architecture in March 2025 as a full-service AI-powered architecture practice after a claimed $25 million and a decade of R&D. Its reported results on a 15-unit Atlanta housing project:
Source: cove.tool, reported by Bisnow and Multi-Housing News (2025).45 Handle with care: these are company-reported figures on a single project from a firm roughly eighteen months old, and neither account clarifies whether Cove Architecture is a registered licensed firm or how it was staffed. They are the best available signal that the magnitude is real, and they are not proof.
The adjacent professions read the same way, and Crosby is the cleanest instance. It reviewed more than 1,000 commercial contracts within roughly five months of launch with about 19 people; by its March 2026 Series B the figure was reportedly around 13,000 contracts with roughly 30 lawyers.40 That is the leverage curve moving in real time, in a licensed profession, with the firm explicitly stating: "We stand behind our work and take liability for it."40 The UK regulator authorized Garfield.Law on the condition that named regulated solicitors remain ultimately accountable for all system outputs.41 That is the QA model, licensed and operating, in a profession with the same liability structure as engineering.
The hard part: making review scale
Here is the honest problem with the claim, and it is a real one. If two engineers must genuinely review the output of a system producing at the rate of a hundred, review throughput becomes the binding constraint and rubber-stamping becomes the failure mode. A reviewer who cannot actually review is not in responsible charge, whatever the org chart says — and in New York, that reviewer owes a retained written evaluation for six years on work prepared by others.59
So the leverage ratio is not won by making the AI better at producing. It is won by making the output cheaper to verify. That is a different engineering problem, it is the actual product, and it is what Figure 6 describes.
Stated as a design principle: an engineer should never review a drawing. An engineer should review a small, ranked set of exceptions, each one carrying its own provenance, against a package where everything unexceptional has already been proven by deterministic check. Get that right and L climbs. Get it wrong and you have built a machine that manufactures liability at scale.
A fragmented industry, an adoption gap, and a consolidation window.
Three market conditions have to hold simultaneously for this to be a business rather than a thesis. All three currently hold.
The industry is extraordinarily fragmented
Start with the concentration figure, because it is the cleanest. ENR's 2026 Top 500 Design Firms — the five hundred largest design practices in the country — booked $136.3 billion in US domestic design revenue in fiscal 2025.12 Census's Quarterly Services Survey puts total 2025 revenue for NAICS 5413, architectural, engineering and related services, at roughly $494.5 billion.12b The five hundred biggest firms in America therefore account for something like 28% of the industry. The other 72% is spread across tens of thousands of firms — and within the Top 500 itself, the top ten alone take 32.2% of the revenue.12
Look inside the long tail and it is smaller than most people assume. AIA's 2024 Firm Survey finds that 28% of US architecture firms are sole practitioners, 32% have two to four employees, and about 75% have fewer than ten.7
Fragmentation matters for a specific reason: a firm with eleven people cannot build an AI operating layer, and cannot buy one that fits, but it can be joined to one. And the concentration trend says the small firms are already losing ground — the operating-model gap is the mechanism. And these firms are already changing hands. 2025 was the first year on record with more than 500 completed domestic A/E transactions; 72% of those deals involved firms under $10 million in revenue, and private-equity-backed acquirers or recapitalizations now account for more than half of all design and environmental consulting deals.13 A second tracker, using a narrower sector definition, counts 361 transactions in 2025 with PE at 38.3% of deal volume — up from 22.3% in 2018.14 The definitions differ; the direction does not.
The talent supply is structurally short
ACEC's workforce research puts the 2022 net engineering shortage at roughly 18,000 engineers, with 8,411 of that in the core civil, mechanical, and electrical disciplines.15 Engineering degrees awarded peaked near 214,000 in 2019 and have fallen by more than 10,000 since.15 As of Q2 2026, 33% of engineering firms had turned down work in the prior six months because they were short-staffed — down from 51% in Q4 2024, but still a third of the industry declining revenue it cannot staff.16 Median backlog stands at 11 months, with 49% of firms reporting a year or more.
This is the demand-side case, and it is stronger than the cost-side case. A firm that turns down work is not looking to cut headcount — it is looking for throughput. That reframes the entire sale.
Adoption is real but shallow — and the published numbers are a mess
Here I have to be careful, because this is where AEC-technology writing is least honest. Published AI adoption rates for this industry range from 6% to 98% depending entirely on what was asked and of whom.
The most defensible signal is not any single figure but the one series that asked the same question of the same industry three years running. Deltek's A&E study went from 38% to 53% to 70%.46
That gap — 70% of firms touching AI but only 38% reporting measurable business impact,4 while 52% of AEC respondents still use paper during the design phase and 43% still depend on physical signatures9 — is the whole opportunity. Adoption is a mile wide and an inch deep. Nobody has rebuilt the enterprise.
Seven objections, at full strength.
A position paper that only argues one side is marketing. Below are the seven strongest arguments against Oxvelo doing this — stated the way a hostile, well-informed reader would state them, then answered. Three of them are substantially correct, and they change the strategy rather than defeat it.
An AI cannot be in responsible charge of engineering work, cannot seal a drawing, and cannot hold professional accountability. Everything valuable in an engineering firm terminates in a licensed human's signature. You are automating the parts that don't matter.
The first half is simply true, and Oxvelo should say so louder than its critics do. NCEES Model Rules §240.20 D require that a licensee's seal and signature be placed on work "only when it was under the licensee's responsible charge,"20b with responsible charge defined since August 2025 as exercising "full professional knowledge of and control over work."20 New York goes further: sealing documents the licensee did not personally prepare or "thoroughly review" is defined as unprofessional conduct, and a licensee reviewing another's work must retain a written evaluation for six years.59
The second half — "the parts that don't matter" — is where the objection fails, and Figure 2 is the answer. If 41% of paid hours are already non-billable and every billable labor dollar carries $1.61 of overhead, then the unsealed work is the economics of the firm. The seal is the product's warranty. It is not the product's cost structure.
The regulators, notably, have already answered the narrow question. The Texas Board of Professional Engineers and Land Surveyors issued Policy Advisory Opinion PAO-71 in November 2024 finding that AI software is not prohibited, that "licensees are ultimately responsible for any work product they sign and seal," and that existing rules already govern AI adequately — declining to write AI-specific rules.22 NSPE's Board of Ethical Review reached the parallel conclusion in Case 24-2: using an AI drafting tool "was not unethical per se"; failing to maintain responsible charge over its output before sealing was.21 Florida's board put it in one sentence: "AI can assist your work, but it cannot replace your professional judgment or accountability."23
You are a non-licensee in New York. Under New York law you cannot own an engineering firm, cannot be its President, cannot be its Chairman, and cannot be its CEO. The entity you are proposing to build is illegal in your own state.
This objection is correct, and it is the single most consequential finding in this paper. It does not defeat the strategy. It dictates the corporate structure.
New York's baseline rule for an ordinary professional service corporation practising engineering is that every shareholder, officer, and director must be licensed in New York, and "no other entity or individual except those described in the preceding may practice professional engineering in New York State."25 The one relief valve is the Design Professional Service Corporation under BCL §1503(b-1), and its conditions are tighter than the summary usually given:24
- Greater than 75% of outstanding shares must be owned by design professionals and an ESOP;
- Greater than 75% of directors and greater than 75% of officers must be design professionals;
- The president, the chairperson of the board, and the chief executive officer or officers must all be design professionals;
- The largest shareholder must be a design professional or a qualifying ESOP;
- And critically — the sub-25% tranche is available to employee stock ownership plans and to employees of the corporation who are not design professionals. An outside, non-employee investor entity does not qualify for it at all.
So the position is worse than "a non-licensee may hold up to 25%." In New York, an outside holding company has no route to equity in the licensed design entity. Other states are far more permissive: California requires that a licensed engineer be an owner, partner, or officer in charge and bars sole ownership by a non-licensee,27 and Texas's firm-registration rule is silent on ownership entirely, requiring only that at least one full-time active licensee be employed and supervise the licensed work.26
Three separate regulated professions have independently converged on the same solution to this exact problem, which is strong evidence it is the right one.
In medicine, the corporate practice of medicine doctrine bars companies from employing physicians to deliver care. The universal workaround is the MSO/PC structure: a physician-owned professional corporation employs the clinicians and controls all clinical decisions; a management services organization — the technology company — provides platform, scheduling, billing, HR, and marketing under a management services agreement.55 In law, Crosby — which raised a $60M Series B led by Lux and Index in March 2026, with Sequoia, Elad Gil, 01 Advisors and Bain Capital Ventures participating — runs precisely this shape: "Crosby Legal PLLC is a law firm……powered by tech from Crosby Legal, Inc."40 In accounting, PE platforms like Alpine's Ascend split each acquired firm into an attest entity and a tax/advisory entity — the "alternative practice structure" — because CPA-majority ownership is required for attest work.44
Engineering's constraints map onto healthcare's almost one-to-one: state-by-state licensed entities, a named individual who must hold professional authority (the engineer in responsible charge, the physician), and ownership restrictions that vary from strict to permissive by state.
Language models are demonstrably bad at exactly the reasoning engineering requires. They cannot reliably read a code table, they fail at multi-step calculation, and hallucination is a structural property of how they are trained — not a bug awaiting a patch.
Accurate, and worth quantifying rather than hand-waving. AECBench — a peer-reviewed benchmark of nine LLMs against 4,800 expert-validated questions across a five-level cognitive framework in architecture, engineering, and construction — found that models handle foundational knowledge adequately but show significant deficits in interpreting knowledge from tables in building codes, in complex reasoning and calculation, and in generating domain-specific documents, with performance declining steadily as cognitive demand rises.38 And hallucination is not incidental: OpenAI and Georgia Tech researchers argue models hallucinate "because the training and evaluation procedures reward guessing over acknowledging uncertainty" — a statistical pressure, not a defect.39
But the AECBench paper contains a finding its own summary buries, and it is the single most useful result in this entire literature for Oxvelo's purposes. The building-code-table failure is an encoding problem, not a knowledge problem. When the researchers converted code tables into natural-language descriptions, model accuracy on those questions rose from 45.27% to 98.94%. Rendering the same tables as HTML lifted accuracy from 40.65% to 87.27%.38
Read that again. The same model, on the same code provisions, goes from failing more than half the time to answering almost perfectly — purely by changing how the source material is presented to it. That is not a capability ceiling. It is a data-engineering task, and it is precisely the kind of work an SME-built corpus does. It is also the strongest available evidence that the constraint in this domain is the preparation of knowledge rather than the reasoning over it.
Notice too what AECBench measured: unaided models answering engineering questions from parametric memory. That is not the proposed architecture. The proposed architecture is retrieval against the firm's own standards and prior work — restructured for machine legibility, as the 45%→99% result demands — with deterministic calculation in code rather than in the model, structured templates carrying prescribed checks, and mandatory escalation to a licensed reviewer. A system that looks up the code table and shows the citation is a different failure surface from a system asked to recall it.
"Two engineers instead of a hundred" does not eliminate the bottleneck — it relocates it. If AI produces at a hundred engineers' rate, two humans must read at a hundred engineers' rate to be genuinely in responsible charge. They cannot. So either the leverage is fictional, or the review is, and if it is the review then you have built a plan-stamping operation with better branding.
This is the correct objection to the boldest claim in the paper, and the answer is not to soften the claim — it is to say precisely what has to be true for it to hold.
The objection assumes review effort scales linearly with output volume. It does not have to. Review effort scales with the number of things a human must actually adjudicate, and that number is a design variable. A package where every dimension has been checked deterministically against the model, every code citation carries a retrievable provenance link to the clause it came from, and every deviation from firm standard is surfaced as a ranked exception is a fundamentally cheaper object to review than a stack of drawings — even though it represents the same volume of output. That is the entire content of Figure 6, and it is why the engineering effort goes into verification tooling rather than into generation quality.
The regulatory frame supports this reading, and after August 2025 it supports it more directly than before. NCEES's current four duties are the authority to review, change, reject or approve the work; personal awareness of its scope; the capability to answer questions about it; and acceptance of full responsibility.20 None of the four is "read every line." A principal sealing a junior's work has never re-run every calculation; they have applied judgment to a reviewable artifact. The question a board would ask is whether you exercised professional judgment over this work and can demonstrate it — and an exception-based review with a complete audit trail answers that better than the status quo does, not worse.
The claimed gains do not survive measurement. METR ran a randomized controlled trial and found experienced developers were 19% slower with AI tools while believing they were 20% faster. MIT reported that 95% of enterprise GenAI pilots produced no measurable P&L impact. You are selling a productivity story the evidence does not support.
Both studies are real, and both have been materially undercut since — which a reader in August 2026 needs to know, and which most people citing them do not mention.
METR redesigned the experiment and got a different answer. The July 2025 study did find a 19% slowdown across 16 developers and 246 tasks, with a confidence interval of +2% to +39%.34 In February 2026, in a post titled We Are Changing Our Developer Productivity Experiment Design, METR reported that in the original design "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI" — a selection bias that removes exactly the tasks AI would handle best. A later study produced an 18% slowdown for the subset of original developers who took part, with a confidence interval of −38% to +9%, and a 4% slowdown for new developers, at −15% to +9%. Both intervals cross zero. METR's own current position: "we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025."35
Two honest notes, because this is where people get sloppy in both directions. METR did not retract or revise the 19% figure — it stands as published; the later numbers come from a separate study. And METR says its own data give only "very weak evidence" for the size of the improvement. The correct reading is not "AI makes developers faster." It is that the headline number is a great deal less settled than the people quoting it at you will suggest.
The MIT "95%" figure is weakly sourced. It comes from an MIT NANDA report built on structured interviews with representatives of 52 organizations and 153 survey responses gathered at four industry conferences, whose own limitations section concedes the figures are "directionally accurate based on individual interviews rather than official company reporting."36 Wharton's Kevin Werbach, on the 95%: "I've read through the document multiple times, and I still can't understand where it comes from." Futuriom's assessment is that the report "lacks thorough sourcing" and is "not a traditional academic research paper," and notes that "some of the authors work at some of these vendors."37
This industry has absorbed CAD, BIM, cloud collaboration, and a decade of construction SaaS. Productivity went down. Why would AI be different — and why would an AI company succeed where Autodesk, with a twenty-year head start and every firm as a customer, has not?
This is the best objection in the list, and Figure 1 is the evidence for it, not against it. Structural engineers have said it out loud in their own trade press: BIM "is revolutionizing design," modular assembly is doing the same for construction, and yet aggregated industry data shows productivity "in a neutral to negative trend for many years."6b
But read the failure precisely. Every one of those technologies was sold as a tool, to a firm, which then had to reorganize itself to capture the value — and did not. Tools do not change operating models; they get absorbed into existing ones. A firm that buys BIM and keeps the same proposal process, the same document control, the same handoffs, and the same overhead structure gets a better drawing at the same cost.
That is precisely the argument for owning the firm rather than vending to it. Elad Gil put it as cleanly as anyone: "If you own the asset, you can [transform it] much more rapidly than if you're just selling software as a vendor."48 (The brackets are TechCrunch's — its editorial substitution, not Gil's word.)
The reverse trade is also instructive. On 27 November 2025, AECOM announced the acquisition of the Norwegian AI startup Consigli — self-described as "The Autonomous Engineer" — for $390 million. ENR's analysis identified four headwinds: "the internal user base — even at AECOM's scale — may be too small to justify long-term platform investment"; "competing vendors are embedding similar capabilities into off-the-shelf tools, eroding differentiation"; "data readiness remains uneven"; and "the economics of time-and-materials contracts limit how much efficiency can be monetized."46 ENR also characterized "the immediate drop in AECOM's share price" as suggesting investor skepticism — that reading is ENR's, and I could not corroborate the price move independently, so treat it as their assessment rather than established fact.
That fourth headwind is the deepest point in this entire paper: a firm billing hourly has a structural disincentive to reduce hours. An AI-native firm that prices on outcomes does not.
Engineering work is not won on price or efficiency. It is won on qualifications, named credentialed individuals, and reference projects you do not have. And your professional liability insurer has not decided how it feels about any of this.
The procurement half is correct and, for a new entrant, close to disqualifying — which is why it drives the recommendation in §9 rather than being waved away.
Under the Brooks Act (40 U.S.C. §§1101–1104), federal agencies must publicly announce A/E requirements, discuss with at least three firms, rank at least three on "demonstrated competence and qualification," and only then negotiate price with the highest-ranked firm.28 Price is not a selection criterion. Forty-six states had adopted some form of qualifications-based selection as of NSPE's 2018 count.28b A cost-advantaged entrant cannot buy its way onto the shortlist.
Worse, the federal instrument itself is built around named humans. Standard Form 330 requires a resume for each key person, their current professional registration by state and license number, their years of relevant experience, and a matrix cross-referencing each named individual against the specific past projects they personally worked on in a similar role.29 There is no field for "our model." State DOT prequalification is stricter still: FDOT Rule 14-75 requires that listed personnel be bona fide employees, that engineers hold Florida registration, that experience thresholds be met personally, and it formally grades past performance on schedule, management, and quality — with unsatisfactory grades triggering suspension.30
On insurance, the picture is less alarming than commonly claimed. Ames & Gough's 2026 survey of fifteen leading A/E professional liability carriers found 80% view AI adoption as "a potential disruptor of the overall professional liability market," with 73% "planning modest rate increases largely in the single digits."31 The AIA Trust's 2025 review found carriers "focused less on the technology itself and more on quality control, professional judgment, and potential copyright concerns," and concluded that "negligence is evaluated based on outcomes and professional responsibility, regardless of the tools used."32 AI-specific exclusions certainly exist — Berkley's "absolute" AI exclusion across D&O, E&O and fiduciary; Hamilton's generative-AI exclusion in professional liability; ISO's optional January 2026 endorsements CG 40 47, CG 40 48 and CG 35 08 — but all of those are adjacent lines, not design professional E&O.33 No professional liability claim over AI-generated design work appears in the public record as of this writing. And construction counsel have flagged the opposite possibility as an open question: "whether the use of AI will be required to meet the applicable standard of care for designers."32b
The industry will not do this to itself. That is the opportunity.
There is a common view that engineering is a reactionary field — dominated by old hats who believe that if it ain't broke, don't fix it. The view is directionally right about the outcome and wrong about the cause, and the difference matters enormously, because the real cause is not temperament. It is written into the law of the profession.
Engineering is not conservative about technology. It is conservative about accountability.
The evidence against the simple "dinosaurs" story is easy to find: this is a field that absorbed finite element analysis, performance-based seismic design, high-performance concrete, parametric modeling, and drone-based survey without much drama. Engineers are not afraid of new methods. What the profession is structurally slow about is changing who is accountable for what, and how that accountability is demonstrated.
And there is a specific legal mechanism that produces this, which almost nobody outside the professional-liability world names out loud. The standard of care — the thing a design professional is legally measured against — is not codified in a statute; it is a common-law test, commonly restated as "the degree of care and skill ordinarily exercised by practicing professionals performing similar services under similar circumstances."32c It reaches contract in black and white: AIA B101-2017 §2.2 obliges the architect to perform "consistent with the professional skill and care ordinarily provided by architects practicing in the same or similar locality under the same or similar circumstances."32c
Read that carefully. Correct practice is legally defined as what everybody else is already doing. That is a conservatism engine, and it runs on its own. It makes innovation asymmetrically risky in a way it is not in almost any other industry: follow the herd and fail, and you met the standard of care; innovate and fail, and you did not. A prudent principal facing that payoff structure waits — not because they are a dinosaur, but because waiting is the rational play under the rules they are actually judged by.
What the resistance actually looks like in the data
The result is a profession that overwhelmingly believes the change is coming and overwhelmingly has not made it. This gap is measurable, and it is remarkably consistent across three independent surveys of three different populations.
The texture underneath those numbers is worth stating plainly:
- 52% still use paper during the design phase, 49% during planning, and 43% still depend on physical signatures — and only 11% describe themselves as fully digital organizations.9 This is not a field that skipped digitization and is leaping to AI. It is a field that has not finished digitizing. Note again that these are technology decision-makers at manager level or above — the digitally most engaged slice of the industry.9b2
- 90% of architectural professionals report concern about a cluster of five issues — inaccuracy of AI outputs, unintended consequences, security, authenticity, and transparency.8 (AIA reports the 90% across the bundle, not for inaccuracy alone.) Enthusiasm and alarm coexist in the same respondents.
- Sentiment has cooled, not warmed. Autodesk's 2025 survey found 69% of leaders believed AI would enhance their industry — down twelve points from the prior year — while the share saying AI would destabilize their industry rose seven points to 48%.11
- Capability is concentrated in exactly the firms that need it least. 61% of large architecture firms use AI in day-to-day work against 27% of small ones7 — and 89.1% of US A/E firms have fewer than twenty employees.12 The overwhelming majority of the industry has neither the budget nor the staff to build any of this.
- The billing model punishes efficiency. ENR's analysis of AECOM's $390M Consigli acquisition identified monetization constraints from time-and-materials contracts as a structural headwind.46 A firm that bills by the hour and halves the hours has cut its own revenue. There is no version of the incumbent solving this from inside a T&M contract.
- Regulatory silence reads as risk. 69% of surveyed AEC professionals said regulatory uncertainty affected their plans;9 NCEES's Model Law and Model Rules, in their most recent editions, contain no mention of artificial intelligence at all.20 To a licensee, an unaddressed question is not permission — it is exposure.
And behind all of it sits the demographic fact. An estimated 184,175 engineers retired or left the field in 2022, 85,175 from the core civil, mechanical, and electrical disciplines, while degrees awarded have fallen more than 10,000 from their 2019 peak of about 214,000.15 The people leaving hold the institutional knowledge. The people who would rebuild the operating model are not arriving in sufficient numbers to do it.
The strategic reading: this is the moat, not the obstacle
Every fact above is usually presented as a reason to avoid this industry. It is the opposite. A market that changes fast has no room for an outsider — by the time you arrive, the incumbents have already done it. A market with a conservatism engine wired into its liability regime, a billing model that penalizes efficiency, a size distribution where 89% of firms cannot fund a platform, and a retirement wave draining the knowledge base is a market where the window stays open long enough to matter.
It also explains why the entrant should be a generalist rather than a specialist. Firms saturated with engineers optimize the engineering. The unexploited surface — contracts, proposals, project controls, document management, the entire production system connecting them — is invisible from inside the discipline precisely because everyone inside it was trained to look somewhere else. Seeing the whole operating system as the target requires standing outside it.
The consolidation data confirms the window is open and also that it is not open forever. Record transaction volume, 72% of 2025 deals involving firms under $10 million in revenue, and private equity in more than half of all design and environmental consulting deals13 means the founder-owners are already selling. The question is only who they sell to, and whether the buyer brings an operating model or just a balance sheet.
The line this argument must not cross
Read this section as "the incumbents are dinosaurs and we will route around them" and you will build something dangerous. The caution is not irrational — it is calibrated to consequences. A bad blog post is embarrassing; a mis-designed beam kills people. The profession's conservatism is the reason bridges mostly stand up, and any entrant who treats it as an obstacle to be defeated rather than a requirement to be satisfied will eventually earn the enforcement action they deserve.
The correct response to a conservatism engine is not to argue with it. It is to feed it what it wants: verifiable output, complete provenance, retained review records, named licensed professionals in genuine responsible charge, and a measurable track record. Do that and the standard of care stops being a wall. Construction counsel have already raised the reverse possibility — that AI use may eventually be required to meet the standard of care.32b The firm that has been documenting its verification discipline for three years when that turns is not the laggard. It is the new baseline.
Build a team, or partner with an established firm?
This is the question the paper exists to answer. The honest answer is that framing it as either/or is the mistake — and the evidence points clearly to a staged structure in which the first move is partnership and the second is ownership of the platform, not the practice.
What building from zero actually costs — in months, not dollars
The individual steps to stand up a licensed engineering firm are each modest. The problem is that they are largely sequential, and one of them cannot be accelerated at any price.
A correction worth making explicitly, because the commonly repeated version is wrong. FDOT professional services prequalification requires a current overhead (FAR) audit — not three years of audited financial statements. The three-year audited-financials requirement belongs to FDOT's separate construction contractor prequalification under Ch. 337.14, Fla. Stat.30 The two programs are routinely conflated in secondary sources. The real time-lock is different and stronger: a new firm has no cost history to audit, and — decisively — no graded past-performance record and no completed projects to map to named staff. SF 330's Section G matrix demands exactly that cross-reference,29 and FDOT grades past work on Schedule, Management and Quality on a 1–5 scale.30 You cannot manufacture completed projects. Multiply the whole stack by each additional state.
Stack the published ranges and it is roughly six to thirteen months before a new entity can bid public work in a single state — with a hard floor set by financial history that no amount of capital compresses. Against that: BizBuySell's 2025 architecture and engineering benchmarks put the median small-firm sale price at roughly $825,000 on median seller's discretionary earnings of about $451,450, while Zweig Group's 2025 data give a median of 4.28× EBITDA across all size bands, with small firms transacting nearer 2.5–3.5× on a blended SDE/EBITDA basis.51 You can buy licensure, registration, prequalification, audited history, a client list, and a data corpus of prior projects for less than a seed round.
What the AEC precedent actually shows
Almost nothing. That is the finding, and it cuts both ways. AI funding into AEC startups reached $616 million across 46 rounds in the first half of 2026 alone, versus $318 million across 31 startups in all of 2025 — and virtually all of it went to companies selling software to firms, targeting what one sector analysis called "point solutions" addressing "specific operational failures."52 The traffic runs the other direction: AECOM bought Consigli,46 Autodesk bought MaintainX for $3.6 billion.46b
The one greenfield precedent is cove.tool, which in March 2025 launched "Cove Architecture," billed as the first full-service AI-powered architecture practice, on the back of a claimed "over $25 million" and "nearly a decade" of R&D. Its inaugural project was a 15-unit housing complex in Atlanta's West End, with claims of 60% shorter design timelines and completion in 15 days.45 It is roughly eighteen months old with a small public record; no source states the practice's firm-level registration, though Cove's own team page lists two AIA-credentialed principals, so it is not unstaffed by credentialed architects.45b It is a signal, not a proof. Its CEO's own framing is notably close to §5's argument: "You need experienced individuals who understand the outputs. You don't want an intern pressing a button and making a building."45
Outside AEC the precedent is much richer, and it points one way. Eudia — backed by a General Catalyst-led Series A of up to $105M — acquired Johnson Hana, a 300-plus-person legal services provider, in July 2025, with co-founder and CEO Omar Haroun giving the explicit reasoning: "we've learned that building truly transformative systems requires deep human expertise embedded within the technology itself."42 Crete Professionals Alliance committed over $500 million to acquiring US accounting firms, backed by Thrive Capital, and already spans 20+ businesses and $300M+ in revenue.43 The UK's Solicitors Regulation Authority authorized Garfield.Law in May 2025 as the first AI-driven law firm — with conditions: the system cannot propose case law, clients must approve actions, and named regulated solicitors remain ultimately accountable for all outputs.41 That conditions list is a preview of what an engineering board would impose.
And in insurance, the staged path is well established: startups begin as managing general agents writing business on a licensed carrier's paper, prove the model, then move to full licensure — Next Insurance, Clearcover, Metromile and Kin all made that transition.57 (Not universal, and worth saying so: Lemonade and Root launched as licensed carriers from the start, so the pattern is a common route rather than the only one.) The closest structural parallel to Oxvelo's question is Pie Insurance, which raised $127 million in 2020 and earmarked $100 million of it to "form and purchase licensed insurance companies" through a dedicated affiliate.57b Note what Pie did: it capitalized the build-or-buy decision rather than assuming an answer.
The comparison, scored
| Requirement | Build a team from zero | Partner with / acquire a firm | Staged hybrid |
|---|---|---|---|
| PE licensure & certificate of authorization | 8–14 mo | Day one | Day one |
| Auditable cost history for FAR overhead audit | None to audit | Inherited | Inherited |
| Graded past performance & reference projects (SF330, DOT) | None | Existing record | Existing record |
| E&O history & insurability | Unrated | Established | Established |
| NY ownership compliance for a non-licensee founder | Blocked | Minority only | MSO owns platform |
| Proprietary data corpus for the AI layer | Empty | Decades of projects | Decades of projects |
| Freedom from legacy process & hourly-billing incentives | Total | Inherits both | Isolated in the MSO |
| Enterprise value capture / multiple | Services multiple | Services multiple | Platform multiple |
| Capital required to first revenue | High, slow | ~$0.8–3M | Low — MSA first |
| Regulatory precedent for the structure | Conventional | Conventional | Proven in 3 professions |
Scoring is Oxvelo's judgment. Every row's underlying constraint is cited elsewhere in this paper; the ratings are not.
The recommendation
Partner first, own the platform, acquire second. Building an engineering team from zero fails on the one requirement that cannot be compressed — three years of audited financials and a graded past-performance record. Acquiring a firm outright as the opening move buys those, but spends scarce capital on a services multiple before the AI layer has proven it changes the economics, and — in New York — cannot legally be controlled by a non-licensee founder anyway.
The staged path resolves all three. Stand up Oxvelo as the management and technology entity. Contract with an existing licensed design firm under a management services agreement, exactly as AI-native law and accounting platforms have done. Prove the operating-model change on their book of business, against their data, under their seal. Then decide — with evidence rather than conviction — whether to acquire the practice, replicate the structure across states, or stay a platform.
Six moves, in order.
- 1
Get the New York structure ruled on before anything else
Retain a New York professional-licensing and corporate attorney to answer two questions in writing: whether an outside non-employee entity may hold the sub-25% BCL §1503 tranche, and whether a percentage-of-revenue management fee from an engineering MSO would be characterized as unlawful fee-splitting.2456 Everything downstream depends on these answers. This is a four-figure decision that determines a seven-figure structure.
- 2
Write the constitution before writing the software
A founding operating document that defines the AI workforce, its permissions, its escalation rules, and — most importantly — the human control points and the audit record each one produces. In New York that record is not optional: a licensee reviewing work prepared by others must retain a written evaluation for six years.59 Build the audit trail as a first-class feature, not a compliance retrofit. It is also the artifact that makes the firm insurable and the structure defensible.
- 3
Map one firm's complete value stream, from inquiry to closeout
Client acquisition, opportunity qualification, proposals, estimating, contract negotiation, project setup, technical production, QA/QC, project controls, construction support, billing, collections, closeout. For each stage: tasks, decisions, documents, handoffs, responsible parties, recurring pain points, data inputs. This is the automation blueprint, and it is also the diligence document for choosing a partner.
- 4
Find the partner firm — and screen for the right one
Target profile: 15–60 people, one to three states, an owner within five years of exit, a strong past-performance record, and — critically — a principal who is frustrated rather than complacent. Avoid firms whose economics depend on maximizing billable hours on time-and-materials contracts; that is the structural disincentive ENR identified in the AECOM/Consigli deal.46 Look for firms with fixed-fee or lump-sum work, where reducing hours increases margin instead of reducing revenue.
- 5
Hire for judgment, not for production capacity
Two or three unusually senior licensed professionals, hired explicitly into the three roles in §5: knowledge architect, reviewer in responsible charge, and creative on the novel. Screen for people who can articulate why a detail is right, not just draw it — the encoding work requires someone who can externalize judgment, which is a rarer skill than exercising it. Do not hire a production team; the entire thesis is that the production team is the thing being replaced. Note that the licensed reviewer must sit inside the professional entity, not the management entity, or the structure in Figure 11 fails on its own terms.
- 6
Build one end-to-end demonstration: opportunity to project kickoff
Identify an opportunity, draft the proposal, analyze the contract against a company standard, build the fee and schedule, create the project workspace, establish the deliverable register, generate the kickoff package. Multiple AI workers collaborating, one human approving at each defined control point, every step logged. Measure the cycle time against the firm's historical baseline — measured, not self-reported, for the reason METR's participants demonstrated.35 One credible before-and-after number on a real project is worth more than the rest of this document.
What would prove this wrong.
A thesis that cannot be falsified is a slogan. Four specific outcomes would mean this strategy is wrong, and they should be watched deliberately rather than rationalized:
- The kickoff demonstration does not beat the baseline on measured cycle time. If a full AI workforce cannot compress opportunity-to-kickoff on a real project against that firm's own history, the operating-model claim is false and no amount of structure fixes it.
- The leverage ratio stalls near 2. The whole §5 argument rests on L climbing well past the doubling that McKinsey's "50% of non-physical work" implies.2 If measured, audited throughput on repeatable packages plateaus at 2–3× with an honest review standard held constant, then this is a good consulting business and not a revolution — and the paper's boldest claim should be retracted rather than re-explained.
- Reviewer catch rates fall as volume rises. The cleanest single early-warning metric in the whole model. If seeded-defect detection degrades as packages-per-reviewer climbs, the leverage is buying itself with silent risk, and the cap has already been exceeded.
- The partner firm's staff route around the system. The Bluebeam data showing 52% of AEC professionals still using paper in the design phase and 43% still relying on physical signatures9 is a warning about workflow gravity, not just about digitization. If the licensed professionals treat the AI layer as an extra step rather than a replacement step, the productivity never materializes — which is precisely how BIM's promise dissolved.
- A state board or insurer moves against the structure. Currently no A/E professional liability carrier is writing AI exclusions,32 and Texas's board has explicitly declined to write AI-specific rules.22 Both could change. NCEES's Model Law and Model Rules, as of their most recent editions, contain no mention of artificial intelligence at all20 — the governing framework is silent, which is currently permissive and could stop being.
- The economics land on a services multiple. If the platform never separates from the practice, Oxvelo has bought a small engineering firm with better software — a fine business, valued at 2.5–3.5× EBITDA,51 and not the one described here. Nines is the documented version of this failure.47
Glossary
This paper spans four vocabularies — engineering licensure, public procurement, private-equity finance, and machine learning — and almost nobody is fluent in all four. Every term below is also tappable inline throughout the paper.
How to read the numbers in this paper.
Not all evidence is equal, and a document that presents a peer-reviewed benchmark and a vendor press release in the same typeface is quietly misleading its reader. Every figure in this paper falls into one of five tiers. Where a claim rests on tier 3 or below, the text says so at the point of use.
Two figures deliberately excluded
The widely quoted "35% of construction professionals' time is non-productive, costing $177 billion annually" is real and well-sourced, but its survey population was 49% general contractors and 36% specialty trades — not design professionals.18 Applying it to engineers and architects would be the kind of small dishonesty that costs a reader's trust in everything else. The US Census SUSB firm-size figures for NAICS 5413 were removed after they could not be verified from a retrievable source; the fragmentation argument was rebuilt on AIA and ENR data instead.
The leverage arithmetic, worked.
Figure 5 compresses a calculation into a curve. Here it is in full, so a reader can disagree with a specific assumption rather than with the shape of a line.
Definitions
one licensed professional can review and take responsibility for
B = baseline net revenue per employee = $143,000 (AIA, 2023 data, mean)
C = total labor cost per employee = $119,511 (Deltek Clarity 47th)
O = overhead rate on direct labor = 161.3% (Deltek Clarity 47th)
U = utilization = 58.9% — share of paid hours that are billable
The curve
L = 1 → $143,000 (the industry today)
L = 2 → $286,000 (McKinsey-implied: 50% of non-physical work)
L = 5 → $715,000 (conservative near-term target)
L = 10 → $1,430,000
L = 25 → $3,575,000
L = 50 → $7,150,000 (2 reviewers carrying a 100-engineer production load)
The four assumptions, and what breaks each
- Price per deliverable holds constant. It will not, in a competitive market. But note what happens if it falls: revenue per professional rises more slowly while margin still rises, because cost per deliverable falls faster than price. The leverage shows up in the P&L either way — the question is only whether the client or the firm captures it.
- The contract is fixed-fee or lump-sum. This is the assumption that actually decides the business. On time-and-materials, halving the hours halves the revenue and the entire curve inverts into a loss. ENR identified precisely this as a constraint on AECOM's Consigli acquisition.46 Partner selection must screen for fee structure.
- Review effort does not scale linearly with output. If it does, L is capped near 1 and the thesis fails outright. The entire verification architecture in Figure 6 exists to break that linearity, and the assurance layer exists to detect it if we are fooling ourselves.
- The work is templatizable. L applies to repeatable production against established standards. On first-of-its-kind design, L stays near 1 and should.
What the model does not claim
It does not claim a hundred-person firm becomes two people. A hundred-person engineering firm is not a hundred people drafting — it contains project managers, construction administration staff, principals, finance, and business development. The curve applies to the design production function only. The other functions compress on their own separate curves, described in Figure 3, and none of them reaches L = 50.
It also does not claim any of this has been achieved. L is a measurement to be taken, not a result to be assumed — which is why §11 names a stalled leverage ratio as the first thing that would falsify the whole paper.
Regulatory quick-reference.
The three states this paper examines, side by side. Nothing here is legal advice — it is a starting point for a conversation with counsel, and the New York column in particular has a load-bearing ambiguity that only an attorney's opinion can close.
| Question | New York | California | Texas |
|---|---|---|---|
| Governing provision | Educ. Law §§7202, 7210; BCL §1503(b-1) | Bus. & Prof. Code §6738 | 22 Tex. Admin. Code §137.77 |
| Non-licensee ownership | Ordinary PSC: none. DPC: <25%, but reserved for ESOPs and non-licensee employees — an outside investor entity does not qualify | Permitted, but a non-licensee may not be sole owner | No ownership restriction in the rule |
| Licensed leadership required | President, chairperson and CEO must all be design professionals; largest shareholder too | A licensed engineer must be owner, partner, or officer in charge | At least one full-time active licensee employed |
| Board / director composition | >75% design professionals | Not specified | Not specified |
| LLC permitted? | PLLC, all members licensed | No — Corp. Code §17701.04(e) | Yes |
| Branch coverage | All engineering | §6738 reaches civil, electrical, mechanical only | All engineering |
| Practical read for Oxvelo | MSO/PC required. No equity route into the licensed entity | Workable with a licensed principal in charge | Most permissive of the three |
The questions counsel must answer before any structure is committed
- Can an outside, non-employee entity hold any equity in a New York design professional service corporation? The statutory text points to no. If that is right, the MSO/PC split is not a preference — it is the only lawful structure.
- Would a percentage-of-revenue management fee from an engineering MSO be characterised as unlawful fee-splitting, as it is in the New York healthcare analogue? If yes, the fee must be flat, time-based, or cost-related, and the financial model must be built that way from the start.
- Does the six-year written-evaluation retention duty under 8 NYCRR §29.3(a)(3)(i) attach to AI-produced work reviewed by a licensee? Assume yes and build the audit trail accordingly; it is cheap insurance either way.
- In which states does the licensed entity need a certificate of authorization before it can be named on a proposal, and what is the current processing time in each?
The client question — which nobody has asked clients.
Every principal who reads this paper will arrive at the same objection within five minutes: will my clients accept it? The paper owed an answer. Researching it produced a finding more interesting than the answer itself.
The finding: the question has never been put to clients
We checked every substantial AEC survey on artificial intelligence we could find — RICS (2,200+ professionals), RIBA's 2025 and 2026 reports, Chaos (~800 architects), Arup (5,000 professionals across ten countries), Mastt, Deltek Clarity (896 firms), and the AIA/Deltek/ConstructConnect study.60 Every one of them surveys practitioners. Not one asks clients whether they want their design consultants using AI.
The nearest owner-side data is Dodge Construction Network's work with the National Institute of Building Sciences, which surveyed roughly 200 US owners each managing $5M+ of annual construction. It found 28% of owners already using AI themselves, with the fastest growth rate of any technology studied, and delivered this warning to their consultants: data capability "is a competitive advantage now, but it is only a matter of time before it is a requirement."61 That measures owners' own adoption. It does not measure their attitude to yours.
Why an absence is worth publishing
An industry has spent three years debating whether AI belongs in design practice without asking the people who pay for design. That is not a small oversight — it is a measurement gap sitting exactly where the commercial risk lives. It is also, for a firm willing to do the work, an opportunity: the study we are fielding asks it directly.
What is arriving is not client objection
Three real forces are moving, and none of them is a client refusing AI-assisted work.
1. Public buyers are mandating disclosure
California's Department of General Services amended its State Contracting Manual (§2302) to require that a bidder notify the State in writing if it intends to use generative AI to complete work that materially affects system functionality, risk to the State, or contract performance. It applies to all written solicitations regardless of acquisition type, and non-disclosure "may result in bid disqualification."62 This is the only binding, generally applicable US buyer-side disclosure requirement we could verify.
Federal policy is softer. OMB M-25-22 says agencies "are advised to consider" requiring contractors to disclose AI use — advisory, not mandatory.63 GSA circulated a draft clause in March 2026 requiring disclosure of all AI systems used in contract performance; comments closed in April and it was not included in the subsequent refresh.64 And notably, Caltrans' own scan of state GenAI procurement frameworks identified no state DOT with a binding policy requiring consultants to disclose AI use in deliverables.65
2. Insurers are asking questions — buyers are not
NSPE's professional-liability reporting found that most underwriters "are not yet incorporating AI into underwriting criteria, but they are beginning to ask firms how they supervise and review AI-assisted work," with "no widespread exclusions or special policy modifications yet evident."66 Ames & Gough's carrier survey adds that firms with proactive AI governance will "gain underwriter confidence."31 We found no evidence of AI questions appearing in client prequalification questionnaires. The scrutiny is coming from the people carrying the risk, not the people buying the service — which is a meaningful distinction, and an argument for building the audit trail early.
3. No standard-form AI clause exists yet
Neither AIA, EJCDC, nor ConsensusDocs has published an AI clause or addendum for design agreements as of August 2026.67 Where AI provisions are appearing, the direction is instructive: they run contractor-to-owner as protective disclaimers, not owner-to-designer as restrictions. One construction attorney describes drafting a provision stating that where AI is used for scheduling or preliminary estimates, "the customer cannot rely on those until such time as boots on the ground have determined whether those field measurements are accurate."68
The real commercial risk is the fee, not the refusal
The threat this paper takes most seriously is not that clients reject AI-assisted work. It is that they accept it and expect to pay less for it.
In professional services broadly, that is already measurable: a 2026 survey of 1,000+ senior business leaders found 66% reporting that clients are becoming more demanding while less willing to pay, with two-thirds expecting AI to reduce costs, and a third saying clients now see AI as a route to bringing work in-house rather than hiring specialists.69 That study is cross-sector, not A/E — the transfer is an assumption, and we label it as one.
In A/E specifically, what is documented is a shift in the contracting model rather than the fee level. AECOM's CEO Troy Rudd describes clients themselves raising it: moving "away from something like cost plus to something that looks more like a fixed fee."70 And ACEC's work with Virginia Tech states the underlying danger with unusual bluntness:
That is the single strongest external corroboration of this paper's commercial argument, and it arrives from the engineering industry's own trade body rather than from a technology vendor. It is also why the leverage calculator makes fee mix the setting that dominates every other input.
What we could not find at all
- No published case of a client rejecting or disputing AI-produced design deliverables. The nearest real disputes are bid protests dismissed by the GAO for AI-fabricated legal citations — five decisions in 2025 — plus a contested Army source-selection where the buyer's own AI tool was alleged to have hallucinated evaluation findings on a $449M engineering services award.71 In both categories the AI failure sits in the procurement paperwork, not the engineering.
- No evidence that any owner scores AI capability in A/E selection. Every search on this returned results about using AI to write proposals — the opposite question.
- No A/E-specific instance of a client demanding a fee reduction on AI grounds. The mechanism is credible and the adjacent-sector evidence is quantified, but the A/E case has not been documented.
Three genuine blanks. In a paper that concedes three of its seven objections, it would be inconsistent to dress absences up as answers.
Oxvelo's AI disclosure policy.
NSPE's Board of Ethical Review found there is no professional or ethical obligation to disclose AI use to a client unless the contract requires it — while adding that "ethical principles favor transparency when AI plays a substantial role in generating work products."21 That is permission to say nothing. We think saying nothing is the wrong call for a firm whose entire proposition is verifiable accountability, so this is our standing policy, published in advance of having clients.
1. We disclose by default
Any deliverable produced with material AI involvement carries a statement of that fact, whether or not the contract requires it and whether or not the client asks. Disclosure is not an admission; it is a description of method, in the same way a firm describes which code edition it designed to.
2. We name the human, not the model
Every sealed deliverable identifies the licensed professional in responsible charge by name and licence number. No document will ever suggest that a system approved anything. Responsibility is personal, it is legally required to be personal, and diffusing it across a "platform" would be the first step toward the plan-stamping failure this paper warns about.
3. The review record is available to the client on request
The exception set a reviewer adjudicated, the deterministic checks that ran, and the provenance links behind cited code provisions are retained and disclosable. In New York this is already obligatory — a licensee sealing work prepared by others must retain a written evaluation for six years59 — so we treat it as a client-facing feature rather than a filing-cabinet duty.
4. Client data does not train anything
Client project data is not used to train third-party models, and is not routed to any service that reserves the right to do so. NSPE's Case 24-2 found the engineer at fault partly for disclosing private client information to an AI tool without permission — a failure of data handling, not of design.21 Where a client's contract restricts data routing, that restriction governs model selection, not the other way around.
5. We answer the public-buyer question before it is asked
Any bid into a jurisdiction with a disclosure regime — California's DGS requirement being the clearest62 — carries the disclosure whether or not the solicitation demands it. The downside of over-disclosing is a conversation. The downside of under-disclosing is disqualification.
6. We do not claim what we cannot show
No efficiency, accuracy or cycle-time claim is made to a client without a measured baseline behind it. This is the same standard we hold our own published research to, and it exists because METR's participants believed they were 20% faster while measuring slower.35
Why publish this at all
Because a policy written after the first difficult conversation is a rationalisation, and everyone can tell. Publishing it now — before there is a client, a claim, or an incentive to soften it — is the only moment at which it costs nothing and means something. If we ever depart from it, this page is the evidence.
Worked example: opportunity to project kickoff.
Read this correctly
This is a design specification, not a result. No project has been run through this process. Nothing here is a measurement, a case study, or a claim about performance. It is the blueprint for the demonstration described in §10 — published so that when the demonstration runs, it can be judged against what was specified beforehand rather than against a narrative written afterwards.
The paper argues its mechanism at length and never shows it. Here is one complete workflow, stage by stage, with the human control point and the audit artifact named at each step.
| Stage | AI workforce action | Human control point | Audit artifact produced |
|---|---|---|---|
| Opportunity intake | Parses the RFP or inquiry; extracts scope, deliverables, deadlines, submission requirements, evaluation criteria and disqualifiers | None — retrieval only | Structured opportunity record with every extracted term linked to its source page |
| Go / no-go | Scores fit against the firm's licence coverage, prequalification status, staff availability and past-performance record; flags anything the firm cannot legally satisfy | Principal decides. The system may not advance an opportunity | Dated decision record with the reasoning presented and the decision taken |
| Fee and scope | Retrieves comparable prior projects, builds a first-pass effort model by discipline and phase, drafts the scope narrative from firm-standard language | Principal sets the fee. Non-negotiable — this is a commercial commitment, not a calculation | Effort model with each assumption traced to the prior project it came from |
| Proposal assembly | Assembles the SF 330 or private-format response; populates named-staff resumes, registrations and matched project experience from the personnel record | Named individuals confirm their own resume entries and registration status | Per-person confirmation log |
| Contract review | Compares the proposed agreement against the firm's standard positions; ranks every deviation by risk with the firm's standard clause shown beside the client's | Principal accepts or rejects each flagged deviation. Signature is a human act | Deviation register with disposition and rationale per item |
| Project setup | On execution, creates the workspace, budget structure, phase and task breakdown, and compliance log | Project manager confirms structure | Setup manifest, timestamped against contract execution |
| Deliverable register | Derives every required deliverable from the contract scope and schedule, with owner, due date and dependency | Project manager confirms completeness | Register with each line traced to the contract clause requiring it |
| Kickoff package | Drafts agenda, roles, standards to be applied, code editions in force, submittal schedule and communication protocol | Project manager approves and issues | Issued package with revision history |
What is deliberately absent
No stage above produces sealed engineering work, and none advances without a named human decision at the points where commitment or liability attaches. The seal appears nowhere in this workflow, because this workflow ends at kickoff — which is precisely the point of entering through the beachhead in Figure 4.
The measurement that decides whether any of this is true
One number, taken the same way twice:
across the partner firm's last 10 comparable pursuits
Measured = the same interval, same definition, on the next 5 pursuits run through the system
Also tracked: professional hours consumed per pursuit · win rate
exceptions raised per pursuit · exceptions overturned on review
Elapsed time is the headline, but exceptions overturned on review is the number that matters most. If reviewers almost never overturn what the system flags, the system is not surfacing genuine judgment calls and the review is decorative. A healthy exception rate is not zero.
Win rate is included because it is the one metric that could invalidate everything while every other number improves. A faster, cheaper proposal that wins less work is not an operating-model improvement. It is a worse business with better dashboards.
Verification log.
Every claim in this paper was checked against its primary source in a dedicated fact-checking pass. Corrections made in that pass are recorded here, because a reader is entitled to know which way the errors ran before they were fixed.
Corrections that changed the argument
- NCEES redefined responsible charge in August 2025. The widely quoted "direct control and personal supervision" formulation is superseded and appears nowhere in the current Model Law. §5 now argues from the current four-duty test — authority to review/change/reject/approve, personal awareness of scope, capability to answer questions, acceptance of full responsibility — which is more favourable to this thesis than the language it replaced.
- The Census fragmentation figures could not be verified from any retrievable source and were removed. Figure 7 now uses AIA's published firms-and-billings table; the concentration claim was rebuilt on ENR Top 500 revenue against Census Quarterly Services Survey industry revenue.
- The "three years of audited financials" prequalification claim was wrong. That requirement belongs to FDOT's construction contractor programme, not professional services, which requires a current FAR overhead audit. The binding constraint was restated as the graded past-performance record — better sourced, and harder to overcome anyway.
- New York's ownership position is worse than first drafted. BCL §1503(b-1) reserves the residual tranche for ESOPs and employees; an outside investor entity does not qualify at all. The effective ceiling for a holding company is zero, not 25%.
- METR did not retract its 19% finding. The February 2026 figures come from a separate study, not a revision. The earlier draft overstated the walk-back in this paper's favour.
Corrections of attribution and precision
- $143,000 net billings per employee is a mean; AIA publishes no median.
- AECBench is Liang et al., not Zhao et al. Zhao is the last author.
- MGI's $228 billion covers the entire US AEC industry and includes "other automation technologies" — it is not the dollar value of the 50% automation figure, and the two are frequently conflated.
- ACEC's 14% preparedness figure is estimated by mid-career and older practitioners, not executives; the 34% student figure comes from a 195-person convenience sample.
- Bluebeam's 27% surveyed technology decision-makers at manager level or above in five countries — not "AEC professionals globally."
- Dodge's 87% concerns AI's impact on construction as an industry, not on the respondent's own business.
- AIA's 90% covers a bundle of five concerns, not inaccuracy alone.
- Lemonade and Root launched as licensed carriers — they are counterexamples to the insurtech staged path, not examples of it.
- Engineering degree-completion figures come from NCES via ACEC, not ASEE, which counts a different universe.
- The Elad Gil quotation contains TechCrunch's editorial insertion and is now printed with the brackets intact.
- ENR's characterisation of AECOM's share-price move could not be independently corroborated and is now attributed to ENR rather than stated as fact.
What remains unverified, by design
Two sets of figures are company-reported and unaudited: cove.tool's project metrics (60% timeline reduction, 95% estimate accuracy, 40% iteration cost, 15-day completion) and Crosby's contract volumes. Both are flagged in place and used as existence proofs, not measurements. One structural claim — that AI-specific exclusions are not standard in A/E professional liability policies — is an inference from the absence of documented exclusions in that line, not a positive finding by any carrier or broker. It is labelled as such in §7.