Position paper · Strategy · Engineering & Architecture

The AI-augmented engineering firm.

Engineering is about to be disrupted. Not by the firms already in it.

Two-panel illustration of the same modern engineering office. Left, headed Today's Engineering: dozens of engineers packed at CAD workstations behind towering stacks of printed check sets, a queue of people waiting at one senior reviewer's desk, and a wall calendar with weeks crossed off. Right, headed AI-Augmented Engineering: the same room nearly empty, two engineers at a single desk reviewing a complete AI-produced drawing set on three monitors — one circling a flagged exception in red on a tablet, the other pressing a professional engineer's seal onto a sheet.
The same office, before and after. Not two eras and not two companies — one engineering organization, reorganized. The left panel is current practice: modern tools, but headcount scaled to production and a single sequential review gate. The right panel is the operating model this paper argues for: the drawing set arrives finished, and the licensed professionals spend their hours on the two things only they can do — judging the exceptions and carrying the seal. Illustration generated with OpenAI’s image model to an Oxvelo art direction; it depicts the argument, not a measured workplace.

Every serious attempt to bring AI to engineering has tried to build a better drafting tool. Oxvelo's thesis is that the tool was never the bottleneck — the enterprise around it was — and that the engineer's job is no longer to produce the design but to encode the judgment, review the output, and carry the accountability. This paper makes that case, states the seven strongest objections to it at full strength, answers them with the evidence, and recommends how Oxvelo should actually enter: not by building a firm from zero, and not by selling software to firms, but by a staged structure that has already been proven in law, accounting, and medicine.

The premise, stated plainly. This paper takes as its premise that AI is the most disruptive force engineering has faced — more consequential than CAD, and arriving on a far shorter timeline. That is a conviction, not a finding, and it is marked as one. What follows is not an argument about whether the disruption is real. It is an argument about where it lands — on the firm rather than the drawing — and why that is precisely what the incumbents cannot capture.

What this is, and is not. A position paper: a thesis argued from 85 published sources, with original analysis — the leverage model, the automation-versus-liability framework, and a systematic search establishing that nobody has asked clients this question. It presents no new survey data. The study that would supply it is pre-registered and open.

Oxvelo Research · Founding position paper
Published by Oxvelo LLC, Baldwin, New York · oxvelo.com
Powered by Oxvelo — AI-Augmented FinTech. Intelligence that compounds.
Compliance Platinum · 96–100 Oxvelo Certified

Executive summary

Engineering does not have an AI problem. It has an operations problem. The industry absorbed CAD, BIM and a decade of construction software, and sector productivity fell. US construction value added per worker was about 40% lower in 2020 than in 1970. Tools sold into an unchanged operating model get absorbed by it.

The economics are in the unsealed work. The median architecture and engineering firm bills 58.9% of the hours it pays for and carries $1.61 of overhead on every billable dollar. Proposals, estimating, contract review, scheduling, submittals, RFIs and document control require no professional seal — and consume licensed professionals' hours.

The engineer's job inverts. Not to produce the design, but to encode the judgment, review the output and carry the accountability. NCEES rewrote the definition of responsible charge in August 2025 — from "direct control and personal supervision" to exercising "full professional knowledge of and control over work." The new test asks what a licensee knows and controls, not who they supervised.

Subject-matter expertise becomes the scarce input, not the surplus one. AECBench found model accuracy on building-code tables rising from 45.27% to 98.94% purely by restructuring how the tables were presented. The constraint is the preparation of knowledge, which only senior practitioners can do.

The incumbents structurally cannot build this. The standard of care is defined as what practitioners ordinarily do, making innovation asymmetrically risky; and a firm billing time-and-materials that halves its hours halves its revenue. That is why the disruption arrives from outside — and why the recommended entry is a staged structure, partnering with a licensed firm before owning one, exactly as AI-native practices in law, accounting and medicine have done.

This summary is translated. The full analysis below remains in English: it quotes statutes, board rules and contract language verbatim, and translating quoted law would risk changing its meaning. Citations, figures and sources are language-independent.

Terms with a dotted underline carry a definition — hover or tap for a short explanation, then open a deep dive. All 90 are collected in Appendix A.

01 — The premise

Engineering does not have an AI problem. It has an operations problem.

The instinct, when a technologist looks at an engineering firm, is to look at the calculations. That is the visible, technical, impressive part — and it is the wrong target. It is the part of the business with the least slack in it, the highest liability, and the smallest share of the payroll.

The rest of the firm is where the money leaks. Proposals. Estimating. Contract review. Scheduling. Submittal logs. RFIs. Meeting minutes. Specification cross-checks. Document control. Invoicing. Collections. Closeout. None of this requires a professional seal. All of it consumes licensed professionals' hours. And it is exactly the work that a coordinated set of AI agents is now good at.

The industry-level evidence that something is structurally wrong here is not subtle, and it is not new. In McKinsey's words: "Construction productivity improved by only 10 percent, or 0.4 percent annually, from 2000 to 2022, compared with 90 percent, or 3.0 percent annually, in manufacturing."2 In the United States the picture is worse than stagnation: Goolsbee and Syverson found that value added per construction worker was about 40% lower in 2020 than in 1970, a decline large enough to measurably drag on aggregate US productivity growth.3

Figure 1 · Two sectors, twenty-two years
Cumulative labor-productivity change, indexed to 100 in 2000. Construction gained about 10% over the full period; manufacturing gained about 90%. The curves between the endpoints are drawn from the published compound annual rates (0.4% and 3.0%) — the endpoints themselves are McKinsey's figures, the path between them is modeled.
Manufacturing Construction
Source: McKinsey & Company, How agentic AI is transforming the AEC industry (July 2026).2 Corroborating the same direction: McKinsey Global Institute, Reinventing Construction (2017), which put construction productivity growth at ~1%/yr over the prior two decades against 2.8% for the world economy, and valued closing the gap at $1.6 trillion a year.1

Two decades of CAD, BIM, cloud document management, and project-management SaaS were sold into this industry over exactly the window in the chart above. The line did not move. That is the single most important fact in this paper, and I will come back to it twice — once as the argument for why the opportunity is real, and once as the strongest objection to Oxvelo entering at all.

02 — Where the margin actually is

Two out of five paid hours are already not billable.

Zoom from the sector to the firm. Deltek's Clarity study is the industry's longest-running financial benchmark of architecture and engineering firms; the 47th edition surveyed roughly 900 of them.4 Its numbers describe a business model under compression.

58.9%
Median utilization — the share of paid staff hours that are billable
▼ 2.2 pts from 61.1%
161.3%
Median overhead rate — every $1.00 of direct labor carries $1.61 of overhead
▲ a 10-year high
16.7%
Operating profit on net revenue
▼ 4.7 pts from 21.4%
13.8%
Annual staff turnover

Read those together and the arithmetic is stark. Roughly 41% of paid professional hours are already non-billable before a single overhead-department salary is counted — and then every remaining billable dollar of labor drags $1.61 of overhead behind it. Meanwhile operating profit fell nearly five points in a single year off a ten-year high.

Figure 2 · The engineering payroll hour
Median utilization across ~900 A&E firms, current year versus prior year. The non-billable share is not waste by definition — it contains business development, QA/QC, training, and management — but it is the pool that an AI workforce is aimed at.
Billable hours Non-billable hours
Source: Deltek Clarity 47th Annual A&E Industry Study (2026), ~900 firms.4 Utilization and overhead figures as reported by Full Sail Partners, a Deltek partner, from the full study; Deltek's own public page confirms the 16.7% operating profit, 13.8% turnover, and 70% AI-adoption headlines.5 Note: no credible time-and-motion study measures how architects and engineers split their day between design and administration — utilization rate is the closest published proxy, and it is what is charted here.

A number this paper deliberately does not use

The most-quoted statistic in AEC-technology marketing is that "construction professionals spend 35% of their time — over 14 hours a week — on non-productive activities," at a cost of $177 billion in US labor annually.18 It is real and well-sourced, but the surveyed population was 49% general contractors and 36% specialty trades — not design professionals. Applying it to engineers and architects would be the kind of small dishonesty that makes a reader stop trusting everything else in the document. The 58.9% utilization figure above is the defensible one.

03 — The argument

The firm, not the task, is the unit of transformation.

Engineering companies are staffed almost entirely by engineers, and engineers think like engineers. That is a strength in the technical work and a blind spot everywhere else. It produces firms that will spend two years optimizing a structural analysis workflow and zero years asking why it takes eleven days to turn a signed contract into a project workspace.

The alternative framing is to organize the opportunity by business function rather than by engineering discipline. Every engineering business contains business development, estimating, contracts, project management, technical production, document control, construction administration, finance, people operations, compliance, and knowledge management. Each of these has an AI counterpart. Together they are not a feature set — they are a workforce.

Figure 3 · The AI-native firm, mapped by function
Digital colleagues observe, retrieve, organize, analyze, draft, recommend, coordinate, and escalate. Humans retain decision rights, professional responsibility, approvals, client accountability, and judgment. The control points marked in orange are where a licensed human must act — they are load-bearing, not decorative.
Framework diagram — Oxvelo. The human control points are drawn from the professional-practice constraints documented in §7: NCEES Model Rules §240.20 on sealing and responsible charge,20 NSPE Board of Ethical Review Case 24-2,21 and Texas Board PAO-71.22

Underneath the workforce sit two more layers, and they are where the durable advantage actually lives:

Institutional intelligence. Every project, document, standard, decision, lesson learned, estimate, and historical outcome becomes searchable and reusable. Today this asset exists in every firm and is worth almost nothing, because it is scattered across file servers, inboxes, and the memory of a principal who retires in four years. In 2022 an estimated 184,175 engineers retired or left the field, 85,175 of them from civil, mechanical, and electrical disciplines.15 Every one of them took an unindexed corpus with them.

Workflow orchestration. Once a project is won, the system initiates project setup, schedules, deliverable registers, budget tracking, compliance logs, document structures, and draft work products — escalating to humans at defined control points rather than waiting to be asked.

The product is not engineering calculations or drawings. The product is the operation of an entire engineering enterprise in a fundamentally new way.The founding thesis, stated plainly
04 — The beachhead

Start where no one has to seal anything.

Some engineering work is now so repeatable that the intelligence has already been embedded — in standards, templates, prior designs, and established procedures. What remains is execution, checking, documentation, and sign-off. That is a very different problem from "design a bridge," and it is the correct place to begin.

The right way to sequence entry is to plot every function on two axes at once: how automatable it is, and how much professional liability it carries. Those two axes are not independent — and the safe, valuable territory is the top-left, not the top-right.

Figure 4 · Automation potential versus liability exposure
The horizontal axis is how much of the function AI can plausibly carry today. The vertical axis is how much professional liability attaches to getting it wrong. Blue functions require no professional seal; orange functions do, or feed directly into sealed work. Oxvelo's entry sequence runs left-to-right through the blue band before it touches the orange one.
No professional seal required Sealed work / responsible charge attaches
This is a framework, not survey data. Positions are Oxvelo's editorial judgment, informed by the licensure constraints in §7 and by the one published quantitative anchor available: McKinsey states that "AI has the potential to automate 50 percent of nonphysical work in the architecture and engineering sectors."2 Separately — and this is a distinct claim, not a valuation of the first — the McKinsey Global Institute projects the AEC industry could unlock roughly $228 billion in annual US value by 2030 "through AI and other automation technologies."2 The two are often quoted as one figure; they are not. This chart should be read as a stated hypothesis to be tested, not as a measurement.

Note what this rules out. It rules out an AI that independently approves submittals. It rules out an AI in responsible charge of a design. It rules out selling "the model sealed it." What it rules in is a firm where the proposal was assembled in an afternoon instead of two weeks, the contract deviations were flagged against a company standard before a principal ever opened the file, the deliverable register built itself at kickoff, and the RFI log answered 60% of its own questions by citing the spec section — all of it auditable, all of it escalating to a named licensed human at the points where the law requires one.

05 — The core claim

We no longer need engineers to produce the design. We need them to QA it.

This is the sharpest form of the thesis, and it deserves to be stated without hedging: a design team of a hundred engineers is an artifact of a production constraint that no longer binds. Two engineers, reviewing and taking responsibility for AI-produced output, can carry the production load that a large team once required. The financial, throughput, accuracy, and cycle-time consequences of that are not incremental. They are a different business.

The reason this is not reckless is that the profession already has a legal category for exactly this arrangement, and has had one for decades. It is called review.

Responsible charge never required authorship

Read the rules again with this claim in mind — and read the current ones, because NCEES rewrote the key definition in August 2025 in a way that has gone almost entirely unremarked. The old Model Law defined responsible charge as "direct control and personal supervision of engineering or surveying work." The August 2025 Model Law replaced that with a functional test: responsible charge now means "to exercise full professional knowledge of and control over work," elaborated as four duties — the authority to review, change, reject or approve the work; personal awareness of its scope; the capability to answer questions about it; and acceptance of full responsibility for it.20

That change matters more than any commentary I have seen acknowledges. The old test asked whether you supervised the person doing the work. The new one asks whether you know, control, can defend, and own the work itself. A licensed engineer reviewing AI output can satisfy every one of those four duties. A licensed engineer cannot "personally supervise" a model in any meaningful sense — under the superseded language, the argument was awkward. Under the current language, it is clean.

The rest of the framework points the same way. NCEES Model Rules §240.20 D requires that the seal be placed on work "only when it was under the licensee's responsible charge."20b New York's rule defines unprofessional conduct as sealing documents for which the services "have not been performed by, or thoroughly reviewed by, the licensee" — the disjunction is in the regulation.59 NSPE's Code expressly contemplates an engineer accepting responsibility for coordination of an entire project and sealing its documents, provided each technical segment is signed and sealed by the qualified engineer who prepared it.21b

Every large engineering firm on earth already runs on this. A principal seals work produced by junior staff they did not personally draft. The senior engineer's value was never the drafting — it was the judgment applied to the drafting. AI does not change the legal structure of that relationship at all. It changes the marginal cost of the thing being reviewed from an engineer's salary to a compute bill.

NSPE's Board of Ethical Review confirmed the boundary precisely in Case 24-2. An engineer used AI to draft a groundwater report — which he verified thoroughly, cross-checking against journal articles — and to draft design documents, which he reviewed only cursorily; the client found misaligned dimensions and omitted safety features. The Board's finding was not that using AI was wrong: "The use of AI-assisted drafting tools by Engineer A was not unethical per se. However, Engineer A's misuse of the tool, by failing to maintain Responsible Charge over the AI tool and its output before sealing the document… was unethical."21 The violation was insufficient review, not delegation. That is a green light with a condition attached, and the condition is the entire product.

Two details of that case are worth carrying forward rather than glossing. First, the Board judged the report use "partly ethical, and partly unethical" — thorough verification satisfied the competence obligation, but the engineer separately failed to obtain client permission before disclosing private information to the tool and failed to document required technical citations. Verification alone was not sufficient. Second, the Board held there was no obligation to disclose AI use to the client, while noting that "ethical principles favor transparency when AI plays a substantial role."21 Data handling and citation discipline are therefore product requirements alongside review — which is exactly what the provenance layer in Figure 6 exists to satisfy.

The senior engineer's value was never the drafting. It was the judgment applied to the drafting. AI does not change the legal structure of that relationship — it changes the marginal cost of the thing being reviewed.

The arithmetic of leverage

The economics follow mechanically. US architecture firms generate average net billings of $143,000 per employee — up from $86,000 in 20117 — against a total labor cost of $119,511 per employee at A&E firms.45 That is the entire industry's problem in two numbers: revenue per head and cost per head are almost the same number, and the 161.3% overhead rate sits between them. (The $143,000 is a mean; AIA does not publish a median, and in a distribution this skewed the median is likely lower — which makes the gap worse, not better.)

Now introduce a leverage ratio — call it L, the number of engineer-equivalents of production output that one licensed professional can review and take responsibility for. Today, structurally, L ≈ 1. Every engineer produces roughly their own work. Move L and the revenue-per-professional line moves with it, linearly, while headcount cost does not.

Figure 5 · The leverage curve
Net revenue per licensed professional as a function of the leverage ratio L — how many engineer-equivalents of production output one professional reviews and seals. The baseline point is published data; the curve is arithmetic; the ceiling is unknown and finding it is the business.
Baseline (published): $143,000 average net billings per employee, AIA 2024 Firm Survey Report (2023 data) — a mean, not a median; AIA publishes no median.7 Run this on your own numbers → The curve is arithmetic, not a forecast: revenue per professional = baseline × L, holding price per deliverable constant. Two things bend it in practice. First, competitive pricing: if the savings are passed to clients, revenue per professional rises more slowly — but margin still rises, because cost per deliverable falls faster than price. Second and more importantly, this only converts to profit on fixed-fee or lump-sum work. On time-and-materials contracts, reducing hours reduces revenue — which is exactly the monetization constraint ENR identified in AECOM's $390M acquisition of Consigli.46

The "two engineers instead of a hundred" claim sits at the far right of that curve — L ≈ 50 for the production function specifically. Two clarifications make it defensible rather than merely bold:

  • It applies to production, not to the whole firm. A hundred-person engineering firm is not a hundred people drafting. It contains project managers, construction administration staff, principals, finance, and business development. The claim is that the design production function collapses to a small review team — while the AI workforce in Figure 3 simultaneously compresses the other functions on their own curves. Those are two different arguments and both are in this paper.
  • It applies first to templatized, repeatable work. The blue band in Figure 4 — where the intelligence is already embedded in standards, prior designs, and established procedures, and what remains is execution, checking, and documentation. On a first-of-its-kind structure, L stays near 1 for a long time, and it should.

So why is Oxvelo hiring engineers at all?

The obvious objection to everything above is that it appears to contradict the plan. If engineers are not needed to produce the design, why assemble a team of engineers?

Because the ceiling on what an AI system produces is set by the quality of the subject-matter expertise embedded in it — and that expertise has to come from somewhere. A general-purpose model dropped into an engineering firm produces generic output. The same model, working against a curated corpus of that firm's standards, prior designs, acceptance criteria, checking procedures, failure history, and hard-won conventions, produces the firm's work. The difference between those two systems is not the model. It is the engineers who built the second one.

This is exactly what the capability research implies, read forward instead of backward. AECBench found LLMs weakest at interpreting building-code tables and at complex reasoning38 — which means the fix is retrieval against curated, SME-validated sources and deterministic checks with SME-defined acceptance criteria. Curating those sources and defining those criteria is senior engineering work. Hallucination is a structural product of training objectives that reward guessing;39 the counter is a provenance layer whose ground truth an SME had to establish. Every layer in Figure 6 that makes review cheap is itself a deposit of expert judgment made once and spent many times.

Eudia's leadership said the same thing when they bought a 300-person legal services firm rather than hiring engineers to build around a model: "building truly transformative systems requires deep human expertise embedded within the technology itself."42 cove.tool's AI architecture practice was built on a claimed decade of R&D by practicing architects, not by an AI team that read a code book.45

So the headcount does not go to zero. It changes shape. The engineer's job splits into three, and every one of them is a higher-value job than drafting:

Role 1
Knowledge architect
Encodes the firm's judgment into the system: standards, templates, prescribed checks, acceptance criteria, escalation rules, and what "good" looks like. This is the compounding asset. It is done once and spent on every project afterward.
Role 2
Reviewer in responsible charge
Adjudicates the ranked exception set, exercises professional judgment, seals the work, and carries the accountability. Legally irreplaceable, and the only role the statutes actually name.
Role 3
Creative on the novel
Handles the first-of-its-kind problem, where no template exists and no precedent applies. Where L stays near 1 — and where an engineer's time is worth the most it has ever been worth.

Notice what this does to the hiring profile. It does not call for a large team of production drafters. It calls for a small number of unusually senior people — the ones with the judgment worth encoding. And that is precisely the resource the market is shortest of. ACEC's 2025 workforce survey found that only 34% of engineering students feel well-prepared with the skills most important to succeed in their chosen profession, while mid-career and older professionals estimate that just 14% of their young colleagues are well-prepared on those same skills.15 Set that against 184,175 engineers leaving the workforce in a single year.15 The scarce asset in this industry is not production capacity. It is senior judgment — and senior judgment is the one input AI cannot supply and the one an AI-native firm can leverage furthest.

Two caveats on those figures, since they carry weight here. The 34% comes from a 195-student convenience sample recruited via three universities and ACEC scholarship applicants, not a representative national panel. And the 14% is the estimate of mid-career and senior practitioners — it is a perception of preparedness, not a measurement of it. Both point the same direction; neither is precise.

AI does not replace the subject-matter expert. It multiplies one. The question is no longer how many engineers you can afford — it is how good the two you have are.

The magnitude claims, and what actually supports them

The strongest available quantitative anchor is McKinsey's statement that "AI has the potential to automate 50 percent of nonphysical work in the architecture and engineering sectors."2 Note what that means in leverage terms: it is L ≈ 2 across the whole enterprise — a doubling, not a fiftyfold. The high-L claim is a claim about a narrow, deliberately chosen slice, and it should always be stated that way. (A caution on a figure you will see quoted alongside it: MGI's $228 billion is the projected annual US value for the entire AEC industry by 2030 from "AI and other automation technologies" — it is not the dollar value of the 50%, and it is not AI alone.2)

The most direct operating evidence comes from cove.tool, which launched Cove Architecture in March 2025 as a full-service AI-powered architecture practice after a claimed $25 million and a decade of R&D. Its reported results on a 15-unit Atlanta housing project:

60%
Reduction in design timelines
95%
Cost-estimate accuracy
40%
Cut in design iteration expense
15 days
Project completion, start to finish

Source: cove.tool, reported by Bisnow and Multi-Housing News (2025).45 Handle with care: these are company-reported figures on a single project from a firm roughly eighteen months old, and neither account clarifies whether Cove Architecture is a registered licensed firm or how it was staffed. They are the best available signal that the magnitude is real, and they are not proof.

The adjacent professions read the same way, and Crosby is the cleanest instance. It reviewed more than 1,000 commercial contracts within roughly five months of launch with about 19 people; by its March 2026 Series B the figure was reportedly around 13,000 contracts with roughly 30 lawyers.40 That is the leverage curve moving in real time, in a licensed profession, with the firm explicitly stating: "We stand behind our work and take liability for it."40 The UK regulator authorized Garfield.Law on the condition that named regulated solicitors remain ultimately accountable for all system outputs.41 That is the QA model, licensed and operating, in a profession with the same liability structure as engineering.

The hard part: making review scale

Here is the honest problem with the claim, and it is a real one. If two engineers must genuinely review the output of a system producing at the rate of a hundred, review throughput becomes the binding constraint and rubber-stamping becomes the failure mode. A reviewer who cannot actually review is not in responsible charge, whatever the org chart says — and in New York, that reviewer owes a retained written evaluation for six years on work prepared by others.59

So the leverage ratio is not won by making the AI better at producing. It is won by making the output cheaper to verify. That is a different engineering problem, it is the actual product, and it is what Figure 6 describes.

Figure 6 · How review scales — the verification architecture
Every layer below reduces the amount of output a human must read closely in order to be genuinely in responsible charge of all of it. L is a function of this stack, not of model quality.
Architecture — Oxvelo. The deterministic-check layer exists because AECBench found LLMs weakest at exactly complex calculation and code-table interpretation;38 the provenance layer exists because hallucination is a structural property of the training objective, not a patchable defect;39 the assurance layer exists because automation bias is measurable, survives expertise, and "cannot be prevented by training or instructions."7c7b Each layer answers a named, cited failure mode. Note also AECBench's encoding finding: restructuring building-code tables into natural language lifted model accuracy on them from 45.27% to 98.94%38 — which means the retrieval corpus is not a passive archive but an engineered artifact, and preparing it is the knowledge architect's core job.

Stated as a design principle: an engineer should never review a drawing. An engineer should review a small, ranked set of exceptions, each one carrying its own provenance, against a package where everything unexceptional has already been proven by deterministic check. Get that right and L climbs. Get it wrong and you have built a machine that manufactures liability at scale.

06 — Why now, and why this market

A fragmented industry, an adoption gap, and a consolidation window.

Three market conditions have to hold simultaneously for this to be a business rather than a thesis. All three currently hold.

The industry is extraordinarily fragmented

Start with the concentration figure, because it is the cleanest. ENR's 2026 Top 500 Design Firms — the five hundred largest design practices in the country — booked $136.3 billion in US domestic design revenue in fiscal 2025.12 Census's Quarterly Services Survey puts total 2025 revenue for NAICS 5413, architectural, engineering and related services, at roughly $494.5 billion.12b The five hundred biggest firms in America therefore account for something like 28% of the industry. The other 72% is spread across tens of thousands of firms — and within the Top 500 itself, the top ten alone take 32.2% of the revenue.12

Look inside the long tail and it is smaller than most people assume. AIA's 2024 Firm Survey finds that 28% of US architecture firms are sole practitioners, 32% have two to four employees, and about 75% have fewer than ten.7

Figure 7 · Share of firms versus share of billings, by firm size
The last time a professional body published both halves of this distribution in one table. 81% of architecture firms had fewer than ten staff and earned 21% of the billings; the 1% of firms with 100 or more staff earned 27%. AIA's 2024 survey confirms the left-hand shape is essentially unchanged — and that billings have concentrated further since.
Share of firms Share of billings
Source: AIA, 2012 Survey Report on Firm Characteristics, Figure 1.4 (2011 data) — the most recent published table carrying both axes.7d Why 2011 data: AIA's current Firm Survey publishes the firm-size distribution openly but keeps the billings-by-size breakdown inside the paid report. The 2024 edition confirms the distribution has held (75% of firms under ten employees) and that concentration has increased: between 2015 and 2023 the billings share generated by small firms fell by half, the midsize share fell 40%, and the large-firm share rose 40%.7 Read the chart as a conservative floor on today's concentration, not a current snapshot.

Fragmentation matters for a specific reason: a firm with eleven people cannot build an AI operating layer, and cannot buy one that fits, but it can be joined to one. And the concentration trend says the small firms are already losing ground — the operating-model gap is the mechanism. And these firms are already changing hands. 2025 was the first year on record with more than 500 completed domestic A/E transactions; 72% of those deals involved firms under $10 million in revenue, and private-equity-backed acquirers or recapitalizations now account for more than half of all design and environmental consulting deals.13 A second tracker, using a narrower sector definition, counts 361 transactions in 2025 with PE at 38.3% of deal volume — up from 22.3% in 2018.14 The definitions differ; the direction does not.

The talent supply is structurally short

ACEC's workforce research puts the 2022 net engineering shortage at roughly 18,000 engineers, with 8,411 of that in the core civil, mechanical, and electrical disciplines.15 Engineering degrees awarded peaked near 214,000 in 2019 and have fallen by more than 10,000 since.15 As of Q2 2026, 33% of engineering firms had turned down work in the prior six months because they were short-staffed — down from 51% in Q4 2024, but still a third of the industry declining revenue it cannot staff.16 Median backlog stands at 11 months, with 49% of firms reporting a year or more.

This is the demand-side case, and it is stronger than the cost-side case. A firm that turns down work is not looking to cut headcount — it is looking for throughput. That reframes the entire sale.

Adoption is real but shallow — and the published numbers are a mess

Here I have to be careful, because this is where AEC-technology writing is least honest. Published AI adoption rates for this industry range from 6% to 98% depending entirely on what was asked and of whom.

Figure 8 · "AI adoption in AEC" is not one number
Seven published figures from 2024–2026, each measuring something different. Read the question, not the percentage.
Sources, top to bottom: AIA/Deltek/ConstructConnect (2025) — note a further 20% of firms report implementation in progress;8 same;8 Bluebeam via ASCE (2025);9 AIA 2024 Firm Survey Report;7 Deltek Clarity 46th;6 Deltek Clarity 47th;4 Autodesk State of Design & Make: AI Pulse.10 Two populations to watch especially closely. Bluebeam's 27% is not "AEC professionals" — it surveyed 1,000+ technology decision-makers at manager level or above in five countries only (US, UK, France, Germany, Australia).9b2 Autodesk's 98% spans all three Design & Make sectors — architecture/engineering/construction, product design and manufacturing, and media and entertainment — so it is not an AEC-specific number.10 Both are frequently quoted as though they were.

The most defensible signal is not any single figure but the one series that asked the same question of the same industry three years running. Deltek's A&E study went from 38% to 53% to 70%.46

Figure 9 · The one trend line that controls for the question
Share of architecture and engineering firms reporting use of AI and machine learning, Deltek Clarity A&E Industry Study, three consecutive annual editions of the same panel.
Source: Deltek Clarity A&E Industry Study.46 The 38% is reported inside the 46th study as its prior-year comparison ("up from 38% last year"), not as a headline of the 45th study; the sample base also grows across the three years (650+ → 692 → 896 firms), so this is a repeated cross-section, not a fixed panel. Read the level correctly too: "using AI" is a low bar — any use of AI or ML anywhere in the firm. Deltek's own 47th-edition companion figure is the telling one: only 38% report measurable positive business impact from it.4 The door is open. Nobody has rebuilt the operating model.

That gap — 70% of firms touching AI but only 38% reporting measurable business impact,4 while 52% of AEC respondents still use paper during the design phase and 43% still depend on physical signatures9 — is the whole opportunity. Adoption is a mile wide and an inch deep. Nobody has rebuilt the enterprise.

07 — The case against

Seven objections, at full strength.

A position paper that only argues one side is marketing. Below are the seven strongest arguments against Oxvelo doing this — stated the way a hostile, well-informed reader would state them, then answered. Three of them are substantially correct, and they change the strategy rather than defeat it.

Objection 1 · Licensure

An AI cannot be in responsible charge of engineering work, cannot seal a drawing, and cannot hold professional accountability. Everything valuable in an engineering firm terminates in a licensed human's signature. You are automating the parts that don't matter.

Response

The first half is simply true, and Oxvelo should say so louder than its critics do. NCEES Model Rules §240.20 D require that a licensee's seal and signature be placed on work "only when it was under the licensee's responsible charge,"20b with responsible charge defined since August 2025 as exercising "full professional knowledge of and control over work."20 New York goes further: sealing documents the licensee did not personally prepare or "thoroughly review" is defined as unprofessional conduct, and a licensee reviewing another's work must retain a written evaluation for six years.59

The second half — "the parts that don't matter" — is where the objection fails, and Figure 2 is the answer. If 41% of paid hours are already non-billable and every billable labor dollar carries $1.61 of overhead, then the unsealed work is the economics of the firm. The seal is the product's warranty. It is not the product's cost structure.

The regulators, notably, have already answered the narrow question. The Texas Board of Professional Engineers and Land Surveyors issued Policy Advisory Opinion PAO-71 in November 2024 finding that AI software is not prohibited, that "licensees are ultimately responsible for any work product they sign and seal," and that existing rules already govern AI adequately — declining to write AI-specific rules.22 NSPE's Board of Ethical Review reached the parallel conclusion in Case 24-2: using an AI drafting tool "was not unethical per se"; failing to maintain responsible charge over its output before sealing was.21 Florida's board put it in one sentence: "AI can assist your work, but it cannot replace your professional judgment or accountability."23

What this genuinely constrains: Oxvelo can never market autonomous sealed output, and every workflow that touches a sealed deliverable needs an auditable human-review record by design — in New York, a retained written evaluation. That is a product requirement, and it should be built in from day one rather than retrofitted after the first claim.
Objection 2 · Ownership

You are a non-licensee in New York. Under New York law you cannot own an engineering firm, cannot be its President, cannot be its Chairman, and cannot be its CEO. The entity you are proposing to build is illegal in your own state.

Response

This objection is correct, and it is the single most consequential finding in this paper. It does not defeat the strategy. It dictates the corporate structure.

New York's baseline rule for an ordinary professional service corporation practising engineering is that every shareholder, officer, and director must be licensed in New York, and "no other entity or individual except those described in the preceding may practice professional engineering in New York State."25 The one relief valve is the Design Professional Service Corporation under BCL §1503(b-1), and its conditions are tighter than the summary usually given:24

  • Greater than 75% of outstanding shares must be owned by design professionals and an ESOP;
  • Greater than 75% of directors and greater than 75% of officers must be design professionals;
  • The president, the chairperson of the board, and the chief executive officer or officers must all be design professionals;
  • The largest shareholder must be a design professional or a qualifying ESOP;
  • And critically — the sub-25% tranche is available to employee stock ownership plans and to employees of the corporation who are not design professionals. An outside, non-employee investor entity does not qualify for it at all.

So the position is worse than "a non-licensee may hold up to 25%." In New York, an outside holding company has no route to equity in the licensed design entity. Other states are far more permissive: California requires that a licensed engineer be an owner, partner, or officer in charge and bars sole ownership by a non-licensee,27 and Texas's firm-registration rule is silent on ownership entirely, requiring only that at least one full-time active licensee be employed and supervise the licensed work.26

Figure 10 · How much of a licensed design entity a non-licensee may own
Three states, three regimes. New York is the strictest in the country; it is also where Oxvelo is domiciled. The bar is a simplification of statute — the footnotes carry the real conditions, and every one of them is load-bearing.
Sources: NY Business Corporation Law §1503(b-1) and NYSED Office of the Professions;2425 Cal. Bus. & Prof. Code §6738;27 22 Tex. Admin. Code §137.77.26 Read the New York bar as an outer bound, not an entitlement. The sub-25% tranche is reserved by statute for ESOPs and for employees of the corporation who are not design professionals — an outside, non-employee investor entity does not qualify, so the practical figure for a holding company is zero. California: §6738 reaches only the civil, electrical and mechanical branches; the separate prohibition on engineering LLCs comes from Cal. Corp. Code §17701.04(e), not §6738.27b Texas: silent on ownership, but at least one full-time active licensee must be employed. Nothing in this paper is legal advice; confirm all of it with counsel before committing a structure.
Objection 2, continued · The structural answer

Three separate regulated professions have independently converged on the same solution to this exact problem, which is strong evidence it is the right one.

In medicine, the corporate practice of medicine doctrine bars companies from employing physicians to deliver care. The universal workaround is the MSO/PC structure: a physician-owned professional corporation employs the clinicians and controls all clinical decisions; a management services organization — the technology company — provides platform, scheduling, billing, HR, and marketing under a management services agreement.55 In law, Crosby — which raised a $60M Series B led by Lux and Index in March 2026, with Sequoia, Elad Gil, 01 Advisors and Bain Capital Ventures participating — runs precisely this shape: "Crosby Legal PLLC is a law firm……powered by tech from Crosby Legal, Inc."40 In accounting, PE platforms like Alpine's Ascend split each acquired firm into an attest entity and a tax/advisory entity — the "alternative practice structure" — because CPA-majority ownership is required for attest work.44

Engineering's constraints map onto healthcare's almost one-to-one: state-by-state licensed entities, a named individual who must hold professional authority (the engineer in responsible charge, the physician), and ownership restrictions that vary from strict to permissive by state.

One detail that quietly breaks business models: in New York, MSO compensation in the healthcare analogue must be a flat fee, time-based, or otherwise "reasonably related to the cost of the services" and "unrelated, directly or indirectly, to the dollar amounts billed and collected" — a percentage-of-revenue management fee is treated as unlawful fee-splitting, and the Medical Society of the State of New York has warned members explicitly on this point.56 Whether the same characterization reaches an engineering MSO is an open question that must be answered by counsel before a revenue-share model is built, not after.
Figure 11 · The two-entity structure
The technology company owns the platform, the data, and the enterprise value. The licensed design entity owns the professional judgment, the seal, and the liability. A management services agreement binds them. This is not a loophole — it is the standard, regulator-tested structure in three other licensed professions.
Structure diagram — Oxvelo, modeled on the MSO/PC arrangement documented in healthcare5556 and observed in law (Crosby)40 and accounting (Ascend's alternative practice structure).44 The known failure mode is the "strawman": a nominal licensee figurehead with real operational control sitting on the corporate side. The documented red flags are all client revenue flowing directly to the management entity, and the management entity hiring or firing licensed professionals without professional input.56 If the engineer in responsible charge is not genuinely in responsible charge, the structure is not a structure — it is an enforcement action.
Objection 3 · Model capability

Language models are demonstrably bad at exactly the reasoning engineering requires. They cannot reliably read a code table, they fail at multi-step calculation, and hallucination is a structural property of how they are trained — not a bug awaiting a patch.

Response

Accurate, and worth quantifying rather than hand-waving. AECBench — a peer-reviewed benchmark of nine LLMs against 4,800 expert-validated questions across a five-level cognitive framework in architecture, engineering, and construction — found that models handle foundational knowledge adequately but show significant deficits in interpreting knowledge from tables in building codes, in complex reasoning and calculation, and in generating domain-specific documents, with performance declining steadily as cognitive demand rises.38 And hallucination is not incidental: OpenAI and Georgia Tech researchers argue models hallucinate "because the training and evaluation procedures reward guessing over acknowledging uncertainty" — a statistical pressure, not a defect.39

But the AECBench paper contains a finding its own summary buries, and it is the single most useful result in this entire literature for Oxvelo's purposes. The building-code-table failure is an encoding problem, not a knowledge problem. When the researchers converted code tables into natural-language descriptions, model accuracy on those questions rose from 45.27% to 98.94%. Rendering the same tables as HTML lifted accuracy from 40.65% to 87.27%.38

Read that again. The same model, on the same code provisions, goes from failing more than half the time to answering almost perfectly — purely by changing how the source material is presented to it. That is not a capability ceiling. It is a data-engineering task, and it is precisely the kind of work an SME-built corpus does. It is also the strongest available evidence that the constraint in this domain is the preparation of knowledge rather than the reasoning over it.

Notice too what AECBench measured: unaided models answering engineering questions from parametric memory. That is not the proposed architecture. The proposed architecture is retrieval against the firm's own standards and prior work — restructured for machine legibility, as the 45%→99% result demands — with deterministic calculation in code rather than in the model, structured templates carrying prescribed checks, and mandatory escalation to a licensed reviewer. A system that looks up the code table and shows the citation is a different failure surface from a system asked to recall it.

What this genuinely constrains: it is a hard argument against the orange quadrant in Figure 4, and it should be treated as one. Sealed design production is not the beachhead. Building code interpretation in particular — the single weakest measured capability — should be scoped as retrieval-and-cite, never as generate-and-trust.
Objection 4 · Review capacity

"Two engineers instead of a hundred" does not eliminate the bottleneck — it relocates it. If AI produces at a hundred engineers' rate, two humans must read at a hundred engineers' rate to be genuinely in responsible charge. They cannot. So either the leverage is fictional, or the review is, and if it is the review then you have built a plan-stamping operation with better branding.

Response

This is the correct objection to the boldest claim in the paper, and the answer is not to soften the claim — it is to say precisely what has to be true for it to hold.

The objection assumes review effort scales linearly with output volume. It does not have to. Review effort scales with the number of things a human must actually adjudicate, and that number is a design variable. A package where every dimension has been checked deterministically against the model, every code citation carries a retrievable provenance link to the clause it came from, and every deviation from firm standard is surfaced as a ranked exception is a fundamentally cheaper object to review than a stack of drawings — even though it represents the same volume of output. That is the entire content of Figure 6, and it is why the engineering effort goes into verification tooling rather than into generation quality.

The regulatory frame supports this reading, and after August 2025 it supports it more directly than before. NCEES's current four duties are the authority to review, change, reject or approve the work; personal awareness of its scope; the capability to answer questions about it; and acceptance of full responsibility.20 None of the four is "read every line." A principal sealing a junior's work has never re-run every calculation; they have applied judgment to a reviewable artifact. The question a board would ask is whether you exercised professional judgment over this work and can demonstrate it — and an exception-based review with a complete audit trail answers that better than the status quo does, not worse.

Where the objection lands, and it lands hard: automation bias is real, it is measured, and this model concentrates it. In a controlled experiment with 120 clinicians using electronic prescribing, incorrect decision support raised failure-to-detect errors by 25 to 33 percentage points, and between 51.7% and 65.8% of participants accepted the automated system's false-positive advice.7b Parasuraman and Manzey's review adds the part that should worry anyone building this: automation bias "occurs in both naive and expert participants, cannot be prevented by training or instructions."7c That literature is clinical and aviation, not design review — the transfer is an analogy and should be labelled one — but METR's engineers demonstrated the same signature: measured 19% slower, believed 20% faster.34 The mitigations are therefore not optional: blind sampling of "no-exception" output at a fixed rate, tracked reviewer catch rates against seeded defects, hard caps on packages per reviewer per period, and the plan-stamping red line stated in writing — if a reviewer cannot articulate why the design is correct, the package does not ship, regardless of what the system reported. Texas's rule already treats granting others access to your seal as misconduct;26b a reviewer nominally in charge of output they cannot adjudicate is the same offense wearing a different hat.
Objection 5 · The productivity evidence

The claimed gains do not survive measurement. METR ran a randomized controlled trial and found experienced developers were 19% slower with AI tools while believing they were 20% faster. MIT reported that 95% of enterprise GenAI pilots produced no measurable P&L impact. You are selling a productivity story the evidence does not support.

Response

Both studies are real, and both have been materially undercut since — which a reader in August 2026 needs to know, and which most people citing them do not mention.

METR redesigned the experiment and got a different answer. The July 2025 study did find a 19% slowdown across 16 developers and 246 tasks, with a confidence interval of +2% to +39%.34 In February 2026, in a post titled We Are Changing Our Developer Productivity Experiment Design, METR reported that in the original design "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI" — a selection bias that removes exactly the tasks AI would handle best. A later study produced an 18% slowdown for the subset of original developers who took part, with a confidence interval of −38% to +9%, and a 4% slowdown for new developers, at −15% to +9%. Both intervals cross zero. METR's own current position: "we believe it is likely that developers are more sped up from AI tools now — in early 2026 — compared to our estimates from early 2025."35

Two honest notes, because this is where people get sloppy in both directions. METR did not retract or revise the 19% figure — it stands as published; the later numbers come from a separate study. And METR says its own data give only "very weak evidence" for the size of the improvement. The correct reading is not "AI makes developers faster." It is that the headline number is a great deal less settled than the people quoting it at you will suggest.

The MIT "95%" figure is weakly sourced. It comes from an MIT NANDA report built on structured interviews with representatives of 52 organizations and 153 survey responses gathered at four industry conferences, whose own limitations section concedes the figures are "directionally accurate based on individual interviews rather than official company reporting."36 Wharton's Kevin Werbach, on the 95%: "I've read through the document multiple times, and I still can't understand where it comes from." Futuriom's assessment is that the report "lacks thorough sourcing" and is "not a traditional academic research paper," and notes that "some of the authors work at some of these vendors."37

What survives the correction, and should: the self-report gap. METR's developers believed they were 20% faster while measuring slower. That finding is robust across both versions and is a direct warning about automation bias in professional review — the exact risk in a firm where a licensed engineer is asked to check AI output before sealing it. Measure throughput. Never trust the survey.
Objection 6 · The historical record

This industry has absorbed CAD, BIM, cloud collaboration, and a decade of construction SaaS. Productivity went down. Why would AI be different — and why would an AI company succeed where Autodesk, with a twenty-year head start and every firm as a customer, has not?

Response

This is the best objection in the list, and Figure 1 is the evidence for it, not against it. Structural engineers have said it out loud in their own trade press: BIM "is revolutionizing design," modular assembly is doing the same for construction, and yet aggregated industry data shows productivity "in a neutral to negative trend for many years."6b

But read the failure precisely. Every one of those technologies was sold as a tool, to a firm, which then had to reorganize itself to capture the value — and did not. Tools do not change operating models; they get absorbed into existing ones. A firm that buys BIM and keeps the same proposal process, the same document control, the same handoffs, and the same overhead structure gets a better drawing at the same cost.

That is precisely the argument for owning the firm rather than vending to it. Elad Gil put it as cleanly as anyone: "If you own the asset, you can [transform it] much more rapidly than if you're just selling software as a vendor."48 (The brackets are TechCrunch's — its editorial substitution, not Gil's word.)

The reverse trade is also instructive. On 27 November 2025, AECOM announced the acquisition of the Norwegian AI startup Consigli — self-described as "The Autonomous Engineer" — for $390 million. ENR's analysis identified four headwinds: "the internal user base — even at AECOM's scale — may be too small to justify long-term platform investment"; "competing vendors are embedding similar capabilities into off-the-shelf tools, eroding differentiation"; "data readiness remains uneven"; and "the economics of time-and-materials contracts limit how much efficiency can be monetized."46 ENR also characterized "the immediate drop in AECOM's share price" as suggesting investor skepticism — that reading is ENR's, and I could not corroborate the price move independently, so treat it as their assessment rather than established fact.

That fourth headwind is the deepest point in this entire paper: a firm billing hourly has a structural disincentive to reduce hours. An AI-native firm that prices on outcomes does not.

The cautionary case that must be taken seriously: Nines, a radiology AI company, stood up its own teleradiology practice — Nines Radiology — pairing the models with licensed physicians delivering the service, for exactly the reasons argued here. It did not work. Sirona Medical eventually acquired Nines' clinical data pipeline, two FDA-cleared devices, ML engines and key personnel, but explicitly not the practice: "Sirona Medical did not acquire Nines' teleradiology business, Nines Radiology." CEO Cameron Andrews' stated reason — "Sirona isn't in the radiology services business."47 Owning the licensed entity is a heavier, slower, lower-multiple business than software. Anyone choosing it should choose it with open eyes.
Objection 7 · Procurement and insurance

Engineering work is not won on price or efficiency. It is won on qualifications, named credentialed individuals, and reference projects you do not have. And your professional liability insurer has not decided how it feels about any of this.

Response

The procurement half is correct and, for a new entrant, close to disqualifying — which is why it drives the recommendation in §9 rather than being waved away.

Under the Brooks Act (40 U.S.C. §§1101–1104), federal agencies must publicly announce A/E requirements, discuss with at least three firms, rank at least three on "demonstrated competence and qualification," and only then negotiate price with the highest-ranked firm.28 Price is not a selection criterion. Forty-six states had adopted some form of qualifications-based selection as of NSPE's 2018 count.28b A cost-advantaged entrant cannot buy its way onto the shortlist.

Worse, the federal instrument itself is built around named humans. Standard Form 330 requires a resume for each key person, their current professional registration by state and license number, their years of relevant experience, and a matrix cross-referencing each named individual against the specific past projects they personally worked on in a similar role.29 There is no field for "our model." State DOT prequalification is stricter still: FDOT Rule 14-75 requires that listed personnel be bona fide employees, that engineers hold Florida registration, that experience thresholds be met personally, and it formally grades past performance on schedule, management, and quality — with unsatisfactory grades triggering suspension.30

On insurance, the picture is less alarming than commonly claimed. Ames & Gough's 2026 survey of fifteen leading A/E professional liability carriers found 80% view AI adoption as "a potential disruptor of the overall professional liability market," with 73% "planning modest rate increases largely in the single digits."31 The AIA Trust's 2025 review found carriers "focused less on the technology itself and more on quality control, professional judgment, and potential copyright concerns," and concluded that "negligence is evaluated based on outcomes and professional responsibility, regardless of the tools used."32 AI-specific exclusions certainly exist — Berkley's "absolute" AI exclusion across D&O, E&O and fiduciary; Hamilton's generative-AI exclusion in professional liability; ISO's optional January 2026 endorsements CG 40 47, CG 40 48 and CG 35 08 — but all of those are adjacent lines, not design professional E&O.33 No professional liability claim over AI-generated design work appears in the public record as of this writing. And construction counsel have flagged the opposite possibility as an open question: "whether the use of AI will be required to meet the applicable standard of care for designers."32b

One inference to label as mine, not the sources': no A/E professional liability carrier publishes a statement that AI exclusions are absent from its design-professional forms. What the record supports is narrower — the AIA Trust's review identifies no AI-specific coverage gaps and describes carriers treating AI under existing negligence frameworks, while every documented exclusion I could find sits in an adjacent line. That is suggestive, not dispositive. Read your own policy.
What this settles: Oxvelo cannot originate public-agency work as a new entity. Qualifications, licensed named staff, an auditable cost history, and above all a graded record of completed projects mapped to named individuals are not obstacles that speed and cleverness overcome — they are time-locked. They must be acquired, not built.
08 — The incumbent's dilemma

The industry will not do this to itself. That is the opportunity.

There is a common view that engineering is a reactionary field — dominated by old hats who believe that if it ain't broke, don't fix it. The view is directionally right about the outcome and wrong about the cause, and the difference matters enormously, because the real cause is not temperament. It is written into the law of the profession.

Engineering is not conservative about technology. It is conservative about accountability.

The evidence against the simple "dinosaurs" story is easy to find: this is a field that absorbed finite element analysis, performance-based seismic design, high-performance concrete, parametric modeling, and drone-based survey without much drama. Engineers are not afraid of new methods. What the profession is structurally slow about is changing who is accountable for what, and how that accountability is demonstrated.

And there is a specific legal mechanism that produces this, which almost nobody outside the professional-liability world names out loud. The standard of care — the thing a design professional is legally measured against — is not codified in a statute; it is a common-law test, commonly restated as "the degree of care and skill ordinarily exercised by practicing professionals performing similar services under similar circumstances."32c It reaches contract in black and white: AIA B101-2017 §2.2 obliges the architect to perform "consistent with the professional skill and care ordinarily provided by architects practicing in the same or similar locality under the same or similar circumstances."32c

Read that carefully. Correct practice is legally defined as what everybody else is already doing. That is a conservatism engine, and it runs on its own. It makes innovation asymmetrically risky in a way it is not in almost any other industry: follow the herd and fail, and you met the standard of care; innovate and fail, and you did not. A prudent principal facing that payoff structure waits — not because they are a dinosaur, but because waiting is the rational play under the rules they are actually judged by.

Correct practice is legally defined as what everybody else is already doing. Engineering doesn't have a culture problem. It has a conservatism engine, and it is written into the standard of care.

What the resistance actually looks like in the data

The result is a profession that overwhelmingly believes the change is coming and overwhelmingly has not made it. This gap is measurable, and it is remarkably consistent across three independent surveys of three different populations.

Figure 12 · The belief–action gap
In each pair, the blue point is what practitioners say about AI and the orange point is what they have actually implemented. Three surveys, three populations, the same shape.
What they believe or intend What they have actually implemented
Sources: Dodge Construction Network with CMiC (December 2025) — 87% of contractors believe AI will have "a meaningful impact on construction," while usage is under 15% for 20 of the 23 functions studied (note the belief is about the industry, not the respondent's own business — which sharpens the gap rather than softening it);9b AIA with Deltek and ConstructConnect (March 2025) — 84% see potential to automate manual tasks, 6% regularly use AI in their job;8 Autodesk State of Design & Make: AI Pulse (2026) — 98% of leaders use at least one AI tool, 16% have implemented agentic AI with a further 43% planning to within a year.10

The texture underneath those numbers is worth stating plainly:

  • 52% still use paper during the design phase, 49% during planning, and 43% still depend on physical signatures — and only 11% describe themselves as fully digital organizations.9 This is not a field that skipped digitization and is leaping to AI. It is a field that has not finished digitizing. Note again that these are technology decision-makers at manager level or above — the digitally most engaged slice of the industry.9b2
  • 90% of architectural professionals report concern about a cluster of five issues — inaccuracy of AI outputs, unintended consequences, security, authenticity, and transparency.8 (AIA reports the 90% across the bundle, not for inaccuracy alone.) Enthusiasm and alarm coexist in the same respondents.
  • Sentiment has cooled, not warmed. Autodesk's 2025 survey found 69% of leaders believed AI would enhance their industry — down twelve points from the prior year — while the share saying AI would destabilize their industry rose seven points to 48%.11
  • Capability is concentrated in exactly the firms that need it least. 61% of large architecture firms use AI in day-to-day work against 27% of small ones7 — and 89.1% of US A/E firms have fewer than twenty employees.12 The overwhelming majority of the industry has neither the budget nor the staff to build any of this.
  • The billing model punishes efficiency. ENR's analysis of AECOM's $390M Consigli acquisition identified monetization constraints from time-and-materials contracts as a structural headwind.46 A firm that bills by the hour and halves the hours has cut its own revenue. There is no version of the incumbent solving this from inside a T&M contract.
  • Regulatory silence reads as risk. 69% of surveyed AEC professionals said regulatory uncertainty affected their plans;9 NCEES's Model Law and Model Rules, in their most recent editions, contain no mention of artificial intelligence at all.20 To a licensee, an unaddressed question is not permission — it is exposure.

And behind all of it sits the demographic fact. An estimated 184,175 engineers retired or left the field in 2022, 85,175 from the core civil, mechanical, and electrical disciplines, while degrees awarded have fallen more than 10,000 from their 2019 peak of about 214,000.15 The people leaving hold the institutional knowledge. The people who would rebuild the operating model are not arriving in sufficient numbers to do it.

The strategic reading: this is the moat, not the obstacle

Every fact above is usually presented as a reason to avoid this industry. It is the opposite. A market that changes fast has no room for an outsider — by the time you arrive, the incumbents have already done it. A market with a conservatism engine wired into its liability regime, a billing model that penalizes efficiency, a size distribution where 89% of firms cannot fund a platform, and a retirement wave draining the knowledge base is a market where the window stays open long enough to matter.

It also explains why the entrant should be a generalist rather than a specialist. Firms saturated with engineers optimize the engineering. The unexploited surface — contracts, proposals, project controls, document management, the entire production system connecting them — is invisible from inside the discipline precisely because everyone inside it was trained to look somewhere else. Seeing the whole operating system as the target requires standing outside it.

The consolidation data confirms the window is open and also that it is not open forever. Record transaction volume, 72% of 2025 deals involving firms under $10 million in revenue, and private equity in more than half of all design and environmental consulting deals13 means the founder-owners are already selling. The question is only who they sell to, and whether the buyer brings an operating model or just a balance sheet.

The line this argument must not cross

Read this section as "the incumbents are dinosaurs and we will route around them" and you will build something dangerous. The caution is not irrational — it is calibrated to consequences. A bad blog post is embarrassing; a mis-designed beam kills people. The profession's conservatism is the reason bridges mostly stand up, and any entrant who treats it as an obstacle to be defeated rather than a requirement to be satisfied will eventually earn the enforcement action they deserve.

The correct response to a conservatism engine is not to argue with it. It is to feed it what it wants: verifiable output, complete provenance, retained review records, named licensed professionals in genuine responsible charge, and a measurable track record. Do that and the standard of care stops being a wall. Construction counsel have already raised the reverse possibility — that AI use may eventually be required to meet the standard of care.32b The firm that has been documenting its verification discipline for three years when that turns is not the laggard. It is the new baseline.

09 — The decision

Build a team, or partner with an established firm?

This is the question the paper exists to answer. The honest answer is that framing it as either/or is the mistake — and the evidence points clearly to a staged structure in which the first move is partnership and the second is ownership of the platform, not the practice.

What building from zero actually costs — in months, not dollars

The individual steps to stand up a licensed engineering firm are each modest. The problem is that they are largely sequential, and one of them cannot be accelerated at any price.

Figure 13 · The clock on building a licensed, bid-capable firm — in one state
Each bar is the published range for that step. They stack, because each depends on the prior one completing. The final step carries two requirements a brand-new entity structurally cannot satisfy at any price: an auditable cost history, and a graded record of completed projects mapped to named staff.
Sources and the honest limits of each. NCEES Records: "the average time for completion is 2–3 weeks," accepted by all US boards — solid.50 Comity licensure: the 1–3 month range is illustrative, not typical — no authority publishes a national figure, and state board meeting calendars dominate.50b Certificate of authorization: "34 states" and "4–22 weeks" come from a commercial compliance vendor, not a regulator, under its own taxonomy; roughly 42 states impose some firm-level requirement.49 Prequalification: 3–4 months is a third-party guide's figure.30b

A correction worth making explicitly, because the commonly repeated version is wrong. FDOT professional services prequalification requires a current overhead (FAR) audit — not three years of audited financial statements. The three-year audited-financials requirement belongs to FDOT's separate construction contractor prequalification under Ch. 337.14, Fla. Stat.30 The two programs are routinely conflated in secondary sources. The real time-lock is different and stronger: a new firm has no cost history to audit, and — decisively — no graded past-performance record and no completed projects to map to named staff. SF 330's Section G matrix demands exactly that cross-reference,29 and FDOT grades past work on Schedule, Management and Quality on a 1–5 scale.30 You cannot manufacture completed projects. Multiply the whole stack by each additional state.

Stack the published ranges and it is roughly six to thirteen months before a new entity can bid public work in a single state — with a hard floor set by financial history that no amount of capital compresses. Against that: BizBuySell's 2025 architecture and engineering benchmarks put the median small-firm sale price at roughly $825,000 on median seller's discretionary earnings of about $451,450, while Zweig Group's 2025 data give a median of 4.28× EBITDA across all size bands, with small firms transacting nearer 2.5–3.5× on a blended SDE/EBITDA basis.51 You can buy licensure, registration, prequalification, audited history, a client list, and a data corpus of prior projects for less than a seed round.

What the AEC precedent actually shows

Almost nothing. That is the finding, and it cuts both ways. AI funding into AEC startups reached $616 million across 46 rounds in the first half of 2026 alone, versus $318 million across 31 startups in all of 2025 — and virtually all of it went to companies selling software to firms, targeting what one sector analysis called "point solutions" addressing "specific operational failures."52 The traffic runs the other direction: AECOM bought Consigli,46 Autodesk bought MaintainX for $3.6 billion.46b

The one greenfield precedent is cove.tool, which in March 2025 launched "Cove Architecture," billed as the first full-service AI-powered architecture practice, on the back of a claimed "over $25 million" and "nearly a decade" of R&D. Its inaugural project was a 15-unit housing complex in Atlanta's West End, with claims of 60% shorter design timelines and completion in 15 days.45 It is roughly eighteen months old with a small public record; no source states the practice's firm-level registration, though Cove's own team page lists two AIA-credentialed principals, so it is not unstaffed by credentialed architects.45b It is a signal, not a proof. Its CEO's own framing is notably close to §5's argument: "You need experienced individuals who understand the outputs. You don't want an intern pressing a button and making a building."45

Outside AEC the precedent is much richer, and it points one way. Eudia — backed by a General Catalyst-led Series A of up to $105M — acquired Johnson Hana, a 300-plus-person legal services provider, in July 2025, with co-founder and CEO Omar Haroun giving the explicit reasoning: "we've learned that building truly transformative systems requires deep human expertise embedded within the technology itself."42 Crete Professionals Alliance committed over $500 million to acquiring US accounting firms, backed by Thrive Capital, and already spans 20+ businesses and $300M+ in revenue.43 The UK's Solicitors Regulation Authority authorized Garfield.Law in May 2025 as the first AI-driven law firm — with conditions: the system cannot propose case law, clients must approve actions, and named regulated solicitors remain ultimately accountable for all outputs.41 That conditions list is a preview of what an engineering board would impose.

And in insurance, the staged path is well established: startups begin as managing general agents writing business on a licensed carrier's paper, prove the model, then move to full licensure — Next Insurance, Clearcover, Metromile and Kin all made that transition.57 (Not universal, and worth saying so: Lemonade and Root launched as licensed carriers from the start, so the pattern is a common route rather than the only one.) The closest structural parallel to Oxvelo's question is Pie Insurance, which raised $127 million in 2020 and earmarked $100 million of it to "form and purchase licensed insurance companies" through a dedicated affiliate.57b Note what Pie did: it capitalized the build-or-buy decision rather than assuming an answer.

The comparison, scored

RequirementBuild a team from zeroPartner with / acquire a firmStaged hybrid
PE licensure & certificate of authorization8–14 moDay oneDay one
Auditable cost history for FAR overhead auditNone to auditInheritedInherited
Graded past performance & reference projects (SF330, DOT)NoneExisting recordExisting record
E&O history & insurabilityUnratedEstablishedEstablished
NY ownership compliance for a non-licensee founderBlockedMinority onlyMSO owns platform
Proprietary data corpus for the AI layerEmptyDecades of projectsDecades of projects
Freedom from legacy process & hourly-billing incentivesTotalInherits bothIsolated in the MSO
Enterprise value capture / multipleServices multipleServices multiplePlatform multiple
Capital required to first revenueHigh, slow~$0.8–3MLow — MSA first
Regulatory precedent for the structureConventionalConventionalProven in 3 professions

Scoring is Oxvelo's judgment. Every row's underlying constraint is cited elsewhere in this paper; the ratings are not.

The recommendation

Partner first, own the platform, acquire second. Building an engineering team from zero fails on the one requirement that cannot be compressed — three years of audited financials and a graded past-performance record. Acquiring a firm outright as the opening move buys those, but spends scarce capital on a services multiple before the AI layer has proven it changes the economics, and — in New York — cannot legally be controlled by a non-licensee founder anyway.

The staged path resolves all three. Stand up Oxvelo as the management and technology entity. Contract with an existing licensed design firm under a management services agreement, exactly as AI-native law and accounting platforms have done. Prove the operating-model change on their book of business, against their data, under their seal. Then decide — with evidence rather than conviction — whether to acquire the practice, replicate the structure across states, or stay a platform.

Figure 14 · The staged path
Roadmap — Oxvelo. The specific first moves at each stage are set out in §10.
10 — What to do next

Six moves, in order.

  1. 1

    Get the New York structure ruled on before anything else

    Retain a New York professional-licensing and corporate attorney to answer two questions in writing: whether an outside non-employee entity may hold the sub-25% BCL §1503 tranche, and whether a percentage-of-revenue management fee from an engineering MSO would be characterized as unlawful fee-splitting.2456 Everything downstream depends on these answers. This is a four-figure decision that determines a seven-figure structure.

  2. 2

    Write the constitution before writing the software

    A founding operating document that defines the AI workforce, its permissions, its escalation rules, and — most importantly — the human control points and the audit record each one produces. In New York that record is not optional: a licensee reviewing work prepared by others must retain a written evaluation for six years.59 Build the audit trail as a first-class feature, not a compliance retrofit. It is also the artifact that makes the firm insurable and the structure defensible.

  3. 3

    Map one firm's complete value stream, from inquiry to closeout

    Client acquisition, opportunity qualification, proposals, estimating, contract negotiation, project setup, technical production, QA/QC, project controls, construction support, billing, collections, closeout. For each stage: tasks, decisions, documents, handoffs, responsible parties, recurring pain points, data inputs. This is the automation blueprint, and it is also the diligence document for choosing a partner.

  4. 4

    Find the partner firm — and screen for the right one

    Target profile: 15–60 people, one to three states, an owner within five years of exit, a strong past-performance record, and — critically — a principal who is frustrated rather than complacent. Avoid firms whose economics depend on maximizing billable hours on time-and-materials contracts; that is the structural disincentive ENR identified in the AECOM/Consigli deal.46 Look for firms with fixed-fee or lump-sum work, where reducing hours increases margin instead of reducing revenue.

  5. 5

    Hire for judgment, not for production capacity

    Two or three unusually senior licensed professionals, hired explicitly into the three roles in §5: knowledge architect, reviewer in responsible charge, and creative on the novel. Screen for people who can articulate why a detail is right, not just draw it — the encoding work requires someone who can externalize judgment, which is a rarer skill than exercising it. Do not hire a production team; the entire thesis is that the production team is the thing being replaced. Note that the licensed reviewer must sit inside the professional entity, not the management entity, or the structure in Figure 11 fails on its own terms.

  6. 6

    Build one end-to-end demonstration: opportunity to project kickoff

    Identify an opportunity, draft the proposal, analyze the contract against a company standard, build the fee and schedule, create the project workspace, establish the deliverable register, generate the kickoff package. Multiple AI workers collaborating, one human approving at each defined control point, every step logged. Measure the cycle time against the firm's historical baseline — measured, not self-reported, for the reason METR's participants demonstrated.35 One credible before-and-after number on a real project is worth more than the rest of this document.

11 — Intellectual honesty

What would prove this wrong.

A thesis that cannot be falsified is a slogan. Four specific outcomes would mean this strategy is wrong, and they should be watched deliberately rather than rationalized:

  • The kickoff demonstration does not beat the baseline on measured cycle time. If a full AI workforce cannot compress opportunity-to-kickoff on a real project against that firm's own history, the operating-model claim is false and no amount of structure fixes it.
  • The leverage ratio stalls near 2. The whole §5 argument rests on L climbing well past the doubling that McKinsey's "50% of non-physical work" implies.2 If measured, audited throughput on repeatable packages plateaus at 2–3× with an honest review standard held constant, then this is a good consulting business and not a revolution — and the paper's boldest claim should be retracted rather than re-explained.
  • Reviewer catch rates fall as volume rises. The cleanest single early-warning metric in the whole model. If seeded-defect detection degrades as packages-per-reviewer climbs, the leverage is buying itself with silent risk, and the cap has already been exceeded.
  • The partner firm's staff route around the system. The Bluebeam data showing 52% of AEC professionals still using paper in the design phase and 43% still relying on physical signatures9 is a warning about workflow gravity, not just about digitization. If the licensed professionals treat the AI layer as an extra step rather than a replacement step, the productivity never materializes — which is precisely how BIM's promise dissolved.
  • A state board or insurer moves against the structure. Currently no A/E professional liability carrier is writing AI exclusions,32 and Texas's board has explicitly declined to write AI-specific rules.22 Both could change. NCEES's Model Law and Model Rules, as of their most recent editions, contain no mention of artificial intelligence at all20 — the governing framework is silent, which is currently permissive and could stop being.
  • The economics land on a services multiple. If the platform never separates from the practice, Oxvelo has bought a small engineering firm with better software — a fine business, valued at 2.5–3.5× EBITDA,51 and not the one described here. Nines is the documented version of this failure.47
The vision is not to automate engineering calculations. It is to reinvent how the engineering business operates — and the fastest legal path to doing that runs through someone else's license before it runs through your own.
Appendix A

Glossary

This paper spans four vocabularies — engineering licensure, public procurement, private-equity finance, and machine learning — and almost nobody is fluent in all four. Every term below is also tappable inline throughout the paper.

Appendix B

How to read the numbers in this paper.

Not all evidence is equal, and a document that presents a peer-reviewed benchmark and a vendor press release in the same typeface is quietly misleading its reader. Every figure in this paper falls into one of five tiers. Where a claim rests on tier 3 or below, the text says so at the point of use.

TIER 1
Primary legal text. Statutes, administrative codes, board rules, and standard contract forms — NCEES Model Law and Model Rules, NY Business Corporation Law §1503(b-1), 22 Texas Administrative Code, the Brooks Act, SF 330, FDOT Rule 14-75, AIA B101. These are quoted verbatim and can be checked word-for-word. Treat as settled.
TIER 2
Peer-reviewed research and official statistics. AECBench (Advanced Engineering Informatics), the hallucination paper, the automation-bias literature, NBER working papers, US Census series. Methodology is disclosed and the work has survived review. Treat as reliable within its stated scope — and note the scope, which is often narrower than the headline.
TIER 3
Professional-body and industry surveys. Deltek Clarity, AIA Firm Survey, ACEC Research Institute, Ames & Gough, Autodesk, Bluebeam, Dodge. Large samples and consistent methodology, but self-selected respondents, self-reported answers, and questions written by parties with an interest in the answer. Directionally trustworthy; precise to perhaps a few points, not a decimal.
TIER 4
Trade press and market intermediaries. ENR, Morrissey Goodale, Capstone Partners, Zweig Group, BizBuySell, Memoori, Harbor Compliance. Well-informed, often the only source that exists for a given fact, but commercially interested and rarely methodologically transparent. Definitions vary between them — which is why two credible sources can report 361 and 500+ transactions for the same year and both be right.
TIER 5
Company-reported claims. cove.tool's project metrics, Crosby's contract volumes, vendor case studies. Unaudited, unreplicated, and published to sell something. Used in this paper only as existence proofs — evidence that someone is attempting the thing — never as measurements of what the thing achieves. Every instance is flagged in place.

Two figures deliberately excluded

The widely quoted "35% of construction professionals' time is non-productive, costing $177 billion annually" is real and well-sourced, but its survey population was 49% general contractors and 36% specialty trades — not design professionals.18 Applying it to engineers and architects would be the kind of small dishonesty that costs a reader's trust in everything else. The US Census SUSB firm-size figures for NAICS 5413 were removed after they could not be verified from a retrievable source; the fragmentation argument was rebuilt on AIA and ENR data instead.

Appendix C

The leverage arithmetic, worked.

Figure 5 compresses a calculation into a curve. Here it is in full, so a reader can disagree with a specific assumption rather than with the shape of a line.

Definitions

L = leverage ratio — engineer-equivalents of production output that
    one licensed professional can review and take responsibility for
B = baseline net revenue per employee = $143,000 (AIA, 2023 data, mean)
C = total labor cost per employee = $119,511 (Deltek Clarity 47th)
O = overhead rate on direct labor = 161.3% (Deltek Clarity 47th)
U = utilization = 58.9% — share of paid hours that are billable

The curve

Net revenue per licensed professional = B × L

L = 1  →  $143,000    (the industry today)
L = 2  →  $286,000    (McKinsey-implied: 50% of non-physical work)
L = 5  →  $715,000    (conservative near-term target)
L = 10 →  $1,430,000
L = 25 →  $3,575,000
L = 50 →  $7,150,000  (2 reviewers carrying a 100-engineer production load)

The four assumptions, and what breaks each

  1. Price per deliverable holds constant. It will not, in a competitive market. But note what happens if it falls: revenue per professional rises more slowly while margin still rises, because cost per deliverable falls faster than price. The leverage shows up in the P&L either way — the question is only whether the client or the firm captures it.
  2. The contract is fixed-fee or lump-sum. This is the assumption that actually decides the business. On time-and-materials, halving the hours halves the revenue and the entire curve inverts into a loss. ENR identified precisely this as a constraint on AECOM's Consigli acquisition.46 Partner selection must screen for fee structure.
  3. Review effort does not scale linearly with output. If it does, L is capped near 1 and the thesis fails outright. The entire verification architecture in Figure 6 exists to break that linearity, and the assurance layer exists to detect it if we are fooling ourselves.
  4. The work is templatizable. L applies to repeatable production against established standards. On first-of-its-kind design, L stays near 1 and should.

What the model does not claim

It does not claim a hundred-person firm becomes two people. A hundred-person engineering firm is not a hundred people drafting — it contains project managers, construction administration staff, principals, finance, and business development. The curve applies to the design production function only. The other functions compress on their own separate curves, described in Figure 3, and none of them reaches L = 50.

It also does not claim any of this has been achieved. L is a measurement to be taken, not a result to be assumed — which is why §11 names a stalled leverage ratio as the first thing that would falsify the whole paper.

Appendix D

Regulatory quick-reference.

The three states this paper examines, side by side. Nothing here is legal advice — it is a starting point for a conversation with counsel, and the New York column in particular has a load-bearing ambiguity that only an attorney's opinion can close.

QuestionNew YorkCaliforniaTexas
Governing provisionEduc. Law §§7202, 7210; BCL §1503(b-1)Bus. & Prof. Code §673822 Tex. Admin. Code §137.77
Non-licensee ownershipOrdinary PSC: none. DPC: <25%, but reserved for ESOPs and non-licensee employees — an outside investor entity does not qualifyPermitted, but a non-licensee may not be sole ownerNo ownership restriction in the rule
Licensed leadership requiredPresident, chairperson and CEO must all be design professionals; largest shareholder tooA licensed engineer must be owner, partner, or officer in chargeAt least one full-time active licensee employed
Board / director composition>75% design professionalsNot specifiedNot specified
LLC permitted?PLLC, all members licensedNo — Corp. Code §17701.04(e)Yes
Branch coverageAll engineering§6738 reaches civil, electrical, mechanical onlyAll engineering
Practical read for OxveloMSO/PC required. No equity route into the licensed entityWorkable with a licensed principal in chargeMost permissive of the three

The questions counsel must answer before any structure is committed

  1. Can an outside, non-employee entity hold any equity in a New York design professional service corporation? The statutory text points to no. If that is right, the MSO/PC split is not a preference — it is the only lawful structure.
  2. Would a percentage-of-revenue management fee from an engineering MSO be characterised as unlawful fee-splitting, as it is in the New York healthcare analogue? If yes, the fee must be flat, time-based, or cost-related, and the financial model must be built that way from the start.
  3. Does the six-year written-evaluation retention duty under 8 NYCRR §29.3(a)(3)(i) attach to AI-produced work reviewed by a licensee? Assume yes and build the audit trail accordingly; it is cheap insurance either way.
  4. In which states does the licensed entity need a certificate of authorization before it can be named on a proposal, and what is the current processing time in each?
Appendix E

The client question — which nobody has asked clients.

Every principal who reads this paper will arrive at the same objection within five minutes: will my clients accept it? The paper owed an answer. Researching it produced a finding more interesting than the answer itself.

The finding: the question has never been put to clients

We checked every substantial AEC survey on artificial intelligence we could find — RICS (2,200+ professionals), RIBA's 2025 and 2026 reports, Chaos (~800 architects), Arup (5,000 professionals across ten countries), Mastt, Deltek Clarity (896 firms), and the AIA/Deltek/ConstructConnect study.60 Every one of them surveys practitioners. Not one asks clients whether they want their design consultants using AI.

The nearest owner-side data is Dodge Construction Network's work with the National Institute of Building Sciences, which surveyed roughly 200 US owners each managing $5M+ of annual construction. It found 28% of owners already using AI themselves, with the fastest growth rate of any technology studied, and delivered this warning to their consultants: data capability "is a competitive advantage now, but it is only a matter of time before it is a requirement."61 That measures owners' own adoption. It does not measure their attitude to yours.

Why an absence is worth publishing

An industry has spent three years debating whether AI belongs in design practice without asking the people who pay for design. That is not a small oversight — it is a measurement gap sitting exactly where the commercial risk lives. It is also, for a firm willing to do the work, an opportunity: the study we are fielding asks it directly.

What is arriving is not client objection

Three real forces are moving, and none of them is a client refusing AI-assisted work.

1. Public buyers are mandating disclosure

California's Department of General Services amended its State Contracting Manual (§2302) to require that a bidder notify the State in writing if it intends to use generative AI to complete work that materially affects system functionality, risk to the State, or contract performance. It applies to all written solicitations regardless of acquisition type, and non-disclosure "may result in bid disqualification."62 This is the only binding, generally applicable US buyer-side disclosure requirement we could verify.

Federal policy is softer. OMB M-25-22 says agencies "are advised to consider" requiring contractors to disclose AI use — advisory, not mandatory.63 GSA circulated a draft clause in March 2026 requiring disclosure of all AI systems used in contract performance; comments closed in April and it was not included in the subsequent refresh.64 And notably, Caltrans' own scan of state GenAI procurement frameworks identified no state DOT with a binding policy requiring consultants to disclose AI use in deliverables.65

2. Insurers are asking questions — buyers are not

NSPE's professional-liability reporting found that most underwriters "are not yet incorporating AI into underwriting criteria, but they are beginning to ask firms how they supervise and review AI-assisted work," with "no widespread exclusions or special policy modifications yet evident."66 Ames & Gough's carrier survey adds that firms with proactive AI governance will "gain underwriter confidence."31 We found no evidence of AI questions appearing in client prequalification questionnaires. The scrutiny is coming from the people carrying the risk, not the people buying the service — which is a meaningful distinction, and an argument for building the audit trail early.

3. No standard-form AI clause exists yet

Neither AIA, EJCDC, nor ConsensusDocs has published an AI clause or addendum for design agreements as of August 2026.67 Where AI provisions are appearing, the direction is instructive: they run contractor-to-owner as protective disclaimers, not owner-to-designer as restrictions. One construction attorney describes drafting a provision stating that where AI is used for scheduling or preliminary estimates, "the customer cannot rely on those until such time as boots on the ground have determined whether those field measurements are accurate."68

The real commercial risk is the fee, not the refusal

The threat this paper takes most seriously is not that clients reject AI-assisted work. It is that they accept it and expect to pay less for it.

In professional services broadly, that is already measurable: a 2026 survey of 1,000+ senior business leaders found 66% reporting that clients are becoming more demanding while less willing to pay, with two-thirds expecting AI to reduce costs, and a third saying clients now see AI as a route to bringing work in-house rather than hiring specialists.69 That study is cross-sector, not A/E — the transfer is an assumption, and we label it as one.

In A/E specifically, what is documented is a shift in the contracting model rather than the fee level. AECOM's CEO Troy Rudd describes clients themselves raising it: moving "away from something like cost plus to something that looks more like a fixed fee."70 And ACEC's work with Virginia Tech states the underlying danger with unusual bluntness:

If a consulting engineering firm can deliver the same work product 30% to 50% more efficiently in the future but still charges for its services by the hour, it is fundamentally in a race to the bottom.ACEC with Virginia Tech, 2024

That is the single strongest external corroboration of this paper's commercial argument, and it arrives from the engineering industry's own trade body rather than from a technology vendor. It is also why the leverage calculator makes fee mix the setting that dominates every other input.

What we could not find at all

  • No published case of a client rejecting or disputing AI-produced design deliverables. The nearest real disputes are bid protests dismissed by the GAO for AI-fabricated legal citations — five decisions in 2025 — plus a contested Army source-selection where the buyer's own AI tool was alleged to have hallucinated evaluation findings on a $449M engineering services award.71 In both categories the AI failure sits in the procurement paperwork, not the engineering.
  • No evidence that any owner scores AI capability in A/E selection. Every search on this returned results about using AI to write proposals — the opposite question.
  • No A/E-specific instance of a client demanding a fee reduction on AI grounds. The mechanism is credible and the adjacent-sector evidence is quantified, but the A/E case has not been documented.

Three genuine blanks. In a paper that concedes three of its seven objections, it would be inconsistent to dress absences up as answers.

Appendix F

Oxvelo's AI disclosure policy.

NSPE's Board of Ethical Review found there is no professional or ethical obligation to disclose AI use to a client unless the contract requires it — while adding that "ethical principles favor transparency when AI plays a substantial role in generating work products."21 That is permission to say nothing. We think saying nothing is the wrong call for a firm whose entire proposition is verifiable accountability, so this is our standing policy, published in advance of having clients.

1. We disclose by default

Any deliverable produced with material AI involvement carries a statement of that fact, whether or not the contract requires it and whether or not the client asks. Disclosure is not an admission; it is a description of method, in the same way a firm describes which code edition it designed to.

2. We name the human, not the model

Every sealed deliverable identifies the licensed professional in responsible charge by name and licence number. No document will ever suggest that a system approved anything. Responsibility is personal, it is legally required to be personal, and diffusing it across a "platform" would be the first step toward the plan-stamping failure this paper warns about.

3. The review record is available to the client on request

The exception set a reviewer adjudicated, the deterministic checks that ran, and the provenance links behind cited code provisions are retained and disclosable. In New York this is already obligatory — a licensee sealing work prepared by others must retain a written evaluation for six years59 — so we treat it as a client-facing feature rather than a filing-cabinet duty.

4. Client data does not train anything

Client project data is not used to train third-party models, and is not routed to any service that reserves the right to do so. NSPE's Case 24-2 found the engineer at fault partly for disclosing private client information to an AI tool without permission — a failure of data handling, not of design.21 Where a client's contract restricts data routing, that restriction governs model selection, not the other way around.

5. We answer the public-buyer question before it is asked

Any bid into a jurisdiction with a disclosure regime — California's DGS requirement being the clearest62 — carries the disclosure whether or not the solicitation demands it. The downside of over-disclosing is a conversation. The downside of under-disclosing is disqualification.

6. We do not claim what we cannot show

No efficiency, accuracy or cycle-time claim is made to a client without a measured baseline behind it. This is the same standard we hold our own published research to, and it exists because METR's participants believed they were 20% faster while measuring slower.35

Why publish this at all

Because a policy written after the first difficult conversation is a rationalisation, and everyone can tell. Publishing it now — before there is a client, a claim, or an incentive to soften it — is the only moment at which it costs nothing and means something. If we ever depart from it, this page is the evidence.

Appendix G

Worked example: opportunity to project kickoff.

Read this correctly

This is a design specification, not a result. No project has been run through this process. Nothing here is a measurement, a case study, or a claim about performance. It is the blueprint for the demonstration described in §10 — published so that when the demonstration runs, it can be judged against what was specified beforehand rather than against a narrative written afterwards.

The paper argues its mechanism at length and never shows it. Here is one complete workflow, stage by stage, with the human control point and the audit artifact named at each step.

StageAI workforce actionHuman control pointAudit artifact produced
Opportunity intakeParses the RFP or inquiry; extracts scope, deliverables, deadlines, submission requirements, evaluation criteria and disqualifiersNone — retrieval onlyStructured opportunity record with every extracted term linked to its source page
Go / no-goScores fit against the firm's licence coverage, prequalification status, staff availability and past-performance record; flags anything the firm cannot legally satisfyPrincipal decides. The system may not advance an opportunityDated decision record with the reasoning presented and the decision taken
Fee and scopeRetrieves comparable prior projects, builds a first-pass effort model by discipline and phase, drafts the scope narrative from firm-standard languagePrincipal sets the fee. Non-negotiable — this is a commercial commitment, not a calculationEffort model with each assumption traced to the prior project it came from
Proposal assemblyAssembles the SF 330 or private-format response; populates named-staff resumes, registrations and matched project experience from the personnel recordNamed individuals confirm their own resume entries and registration statusPer-person confirmation log
Contract reviewCompares the proposed agreement against the firm's standard positions; ranks every deviation by risk with the firm's standard clause shown beside the client'sPrincipal accepts or rejects each flagged deviation. Signature is a human actDeviation register with disposition and rationale per item
Project setupOn execution, creates the workspace, budget structure, phase and task breakdown, and compliance logProject manager confirms structureSetup manifest, timestamped against contract execution
Deliverable registerDerives every required deliverable from the contract scope and schedule, with owner, due date and dependencyProject manager confirms completenessRegister with each line traced to the contract clause requiring it
Kickoff packageDrafts agenda, roles, standards to be applied, code editions in force, submittal schedule and communication protocolProject manager approves and issuesIssued package with revision history

What is deliberately absent

No stage above produces sealed engineering work, and none advances without a named human decision at the points where commitment or liability attaches. The seal appears nowhere in this workflow, because this workflow ends at kickoff — which is precisely the point of entering through the beachhead in Figure 4.

The measurement that decides whether any of this is true

One number, taken the same way twice:

Baseline = median elapsed days, opportunity identified → kickoff package issued,
          across the partner firm's last 10 comparable pursuits
Measured = the same interval, same definition, on the next 5 pursuits run through the system

Also tracked: professional hours consumed per pursuit · win rate
                   exceptions raised per pursuit · exceptions overturned on review

Elapsed time is the headline, but exceptions overturned on review is the number that matters most. If reviewers almost never overturn what the system flags, the system is not surfacing genuine judgment calls and the review is decorative. A healthy exception rate is not zero.

Win rate is included because it is the one metric that could invalidate everything while every other number improves. A faster, cheaper proposal that wins less work is not an operating-model improvement. It is a worse business with better dashboards.

Appendix H

Verification log.

Every claim in this paper was checked against its primary source in a dedicated fact-checking pass. Corrections made in that pass are recorded here, because a reader is entitled to know which way the errors ran before they were fixed.

Corrections that changed the argument

  • NCEES redefined responsible charge in August 2025. The widely quoted "direct control and personal supervision" formulation is superseded and appears nowhere in the current Model Law. §5 now argues from the current four-duty test — authority to review/change/reject/approve, personal awareness of scope, capability to answer questions, acceptance of full responsibility — which is more favourable to this thesis than the language it replaced.
  • The Census fragmentation figures could not be verified from any retrievable source and were removed. Figure 7 now uses AIA's published firms-and-billings table; the concentration claim was rebuilt on ENR Top 500 revenue against Census Quarterly Services Survey industry revenue.
  • The "three years of audited financials" prequalification claim was wrong. That requirement belongs to FDOT's construction contractor programme, not professional services, which requires a current FAR overhead audit. The binding constraint was restated as the graded past-performance record — better sourced, and harder to overcome anyway.
  • New York's ownership position is worse than first drafted. BCL §1503(b-1) reserves the residual tranche for ESOPs and employees; an outside investor entity does not qualify at all. The effective ceiling for a holding company is zero, not 25%.
  • METR did not retract its 19% finding. The February 2026 figures come from a separate study, not a revision. The earlier draft overstated the walk-back in this paper's favour.

Corrections of attribution and precision

  • $143,000 net billings per employee is a mean; AIA publishes no median.
  • AECBench is Liang et al., not Zhao et al. Zhao is the last author.
  • MGI's $228 billion covers the entire US AEC industry and includes "other automation technologies" — it is not the dollar value of the 50% automation figure, and the two are frequently conflated.
  • ACEC's 14% preparedness figure is estimated by mid-career and older practitioners, not executives; the 34% student figure comes from a 195-person convenience sample.
  • Bluebeam's 27% surveyed technology decision-makers at manager level or above in five countries — not "AEC professionals globally."
  • Dodge's 87% concerns AI's impact on construction as an industry, not on the respondent's own business.
  • AIA's 90% covers a bundle of five concerns, not inaccuracy alone.
  • Lemonade and Root launched as licensed carriers — they are counterexamples to the insurtech staged path, not examples of it.
  • Engineering degree-completion figures come from NCES via ACEC, not ASEE, which counts a different universe.
  • The Elad Gil quotation contains TechCrunch's editorial insertion and is now printed with the brackets intact.
  • ENR's characterisation of AECOM's share-price move could not be independently corroborated and is now attributed to ENR rather than stated as fact.

What remains unverified, by design

Two sets of figures are company-reported and unaudited: cove.tool's project metrics (60% timeline reduction, 95% estimate accuracy, 40% iteration cost, 15-day completion) and Crosby's contract volumes. Both are flagged in place and used as existence proofs, not measurements. One structural claim — that AI-specific exclusions are not standard in A/E professional liability policies — is an inference from the absence of documented exclusions in that line, not a positive finding by any carrier or broker. It is labelled as such in §7.