AI investment

More AI activity is not the same as more business output.

A recent Harvard study offers a useful warning for CFOs: measure the whole workflow before treating a faster task as a financial return.

In Artificial Intelligence in the Firm: Bottlenecks in Software Production (August 2026), Harvard researchers Fiona Chen and James Stratton studied AI coding tools using approximately 300 million work events across 718 firms. The paper is a working paper about software engineering. It does not measure AI return in manufacturing plants.

What the study found

After firms adopted AI coding agents, coding activity rose: lines of code increased about 30%, commits 20%, and pull requests 23%. But the estimated changes in completed Jira issues and larger projects were positive and not statistically significant. The researchers also could not attribute a significant employment change to adoption. This is not proof that AI has no value; it means the measured increase in coding activity did not translate into a comparably clear increase in completed work or a demonstrated reduction in headcount in this sample.

The study identifies a plausible bottleneck. Average pull-request review time increased 49%; the share requiring changes nearly doubled; comments per review increased 35%. Faster drafting can move work to the people who must verify, correct, approve, and integrate it. These are estimates from the authors’ observational design, not a controlled trial of a manufacturing use case.

The CFO interpretationDo not buy a projected return based only on prompts used, documents drafted, code written, or hours claimed to be saved. Ask the process owner what finished output improved, what review work moved elsewhere, and what new cost or risk appeared.

Apply the lesson to a plant or finance workflow

Proposed AI useActivity metricBusiness outcome to test
Invoice coding or AP triageInvoices classifiedApproved invoices per week, exception rate, rework, close time
Maintenance work-order draftingDrafts generatedTime to approved work order, repeat failures, technician review time
Quality-document searchAnswers producedCorrect answers on a test set, reviewer time, audit findings
Demand-planning supportForecasts generatedForecast error, inventory cost, service level, planner override rate

These are suggested test measures, not findings from the Harvard paper. A plant may have a different bottleneck: data quality, supervisor approval, equipment constraints, or ERP integration.

Ask the team for a bounded pilot

  1. Process owner: define the current baseline, acceptable error rate, and where a human must approve a result.
  2. Finance: calculate the full cost: licenses, usage fees, integration, data cleanup, training, review, and ongoing support. Count realized capacity only if it is redeployed or reduces an actual cost.
  3. IT and security: document data shared with the tool, vendor terms, access controls, logging, retention, and the fallback if the tool fails.
  4. Operations or controller: run a representative test, including exceptions and bad data. Compare finished work, quality, elapsed time, and downstream rework with the baseline.
  5. CFO and sponsor: set a decision date and threshold to expand, redesign, or stop the pilot.

Insurance and accountability

Ask the broker and counsel whether existing cyber, errors-and-omissions, and other relevant policies address the contemplated AI use, and whether any application answer or vendor term changes the risk. Coverage varies by policy; do not assume an AI error or data disclosure is insured. Keep a named human owner for financial reporting, customer commitments, safety-related decisions, and production changes.

NIST’s AI Risk Management Framework provides a voluntary structure for governing, mapping, measuring, and managing AI risk. Use it with the pilot’s actual business controls, rather than as a claim of certification.

One-page investment memoProblem and baseline · process owner · proposed tool and data used · full cost · expected business outcome · review burden · security and insurance questions · pilot evidence · go/stop threshold.

Source and limits

Chen and Stratton, Artificial Intelligence in the Firm: Bottlenecks in Software Production, current version August 4, 2026. The paper analyzes software firms using proprietary platform data and an observational difference-in-differences design. Its results should inform how to measure a manufacturing pilot, not serve as an ROI estimate for that pilot. The accompanying video transcript supplied for this article was used only to identify themes; the claims above were checked against the paper.