In Artificial Intelligence in the Firm: Bottlenecks in Software Production (August 2026), Harvard researchers Fiona Chen and James Stratton studied AI coding tools using approximately 300 million work events across 718 firms. The paper is a working paper about software engineering. It does not measure AI return in manufacturing plants.
What the study found
After firms adopted AI coding agents, coding activity rose: lines of code increased about 30%, commits 20%, and pull requests 23%. But the estimated changes in completed Jira issues and larger projects were positive and not statistically significant. The researchers also could not attribute a significant employment change to adoption. This is not proof that AI has no value; it means the measured increase in coding activity did not translate into a comparably clear increase in completed work or a demonstrated reduction in headcount in this sample.
The study identifies a plausible bottleneck. Average pull-request review time increased 49%; the share requiring changes nearly doubled; comments per review increased 35%. Faster drafting can move work to the people who must verify, correct, approve, and integrate it. These are estimates from the authors’ observational design, not a controlled trial of a manufacturing use case.
Apply the lesson to a plant or finance workflow
| Proposed AI use | Activity metric | Business outcome to test |
|---|---|---|
| Invoice coding or AP triage | Invoices classified | Approved invoices per week, exception rate, rework, close time |
| Maintenance work-order drafting | Drafts generated | Time to approved work order, repeat failures, technician review time |
| Quality-document search | Answers produced | Correct answers on a test set, reviewer time, audit findings |
| Demand-planning support | Forecasts generated | Forecast error, inventory cost, service level, planner override rate |
These are suggested test measures, not findings from the Harvard paper. A plant may have a different bottleneck: data quality, supervisor approval, equipment constraints, or ERP integration.
Ask the team for a bounded pilot
- Process owner: define the current baseline, acceptable error rate, and where a human must approve a result.
- Finance: calculate the full cost: licenses, usage fees, integration, data cleanup, training, review, and ongoing support. Count realized capacity only if it is redeployed or reduces an actual cost.
- IT and security: document data shared with the tool, vendor terms, access controls, logging, retention, and the fallback if the tool fails.
- Operations or controller: run a representative test, including exceptions and bad data. Compare finished work, quality, elapsed time, and downstream rework with the baseline.
- CFO and sponsor: set a decision date and threshold to expand, redesign, or stop the pilot.
Insurance and accountability
Ask the broker and counsel whether existing cyber, errors-and-omissions, and other relevant policies address the contemplated AI use, and whether any application answer or vendor term changes the risk. Coverage varies by policy; do not assume an AI error or data disclosure is insured. Keep a named human owner for financial reporting, customer commitments, safety-related decisions, and production changes.
NIST’s AI Risk Management Framework provides a voluntary structure for governing, mapping, measuring, and managing AI risk. Use it with the pilot’s actual business controls, rather than as a claim of certification.
Source and limits
Chen and Stratton, Artificial Intelligence in the Firm: Bottlenecks in Software Production, current version August 4, 2026. The paper analyzes software firms using proprietary platform data and an observational difference-in-differences design. Its results should inform how to measure a manufacturing pilot, not serve as an ROI estimate for that pilot. The accompanying video transcript supplied for this article was used only to identify themes; the claims above were checked against the paper.