+−生成侧沦陷,验证侧堰塞湖 · AI 与软件开发产业 · 2026-07Generation collapses, verification dams up · AI & software · Jul 2026
AI 把「打字」的成本压到趋近零——可软件的瓶颈从来不是打字。生成侧全面沦陷,账单却整张寄到了验证侧:那里正在筑起堰塞湖
AI drove the cost of «typing» toward zero — but software's bottleneck was never typing. The generation side collapses, and the whole invoice lands on the verification side, where a dam is rising
最小单元=「一次变更」(a diff):一个代码差异从意图出发,经评审测试、进入生产、被监控、可回滚,最终沉淀为组织能力。软件业的本质是「围绕变更的风险管理」,从来不只是代码生产——AI 精确命中的,只有「敲键盘」那一环。The atomic unit is «a change» (a diff): a code difference that travels from intent through review and tests, into production, monitored, revertible, finally settling as organizational capability. Software is «risk management around change», never mere code production — what AI hit with precision is only the «keystroke» step.
「写代码」与「交付软件」正在被重新定价:需求澄清、架构权衡、验证信任、事故责任——这四件本来就贵的事,现在更贵了。DORA 给 AI 的定性值得背下来:「放大器」——它加速健康的系统,也同样忠实地放大既有的失调。«Writing code» and «delivering software» are being re-priced: requirement clarification, architectural trade-offs, verification-trust, incident accountability — the four already-expensive things just got dearer. DORA's verdict on AI is worth memorizing: an «amplifier» — it speeds up healthy systems and just as faithfully amplifies existing dysfunction.
主脊是 SDLC「一次变更」八节点(需求→架构→编码→测试→评审→部署→监控→事故责任),每节点标注传统 vs AI + 吞噬度。另含加速鞭打诚实层、coding agent 四代跃迁、经验断层双轨、五块硬骨头、产品指南(2C 个人开发 / 2B 工程组织)。这是一张批判性行业解剖,不是 AI 颂歌。相邻议题见姊妹图:一人公司 / 独立开发者→startup、AI 生成代码的安全→security、初级岗位断层与培训→edu。
The spine is the eight-node «one change» SDLC (requirements → architecture → coding → testing → review → deploy → monitoring → incident accountability), each tagged traditional vs AI + how far AI has eaten it. Plus the Acceleration Whiplash honesty layer, the coding-agent four-generation leap, the experience-gap dual track, five hard bones, and a product guide (2C solo devs / 2B engineering orgs). A critical dissection, not an AI hymn. Adjacent topics on siblings: the one-person company / indie dev → startup, the security of AI-written code → security, the junior-role gap & training → edu.
传统节点Traditional
生成侧 · AI 吞噬Generation · eaten
验证侧 · 堰塞湖Verification · the dam
agent · 资本重构Agent · capital
加速鞭打 · 经验断层Whiplash · the gap
41–46%
生产环境新代码由 AI 生成(GitHub Copilot 平均 46%、Java 高达 61%,较 2022 的 27% 大涨;预计短期破 60%)——打字成本趋零的直接读数Of new production code, 41–46% is AI-generated (Copilot avg 46%, Java up to 61%, from 27% in 2022; expected to top 60% soon) — the direct readout of typing cost near zero
+441%
PR 评审中位时间上升(Faros「加速鞭打」,2.2 万开发者);同一份报告里生产事故 +242.7%、churn +861%、31.3% 的 PR 无人评审就合并——提速的代价整笔记在验证侧Median PR-review time rose (Faros «Acceleration Whiplash», 22k devs); in the same report, production incidents +242.7%, churn +861%, and 31.3% of PRs merged with no review — the whole cost of speed booked to verification
$2B ARR
Cursor(Anysphere)ARR,2025-01 的 1 亿 → 2026-02 的 20 亿美元——零到 20 亿约三年,有记录以来最快的 B2B 软件公司(D 轮估值 293 亿)Cursor's ARR, from $100M (Jan 2025) to $2B (Feb 2026) — zero to $2B in ~3 years, the fastest B2B software company on record (Series D at $29.3B)
−20%
美国 22–25 岁软件开发者就业,2025-09 较 2022 峰值下降近 20%(Stanford,论文口径)——职业阶梯的第一级正被抽掉。⚠️AI 归因仍有争议US software-developer employment for ages 22–25 fell ~20% by Sep 2025 vs the 2022 peak (Stanford, paper) — the career ladder's first rung is being pulled. ⚠️AI attribution is still contested
口径警告:本页是批判性行业分析,非工具选型 / 投资建议,刻意保留反方证据。AI 生成代码占比分三层读:公司 / 整体级 41–46%(A/D)、预测 60%(预测)、YC 极端团队 95%(非普遍,B/C)。采用率:DORA 90% 与 Stack Overflow 84% 双锚(均 A,样本 / 问法不同)。初级岗位降幅以 Stanford「22–25 岁 −20%」为主锚(论文),其余(−67% 招聘量、SignalFire −65% / −76% vs 2019、Indeed −49%)按基期 / 口径并列。AI 就业归因有争议:NBER(丹麦数据)发现 ChatGPT 采用两年后「精确的零效应」,中国官方多归因经济下行——本页拒绝「AI 单因论」。CodeRabbit「1.7× 重大问题 / 2.74× 漏洞率」是两个不同指标;Cursor「$60B 被 SpaceX/xAI 收购」为未证实传闻(D)。快速变动的营收 / 估值均带日期戳;厂商自述 / 预测打 D 级。每张卡片右上角 A/B/C/D=证据强度;产品指南非推荐背书。DORA 的可比性有断点:2025 年那份把标题从《Accelerate State of DevOps》改为《State of AI-assisted Software Development》,报告约 90% 受访者在开发中使用 AI;但它主要基于问卷与自报绩效,历史上作分析主轴的四项交付指标(交付频率、变更前置时间、变更失败率、失败恢复时间)在该版仅出现在脚注,作者还把 2024 年较谨慎的结论称作「DORA 2024 异常」。当一份基准在结论变化的同时更换了主指标,前后年份就不能直接比——可能是范式更新,也可能是标准漂移。
Basis warning: a critical industry analysis, not tool-selection / investment advice, deliberately keeping counter-evidence. Read AI-code share in three tiers: company/overall 41–46% (A/D), forecast 60% (forecast), YC extreme teams 95% (not typical, B/C). Adoption: DORA 90% and Stack Overflow 84% as dual anchors (both A, different samples/phrasing). Junior-role decline anchors on Stanford's «22–25, −20%» (paper); others (−67% postings, SignalFire −65%/−76% vs 2019, Indeed −49%) shown by base-year/basis. AI job-attribution is contested: NBER (Danish data) finds a «precise zero» two years after ChatGPT adoption; Chinese officials mostly cite the downturn — this page rejects an «AI-only» story. CodeRabbit's «1.7× major issues / 2.74× vuln rate» are two different metrics; Cursor's «$60B SpaceX/xAI acquisition» is an unverified rumor (D). Fast-moving revenue/valuation carry date stamps; vendor claims/forecasts are grade D. Each card's top-right A/B/C/D = evidence strength; the product guide is not an endorsement. DORA has a comparability break: the 2025 edition retitled itself from «Accelerate State of DevOps» to «State of AI-assisted Software Development» and reports that ~90% of respondents use AI in development; but it rests mainly on surveys and self-reported performance, and the four delivery metrics that were historically its analytical spine (deployment frequency, lead time, change failure rate, time to restore) appear only in a footnote there, while the authors label the more cautious 2024 findings the «DORA 2024 anomaly». When a benchmark changes its headline metric in the same year its conclusions change, the years stop being comparable — that can be a paradigm update, or drift.
◆ 诚实层 · 加速鞭打(最重)The honesty layer · the Acceleration Whiplash
一个 diff,两侧命运:生成变便宜,验证变贵One diff, two fates: generation gets cheap, verification gets dear
AI「吞噬」SDLC 的方式,不是从左到右匀速推进,而是双侧分化:生成侧全面沦陷,验证侧筑起堰塞湖。把一次变更画成一个 diff 看得最清楚——绿色的「+」全被 AI 接管了,红色的「−」反而涨价:新的价值中心、新的瓶颈、新的预算,全在那一侧。AI doesn't «eat» the SDLC left-to-right at a steady pace; it splits it in two: the generation side collapses while the verification side dams up. Draw one change as a diff and it's plain — the green «+» has been taken over wholesale, while the red «−» is repricing upward: the new value center, bottleneck and budget all live on that side.
a/software-economics.diff一次变更 · 重新定价one change · re-priced
+生成侧(成本趋零)generation (cost → 0):打字 · 样板 Boilerplate · CRUD · 脚手架 · 胶水代码 · 回归 / 单元测试 · 简单修复 · 文档生成typing · boilerplate · CRUD · scaffolding · glue code · regression/unit tests · simple fixes · docs
−验证侧(反而更贵)verification (gets dearer):需求正确 · 架构权衡 · 系统稳定 · 责任可追 · 成本可控 —— 新瓶颈 = 新价值 = 新岗位 = 新预算correct requirements · architectural trade-offs · system stability · traceable accountability · cost control — new bottleneck = new value = new budget
两点关键修正,防止把「吞噬」读成单调线性:① 需求定义不是被吞噬,而是被抬升为最贵、最高杠杆的环节;② 评审是矛盾节点——AI 能评审,但它同时把待评审代码量放大十倍,评审反而成为吞吐瓶颈。Two corrections against a monotonic reading: ① requirement definition isn't eaten but elevated to the dearest, highest-leverage step; ② review is the paradox node — AI can review, yet it multiplies the code awaiting review tenfold, so review becomes the throughput bottleneck.
▚ 加速鞭打 · Faros AI《2026 工程报告》(2.2 万开发者 / 4000+ 团队)The whiplash · Faros AI 2026 (22k devs / 4000+ teams)
生成提速的代价,全都记在验证侧The cost of faster generation lands entirely on verification
+441%
PR 评审中位时间median PR-review time
+242%
每个 PR 引发的生产事故production incidents per PR
+861%
代码 churn(提交后即被删 / 覆盖)code churn (soon deleted/overwritten)
+54%
人均 bugbugs per developer
31.3%
PR 完全无人评审就合并of PRs merged with no review
交叉印证,四个来源指向同一条鞭子:DORA 2025(约 5000 人)AI 采用率 90%(+14pct)、65% 高度依赖,但 30% 几乎不信任——AI 改善几乎每个维度,唯独软件交付稳定性下降。Stack Overflow 2025:66% 把「几乎正确但不完全对」列为主要挫败,45% 认为调试 AI 代码更花时间。GitClear(2.11 亿行):重构比例 25%→不足 10%、重复率飙约四倍。METR:资深开源开发者用 AI 实际慢 19%,却自认快 24%(43pct 感知偏差)。Four sources, one whip: DORA 2025 (~5000): 90% adoption (+14pt), 65% heavily rely, yet 30% barely trust it — AI improves nearly every dimension except delivery stability. Stack Overflow 2025: 66% cite «almost right but not quite» as a top frustration, 45% say debugging AI code takes longer. GitClear (211M lines): refactoring 25%→under 10%, duplication up ~4×. METR: experienced OSS devs were actually 19% slower with AI while feeling 24% faster (a 43-point perception gap).
≋ 奈奎斯特稳定判据 · AI 是放大器,不是解药The Nyquist criterion · AI is an amplifier, not a cure
DORA 2025 借控制论的奈奎斯特稳定判据作比:任何控制系统,必须以至少两倍于被控系统的速度运行。「生成」提速十倍而「验证 / 控制」原地踏步,系统失控不是风险,是定理——加速鞭打的物理学就这一句。另一条同样硬的证据来自 Anthropic 约 40 万个 Claude Code 会话:人仍主导「做什么(what)」的规划,AI 主导「怎么做(how)」的执行——问题定义权,AI 还没拿走。DORA 2025 borrows control theory's Nyquist stability criterion: any control system must run at at least twice the speed of the system it controls. Speed generation up tenfold while verification/control stands still, and instability isn't a risk — it's a theorem. That one sentence is the physics of the whiplash. The other equally hard evidence comes from Anthropic's ~400,000 Claude Code sessions: humans still drive the «what» (planning) while AI drives the «how» (execution) — the problem-definition rights, AI has not taken.
Reading the MapReading the Map
从这张图看到的五条规律Five patterns this map makes visible
立场声明:本页是批判性、祛魅的行业结构分析——用 A–D 角标区分硬数据与传闻,并刻意保留「AI 就业归因有争议」的反方。不美化、不唱衰,不构成工具选型或投资建议;产品指南是市场地图而非推荐背书。营收 / 估值属快速变动领域,均带日期戳,请以最新披露为准。
Stance: a critical, demystifying structural analysis — A–D badges separate hard data from rumor, and the «AI job-attribution is contested» counter-view is deliberately kept. Nothing glamorized or doom-mongered; not tool-selection or investment advice; the product guide is a market map, not an endorsement. Revenue/valuation move fast and carry date stamps; defer to the latest disclosures.