C/PAdvanced-compute evidence ledgerWhen Controls Raise the Cost

Evidence ledger / 19 records

Every claim should survive inspection.

Each card records the proposition, the best available public source, the exact place to look, the strongest qualification, and what remains unknown.

Confidence describes the scoped claim. “Primary” means original or official for that claim—not independent verification.

  1. EV-01 · ev-deepseek-compute
    SupportsHigh confidence

    Claim

    DeepSeek reports training a 671B-parameter MoE on 14.8 trillion tokens with 2.788 million H800 GPU-hours.

    Technical preprint · PrimaryDeepSeek-V3 Technical ReportDeepSeek-AI
    Published
    2024-12-27
    Accessed
    2026-09-03
    Location
    Abstract; section 1; Table 1

    Quotation or data

    “2.788M H800 GPU hours.”

    Counterevidence

    The run and benchmarks are author-reported and have not been independently reproduced end to end.

    Remaining uncertainty

    The public report does not establish acquisition dates, electricity use, or total program cost.

  2. EV-02 · ev-deepseek-cost
    QualifiesHigh confidence

    Claim

    The widely repeated $5.576 million figure is a rental-equivalent training-compute estimate, not an all-in development cost.

    Technical preprint · PrimaryDeepSeek-V3 Technical ReportDeepSeek-AI
    Published
    2024-12-27
    Accessed
    2026-09-03
    Location
    Section 1, note immediately below Table 1

    Quotation or data

    Prior research and ablation costs “are not included.”

    Counterevidence

    The narrow figure remains useful for comparing the disclosed final run if its scope is stated.

    Remaining uncertainty

    R&D, acquisition, depreciation, staff, data, networking, and failed-run costs are undisclosed.

  3. EV-03 · ev-deepseek-independent
    SupportsModerate confidence

    Claim

    On METR’s autonomy suite, DeepSeek-V3 was comparable to Claude 3.5 Sonnet (Old), while trailing newer frontier models.

    Independent evaluation · Independent / secondaryDetails about METR’s Preliminary Evaluation of DeepSeek-V3METR
    Published
    2025-02-12
    Accessed
    2026-09-03
    Location
    Opening findings; General Autonomous Capabilities

    Quotation or data

    METR reported comparable autonomy performance to Claude 3.5 Sonnet (Old), with weaker results than newer frontier models.

    Counterevidence

    METR used basic elicitation and warned that looping and other failures might understate or distort capability.

    Remaining uncertainty

    Benchmark performance is task- and elicitation-dependent and does not establish frontier parity.

  4. EV-04 · ev-deepseek-dependency
    SupportsHigh confidence

    Claim

    Nvidia identified H800 as a product covered by the October 2023 China licensing restrictions.

    Regulatory filing · PrimaryNVIDIA Corporation Form 10-K for Fiscal 2024NVIDIA Corporation / U.S. SEC
    Published
    2024-02-21
    Accessed
    2026-09-03
    Location
    Item 1, Government Regulations—Global Trade

    Quotation or data

    Restrictions on H800 shipments took effect on 23 October 2023.

    Counterevidence

    DeepSeek may have acquired its H800 fleet before the restriction took effect; acquisition timing is not public.

    Remaining uncertainty

    The size and provenance of DeepSeek’s remaining controlled-hardware inventory are unknown.

  5. EV-05 · ev-deepseek-repro
    QualifiesHigh confidence

    Claim

    The public repository enables inference and inspection of model weights but not reproduction of the full training run.

    Company repository · PrimaryDeepSeek-V3 Official RepositoryDeepSeek-AI
    Published
    2024-12-26
    Accessed
    2026-09-03
    Location
    Repository README, model downloads and inference sections

    Quotation or data

    The release includes weights and inference examples for Nvidia, AMD, and Ascend.

    Counterevidence

    Open weights make broad deployment and independent downstream testing possible.

    Remaining uncertainty

    Training data, the full HAI-LLM stack, and exact full-run configuration are unavailable.

  6. EV-06 · ev-pangu-training
    SupportsModerate confidence

    Claim

    Huawei reports a 718B-parameter MoE run on 6,000 Ascend 910B NPUs and says its system supports all training stages.

    Technical preprint · PrimaryPangu Ultra MoE: How to Train Your Big MoE on Ascend NPUsHuawei
    Published
    2025-05-07
    Accessed
    2026-09-03
    Location
    Paper abstract; sections 1 and 2.2

    Quotation or data

    “MFU of 30.0%” on 6,000 Ascend NPUs.

    Counterevidence

    The run is vendor-reported and has no independent full-training audit or reproduction.

    Remaining uncertainty

    Run duration, energy, capital cost, fleet availability, and yield are not disclosed.

  7. EV-07 · ev-pangu-codesign
    QualifiesHigh confidence

    Claim

    The model and execution plan were reshaped around NPU constraints, with cumulative interventions raising reported throughput 58.7 percent.

    Technical preprint · PrimaryPangu Ultra MoE: How to Train Your Big MoE on Ascend NPUsHuawei
    Published
    2025-05-07
    Accessed
    2026-09-03
    Location
    Section 4.5, Table 5

    Quotation or data

    6,000 NPUs; 30% MFU; hierarchical all-to-all; recomputation and host swapping.

    Counterevidence

    Thirty percent MFU still implies substantial unused peak compute, and no comparable counterfactual run is public.

    Remaining uncertainty

    The incremental burden caused specifically by controls cannot be separated from normal large-model engineering.

  8. EV-08 · ev-pangu-weights
    SupportsHigh confidence

    Claim

    The official model card says Pangu Ultra was trained from scratch; the repository publishes roughly 1.47 TB of weights and documents a minimum 32-card Ascend serving setup.

    Company repository · PrimaryopenPangu Ultra MoE 718B ModelHuawei openPangu
    Published
    Date not verified
    Accessed
    2026-09-03
    Location
    Model card overview and sections 1, 4.1, and 4.4; file tree

    Quotation or data

    62 weight shards; four Atlas 800T A2 nodes; 32 64GB cards.

    Counterevidence

    Serving reproducibility does not reproduce pretraining or validate the disclosed training fleet.

    Remaining uncertainty

    The complete training corpus, code, and run configuration are not public.

  9. EV-09 · ev-pangu-upstream
    QualifiesModerate confidence

    Claim

    BIS assessed that listed Ascend 910-series chips likely involved U.S. design technology, software, or equipment subject to export rules.

    Government guidance · PrimaryGuidance on Application of General Prohibition 10 (GP10) to People’s Republic of China (PRC) Advanced-Computing Integrated Circuits (ICs)U.S. Bureau of Industry and Security
    Published
    2025-05-13
    Accessed
    2026-09-03
    Location
    Pages 1–2, summary, illustrative list, and applicable regulations

    Quotation or data

    The illustrative list names Ascend 910B and 910C.

    Counterevidence

    This is an enforcement assessment, not a bill-of-materials audit of the Pangu fleet or CloudMatrix installation.

    Remaining uncertainty

    The precise origin and legal status of each chip and upstream component are not public.

  10. EV-10 · ev-cloudmatrix-architecture
    SupportsHigh confidence

    Claim

    CloudMatrix384 connects 384 Ascend NPUs and 192 Kunpeng CPUs through an all-to-all Unified Bus architecture.

    Technical preprint · PrimaryServing Large Language Models on Huawei CloudMatrix384Huawei and SiliconFlow
    Published
    2025-06-19
    Accessed
    2026-09-03
    Location
    Abstract; sections 3.2–3.4

    Quotation or data

    384 NPUs, 192 CPUs, disaggregated prefill, decode, and caching.

    Counterevidence

    The paper is authored by Huawei and SiliconFlow researchers and is not peer reviewed.

    Remaining uncertainty

    Public fleet volume and production availability outside the disclosed deployment are unknown.

  11. EV-11 · ev-cloudmatrix-performance
    SupportsModerate confidence

    Claim

    The disclosed DeepSeek-R1 test achieved 1,943 decode tokens per second per NPU at 49.4 ms, below one profiled H800 result but above another published H800 baseline.

    Technical preprint · PrimaryServing Large Language Models on Huawei CloudMatrix384Huawei and SiliconFlow
    Published
    2025-06-19
    Accessed
    2026-09-03
    Location
    Section 5.2, Tables 2–4

    Quotation or data

    1,943 tokens/s/NPU at 49.4 ms; reported efficiency 1.29 tokens/s/TFLOPS.

    Counterevidence

    Tests used 256 of 384 NPUs; decode assumed a 70% MTP acceptance rate; batch sizes and implementations differed; and the highest prefill result assumed idealized expert load balancing.

    Remaining uncertainty

    Like-for-like cost, power, reliability, and hardware-count comparisons are unavailable.

  12. EV-12 · ev-cloudmatrix-limits
    QualifiesHigh confidence

    Claim

    Tighter latency requirements sharply reduced reported per-NPU throughput, exposing the workload sensitivity of the comparison.

    Technical preprint · PrimaryServing Large Language Models on Huawei CloudMatrix384Huawei and SiliconFlow
    Published
    2025-06-19
    Accessed
    2026-09-03
    Location
    Section 5.2, Table 4

    Quotation or data

    Throughput fell from 1,943 to 538 tokens/s/NPU as the target tightened from about 50 ms to 15 ms.

    Counterevidence

    The system still met the stated tighter latency target and maintained useful throughput.

    Remaining uncertainty

    Real traffic mixes, uptime, and long-run service cost are not independently reported.

  13. EV-13 · ev-cloudmatrix-power
    QualifiesLow confidence

    Claim

    A specialist estimate argues that matching a GB200 NVL72-class system requires substantially more Ascend devices, optical components, and power.

    Specialist analysis · Independent / secondaryHuawei AI CloudMatrix 384: China’s Answer to Nvidia GB200 NVL72SemiAnalysis
    Published
    2025-04-16
    Accessed
    2026-09-03
    Location
    Opening comparison; CloudMatrix 384 System Architecture

    Quotation or data

    Estimated 4.1 times the system power of a GB200 NVL72.

    Counterevidence

    The estimate is not metered or audited, and the compared systems differ in chip count and configuration.

    Remaining uncertainty

    No public independent energy, cooling, or total-cost measurement validates the exact multiplier.

  14. EV-14 · ev-gatekeeper-volume
    SupportsHigh confidence

    Claim

    A company and its owner pleaded guilty in a scheme involving at least $160 million of exported and attempted H100 and H200 shipments.

    Government enforcement · PrimaryU.S. Authorities Shut Down Major China-Linked AI Tech Smuggling NetworkU.S. Department of Justice
    Published
    2025-12-08
    Accessed
    2026-09-03
    Location
    Release summary; paragraph beginning ‘According to court documents’

    Quotation or data

    Conduct spanned October 2024 to May 2025; more than $50 million was seized.

    Counterevidence

    The $160 million figure combines completed and attempted exports; as of the 8 December 2025 release, not every person described had been convicted.

    Remaining uncertainty

    The number of chips delivered, ultimate recipients, and resulting AI tasks are not disclosed.

  15. EV-15 · ev-gatekeeper-method
    SupportsModerate confidence

    Claim

    The public record describes straw purchasers, third-country customers, relabeling, false paperwork, and warehousing as elements of the diversion route.

    Government enforcement · PrimaryU.S. Authorities Shut Down Major China-Linked AI Tech Smuggling NetworkU.S. Department of Justice
    Published
    2025-12-08
    Accessed
    2026-09-03
    Location
    Paragraphs beginning ‘According to charging documents’ and complaint allegations

    Quotation or data

    Labels were removed and replaced with a fictitious company name.

    Counterevidence

    Several details come from charging documents and, as of the 8 December 2025 release, remained allegations against defendants who had not pleaded guilty.

    Remaining uncertainty

    The release does not quantify how common comparable networks are.

  16. EV-16 · ev-bis-diversion-guidance
    QualifiesHigh confidence

    Claim

    BIS identified concrete checks that could expose transshipment, opaque customers, foreign IaaS access, and implausible data-center capacity.

    Government guidance · PrimaryIndustry Guidance to Prevent Diversion of Advanced Computing Integrated CircuitsU.S. Bureau of Industry and Security
    Published
    2025-05-13
    Accessed
    2026-09-03
    Location
    Pages 1–5, red flags and due-diligence actions

    Quotation or data

    Red flags include freight forwarders, undisclosed end users, foreign IaaS, and infrastructure inconsistent with ordered chips.

    Counterevidence

    Guidance does not establish that every recommended check is implemented or effective in every jurisdiction.

    Remaining uncertainty

    Detection rates and cross-border enforcement capacity are not public.

  17. EV-17 · ev-state-tax
    SupportsHigh confidence

    Claim

    China continued national tax-preference eligibility for integrated-circuit production, design, materials, packaging, and major projects in 2024.

    Government policy · Primary2024 Integrated-Circuit and Software Tax-Preference Eligibility NoticeNational Development and Reform Commission of China and partner ministries
    Published
    2024-03-22
    Accessed
    2026-09-03
    Location
    Article 1 and attached eligibility criteria

    Quotation or data

    The eligibility notice covers production at 28 nm and below and multiple upstream inputs and services.

    Counterevidence

    Eligibility rules do not disclose recipients, award value, additionality, or a capability outcome.

    Remaining uncertainty

    No firm-level causal link to a reviewed AI-chip workaround is established.

  18. EV-18 · ev-state-hangzhou
    QualifiesModerate confidence

    Claim

    A 2025 Hangzhou policy authorized annual compute vouchers and large-model training support, but no reviewed award record ties that support to DeepSeek-V3.

    Government policy · Primary杭州市人民政府办公厅关于促进杭州市人工智能产业高质量发展若干措施的通知 (Hangzhou measures supporting high-quality AI-industry development; translated title)Hangzhou Municipal People’s Government
    Published
    2025-06-21
    Accessed
    2026-09-03
    Location
    Measures on compute services and large-model training support

    Quotation or data

    Annual compute-voucher ceiling: CNY 250 million; qualifying model-training support: up to CNY 50 million.

    Counterevidence

    The policy postdates the DeepSeek-V3 release and does not prove DeepSeek received a grant.

    Remaining uncertainty

    Recipients, disbursements, hardware origin, and capability produced are not public in the reviewed record.

  19. EV-19 · ev-state-stockpile-caveat
    QualifiesHigh confidence

    Claim

    GAO reports that BIS used interim-final rules partly to reduce pre-rule stockpiling, but this does not establish the scale of any Chinese inventory response.

    Government audit · PrimaryExport Controls: Commerce Implemented Advanced Semiconductor Rules and Took Steps to Address Compliance ChallengesU.S. Government Accountability Office
    Published
    2024-12-02
    Accessed
    2026-09-03
    Location
    Highlights; What GAO Found, paragraphs 1–2

    Quotation or data

    BIS cited avoiding stockpiling as one reason for immediate enforcement.

    Counterevidence

    A policy concern is not evidence of actual inventory, intent, ownership, or use by a named Chinese developer.

    Remaining uncertainty

    Firm-level pre-control and legacy inventories remain largely undisclosed.