Research window: Past 24 hours (2026-09-14 10:18 ~ 2026-09-15 10:18, Beijing time) Sources: GitHub (org: flagos-ai, full verification of pushed_at across 53 repositories + single commit search of 97 entries verified by committer-date + per-repo commits review with 1 additional entry + releases/tags metadata for each repository + direct raw retrieval of the community repository’s 2.2 release manifest and release schedule + commit patch details), Google News RSS (21 sets of Chinese and English query terms, via proxy), HN Algolia, Xinhua Net, Guandian, Sina Finance, Cailian Press, STAR Market Daily, Huanqiu, and other industry reports, plus supplementary Tavily searches (see appendix source list for details)


In This Issue

  • Today’s Focus: FlagOS 2.2 RC2 full-stack synchronized tagging — 25-item manifest re-versioned, 25 rc2.post1 tags landed within 10 minutes (09-14)
  • I. Open Source Project Progress (GitHub Activity)
    • 1.1 2.2 RC2 manifest submitted: 25 items covering all four layers L0–L3, versions uniformly aligned to rc0/rc1 (09-14)
    • 1.2 23 components tagged rc2.post1 within 10 minutes: the second full-stack snapshot of this release cycle (09-14)
    • 1.3 First rc2.post1 Releases: FlagGems v5.4.0 and FlagSparse v0.3.0 (09-14)
    • 1.4 FlagTree: AMD decoupled compilation, XPU fixes, new xpu3.6 dedicated tag (09-14/09-15)
    • 1.5 FlagCX: libflagcx.so into wheel, T-Head PPU backend and CI connected bidirectionally (09-14/09-15)
    • 1.6 FlagAttention v0.4.0: four backends — Enflame, Hygon, MetaX, Moore Threads — landed in one shot (09-14)
    • 1.7 FlagGems-Experimental: 17 Hygon operators merged in bulk, plus 4 Moore Threads and 1 Kunlunxin (09-14)
    • 1.8 FlagGems mainline: Enflame sync, Triton 3.5 compatibility, and GitCode mirror (09-14/09-15)
    • 1.9 FlagGems-vllm and inference plugin line: moe_sum dual-backend, MTT S5000 topk, vLLM 0.24.0 (09-14/09-15)
    • 1.10 FlagQuantum: Twin API frozen, QPU digital twin enters historical sequence management (09-14)
    • 1.11 Frameworks and tooling: FlagScale unified CI, Torch-FL backend routing alignment, TE-FL Hygon interface (09-14/09-15)
    • 1.12 FlagGems-sglang: competition outputs merged in bulk, 10 PRs into main repo same day (09-14)
    • 1.13 build-infra: Hygon vLLM rc2.post1 dual-image tags recorded and re-verified (09-14)
    • 1.14 docs and release-info: documentation line merged into CICD overhaul, release site synchronized (09-15)
  • II. News Coverage and Ecosystem
    • 2.1 Open3D-PIMC open-sourced at China Computing Conference: jointly developed by Tsingmicro and BAAI, completing the 3D computing chip compilation stack (09-14)
    • 2.2 Eleventh consecutive quiet window for component-level search: core hits concentrated in BAAI-side coverage (09-14~09-15)
    • 2.3 Three industry conference threads: domestic computing narratives at CIFTIS, Bund Summit, and China Computing Conference (09-09~09-14)
  • III. Member Deep Dives
    • 3.1 Tsingmicro: dual narrative of Open3D-PIMC and 5000P reconfigurable computing (09-14)
    • 3.2 Hygon: from 17 operators to rc2.post1 image re-verification, top engineering output this window (09-14)
    • 3.3 Moore Threads: attention backend, profiler integration, and CIFTIS market-share framing (09-14)
    • 3.4 MetaX: metax CI and vLLM 0.24.0, Xiyun C700 on the eve of tape-out (09-14)
    • 3.5 Enflame: enflame backend continues merging, post-IPO market cap and customer structure narrative (09-14)
    • 3.6 Iluvatar CoreX: iluvatar backend naming unified, capital markets enter pre-unlock window (09-14)
    • 3.7 Kunlunxin and DAMO XuanTie (external ecosystem players): XPU tags and PPU dual-line integration (09-14/09-15)
    • 3.8 BAAI (lead organization): RC2 governance discipline and rule-setting authority framing via Open3D-PIMC (09-14)
  • IV. Summary and Trend Observations
  • Appendix: Complete Source List

Today’s Highlight: FlagOS 2.2 RC2 Full-Stack Synchronized Tagging — 25-Item Manifest Reissued, 25 rc2.post1 Tags Landed Within 10 Minutes

Date: 2026-09-14 Source: community commit #109 (2.2 RC2 manifest), 2.2 release schedule, FlagGems v5.4.0-rc2.post1

The most significant event in this window was a coordinated action: at 10:35 on 09-14, the community repo merged release/2.2/release-2.2-rc2.yaml (251 lines added, submitted by the Release Manager), announcing that the second release candidate manifest of the 2.2 cycle, RC2, had taken shape; then, within ten minutes from 10:45 to 10:54, 23 component repos across the org successively pushed v{X}-rc2.post1 tags or Releases, from 12:18 to 12:19 the three FlagTree Triton variant tags were filled in, and at 19:59 FlagTree added one more xpu3.6 tag. Including vllm-plugin-FL’s dual version lines, a total of 26 tagging actions covering 22 repos were recorded in this window.

The RC2 manifest is organized into four layers: L0 infrastructure layer (FlagTree’s triton3.6/3.5/3.3 three entries, FlagCX); L1 compute library layer (FlagGems, FlagFFT, FlagSparse, FlagDNN, FlagBLAS, FlagTensor, FlagAudio, FlagAttention, FlagGems-vllm, FlagGems-sglang); L2 framework layer (Torch-FL, vllm-plugin-FL and vllm-plugin-fl-0.2, sglang-plugin-FL, TransformerEngine-FL, Megatron-LM-FL, FlagOS-Compressor); L3 application layer (FlagScale, KernelGen, KernelGenBench, FlagRelease). A total of 25 entries, all pointing to verified rc2.post1 snapshots.

Compared with RC1, the nature of this change is “same structure, new snapshots”: module version numbers remain consistent with rc0/rc1, advancing only in the .postN iteration position — FlagGems from v5.4.0-rc1.post2 to v5.4.0-rc2.post1, FlagCX from v0.14.0-rc1.post2 to v0.14.0-rc2.post1, FlagAttention fixed on the v0.4.0 line, FlagTree from 0.7.0rc1.post1 to 0.7.0rc2.post1. The manifest header spells out the workflow as executable rules: pull rc2 integration branches from each module’s baseline branch, and after verification passes, push an incrementing tag (rc2.post1 → rc2.post2), recording the previous round’s tag in the entry’s rc1 comment; the two fields 分支: and 基线分支: are read directly by manage-release.py, so when the same repo is split into multiple entries each tracking a different upstream line (FlagTree, vllm-plugin-FL), and when a release branch rather than the default branch must be pulled from (flagos-compressor), these must be explicitly stated.

Placed back in the schedule, this tagging occurred in the middle of the “testing and stabilization period (09-01 ~ 09-24)”, 14 days before the 2.2 GA on 09-28. The significance of RC2 lies not in new features — feature freeze closed as early as 08-31 — but in solidifying the fixes accumulated during the testing period into a set of locked versions that are reproducible and batch-pullable, giving the multi-chip acceptance matrix a unified comparison baseline.


1. Open Source Project Progress (GitHub Activity)

Window Overview: Of the 53 repos in the org, 29 had pushes during the window; a single commit search returned 97 in-window commits (4 pages, sorted by committer-date descending), and a per-repo commits recheck added 1 more (TransformerEngine-FL), for a total of 98 commits across 17 repos: FlagQuantum 17, FlagGems-Experimental 20, FlagGems 12, FlagGems-sglang 10, FlagTree 6, FlagCX 5, FlagGems-vllm 5, build-infra 5, FlagAttention 4, docs 4, Torch-FL 2, sglang-plugin-FL 2, vllm-plugin-FL 2, FlagScale 1, FlagPrism 1, community 1, TransformerEngine-FL 1. In addition, Megatron-LM-FL, KernelGen, FlagBLAS, FlagSparse, FlagDNN, FlagOS-Compressor, FlagFFT, FlagTensor, KernelGenBench, FlagRelease, FlagAudio, release-info and other repos had pushes during the window, but their in-window new commits landed on tags or non-default branches (most of which are the rc2.post1 tagging described in section 1.2 below). No new repos.

This window’s shape = “full-stack snapshot turnover” + “external backends stepping up,” advancing on two fronts simultaneously: On the governance side, a single manifest commit and a ten-minute collective tagging pushed 2.2 to its second release candidate; on the engineering side, Hygon, Moore Threads, MetaX, Enflame, Kunlunxin, and DAMO XuanTie stepped up simultaneously across four fronts—operators, attention backends, communication libraries, and compilers—with Hygon’s batch of 17 operators landing in the repo and FlagAttention’s four backends landing on the same day being the two most concentrated pieces of engineering work in this window.

1.1 2.2 RC2 Manifest Commit: 25 entries covering L0–L3, versions all aligned to rc0/rc1 (09-14)

Date: 2026-09-14 Source: community #109

The Release Manager added release/2.2/release-2.2-rc2.yaml (+251 lines) on 09-14 at 10:35, with the commit message stating “cut rc2 branches from each module’s baseline branch, tag v{X}-rc2.post1 per the workflow at the top of the manifest; versions kept consistent with rc0/rc1, with each entry’s rc1 comment recording the previously verified tag.” The manifest is organized into four layers—L0 infrastructure, L1 compute libraries, L2 frameworks, L3 applications—totaling 25 entries, and notes that the default branches of FlagGems, FlagGems-sglang, FlagDNN, and FlagBLAS are master rather than main. This is the second manifest snapshot after rc1, and the only intermediate candidate before 2.2 GA (09-28).

1.2 23 components tagged rc2.post1 within 10 minutes: the second full-stack snapshot of this release cycle (09-14)

Date: 2026-09-14 Source: FlagCX v0.14.0-rc2.post1

The tagging cadence was highly concentrated: at 10:45 FlagCX, FlagGems, and FlagFFT went first; at 10:47 FlagDNN and FlagBLAS; at 10:48 seven in a row—FlagTensor, FlagSparse, FlagAudio, FlagAttention, FlagGems-vllm, FlagGems-sglang, and Torch-FL; at 10:49 vllm-plugin-FL (both the 0.3.0 and 0.2.2 lines), sglang-plugin-FL, and TransformerEngine-FL; at 10:51 Megatron-LM-FL and FlagOS-Compressor; at 10:53 FlagScale, KernelGen, and KernelGenBench; at 10:54 FlagRelease. FlagTree’s three entries—triton3.5/3.6/3.3—were filled in between 12:18 and 12:19. Completing the tagging of 23 components within ten minutes shows that this “manifest first, script execution” release process is now stable and reusable.

1.3 First rc2.post1 Releases: FlagGems v5.4.0 and FlagSparse v0.3.0 (09-14)

Date: 2026-09-14 Source: FlagSparse v0.3.0-rc2.post1

Of the 26 taggings, most landed as tags; only FlagGems (v5.4.0-rc2.post1) and FlagSparse (v0.3.0-rc2.post1) generated GitHub Releases with release notes, both published between 10:45 and 10:48 on 09-14. Comparing against FlagGems’ historical cadence: v5.4.0-rc0.post1 (09-01), v5.4.0-rc1.post1 (09-07), v5.4.0-rc1.post2 (09-11), v5.4.0-rc2.post1 (09-14), with the stable release v5.3.6 stalled at 09-11—showing that 5.4.0’s candidate iteration has reached a density of more than two rounds per week.

1.4 FlagTree: AMD decoupled compilation, XPU fixes, new xpu3.6-specific tag (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagTree #1163 (AMD independent compilation)

FlagTree had 6 commits and added 4 tags during the window. Four engineering threads: first, [Build][AMD] Decouple amd for independent compilation v2 (#1163) decouples the AMD backend from the unified build, allowing independent compilation; second, the XPU thread fixes the XPUDriver device-fetching path, resolving cuda-graph capture reporting err -900 (#1168); third, on the CI side, it connects nvidia3.6’s FlagGems baseline and test jobs to a remote reusable workflow (#1172, 09-15 10:04) and removes the USE_FLAGCX variable from the workflow (#1167); fourth, it integrates FlagPrism’s profiler and debugger into the compiler chain (#1106, Moore Threads direction). On the tag side, besides the three rc2.post1 tags, 0.7.0rc2+xpu3.6 was added at 19:59—the first time a non-Triton accelerator suffix has appeared in this repo’s tag scheme (previously only the +triton3.x form), corresponding to the Kunlunxin XPU line being tagged separately outside the delivery workflow.

1.5 FlagCX: libflagcx.so into the wheel, T-Head PPU backend and CI connected both ways (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagCX #596 (PPU torch plugin and CI)

The communication library had 5 commits this window, divisible into two groups. Delivery side: [UIL] Place libflagcx.so into the flagcx wheel package (#570) packs the shared library into the wheel, continuing the previous window’s direction of “moving from source builds toward installability”; [UIL] Rename the iluvatar_corex build key to iluvatar (#583) unifies Iluvatar CoreX’s build key to the vendor name, reducing special cases in backend naming; plus a dependabot upgrade (#572). New backend side: [CICD] Add T-Head PPU CI workflow (#592) establishes CI first, followed by [UIL] Add PPU torch plugin support and extend PPU CI workflow (#596, 09-15 00:49) adding torch plugin support and extending the workflow—a two-step approach consistent with previous vendor backends. Notably, the PPU integration target is DAMO XuanTie (T-Head), which is not on the existing member list and represents a newly added external ecosystem party.

1.6 FlagAttention v0.4.0: Enflame, Hygon, MetaX, and Moore Threads backends land all at once (09-14)

Date: 2026-09-14 Source: FlagAttention #61 (mthread backend)

FlagAttention merged four isomorphic commits within two minutes, from 14:29 to 14:30, adding Enflame (#58), Hygon (#59), MetaX (#60), and Moore Threads (#61) backends to the attention operators in sequence. Four vendors entering the same module in the same batch and the same minute shows that attention operator backend integration has been abstracted into a batch-replicable template—corroborated by the repo’s version line jumping to v0.4.0-rc2.post1 in this window, as attention operators move from “single-platform implementation” into a “multi-vendor parallel maintenance” phase.

1.7 FlagGems-Experimental: Hygon’s 17 operators land in batch, Moore Threads 4, Kunlunxin 1 (09-14)

Date: 2026-09-14 Source: FlagGems-Experimental #557 (Hygon ormqr)

The top repo by commit count this window (20). The Hygon line merged 17 operators consecutively within one hour, from 17:51 to 18:54, covering special functions (special_legendre_polynomial_p, special_hermite_polynomial_h, special_round_out), linear algebra (linalg_ldl_factor, ormqr), and numerical/indexing operators (norm_scalaropt_dim, masked_scale, scalar_tensor, gcd_, addmm_, as_strided_scatter, scatter_add, feature_dropout, conv_depthwise2d, beam_search_score, thnn_fused_lstm_cell), plus a Hygon _scaled-related fix; the Moore Threads line added cholesky_inverse, linalg_ldl_solve, ormqr, and linear operator specializations; the Kunlunxin line added an mvlgamma specialization (#605). This dense backfilling pattern indicates these operators were previously missing on the corresponding platforms—gaps exposed by the acceptance matrix being filled.

1.8 FlagGems mainline: Enflame sync, Triton 3.5 compatibility, and GitCode mirror (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagGems #6200 (enflame sync)

Of the mainline’s 12 commits, three are most worth recording: first, enflame to flagos 20260911 (#6200), where Enflame syncs its branch’s changes back to the mainline in batches—an instance of periodic merging between member branches and the upstream mainline; second, fix: add Triton 3.5 compatibility for flash_attention_backward (#6252, corresponding to issue #6235), adding compatibility for the attention backward operator under the new Triton version, echoing FlagTree’s multi-Triton-version parallel strategy; third, [User Experience] pre commit config use gitcode mirror (#6231), switching pre-commit’s dependency source to the GitCode mirror—an infrastructure adaptation for domestic developers. The rest are bug fixes (fft res_out, boolean slice views, batch_norm benchmark naming) and KernelGen-side operator additions (Nvidia __iand__, Moore Threads linear, Ascend master_scatter_backward), KernelGen test discovery, and an alias_copy.out registration fix (09-15 09:19/09:20).

1.9 FlagGems-vllm and the inference plugin line: moe_sum dual backends, MTT S5000 topk, vLLM 0.24.0 (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagGems-vllm #758 (MTT S5000 persistent_topk)

Inference-side activity this window centered on the “make a given backend usable” layer. FlagGems-vllm: Hygon added moe_sum (#745), and DAMO XuanTie added a moe_sum variant the same day (#772), the two forming contrasting implementations of the same operator on different platforms; the KMCompiler line switched top_k_per_row’s prefill/decode to the FlagGems suite and fixed two vendor-blocking defects, added an MTT prefill, then added a persistent_topk backend for Moore Threads’ MTT S5000 (#758, 09-15 09:56). Plugin side: vllm-plugin-FL enabled metax CI and bumped vLLM to 0.24.0 (#488), plus a batch of adaptation admission test cases (#492); sglang-plugin-FL decoupled dispatch unit tests from platform configuration (#94) and added a benchmark for plugin preloading (#96).

1.10 FlagQuantum: Twin API frozen, QPU digital twin enters historical sequence management (09-14)

Date: 2026-09-14 Source: FlagQuantum #31 (Twin v1 API freeze)

The second-highest commit-count repo (17), all concentrated within a single day from 10:29 to 22:34, and forming a single-purpose unidirectional chain: first generate identity-bound Twin evidence (#27), then do duplicate-verification aggregation (#28), next land the verification sequence (#29), then at #30 prepare the API freeze and at #31 formally freeze the Twin v1 API, followed by convergence (#32), resuming commits (#33), persisting the QPU digital twin (#34), and in sequence comparing calibration drift (#35), building calibration history (#36/#37), building verification history (#38/#39), extending and aligning history (#40/#41), comparing candidates (#42), and resuming candidate commits (#43). Completing the combination of “freeze the interface + build historical sequences” within one day is the quantum direction’s first time making twin data management a reproducible process within the 2.2 cycle.

1.11 Framework and tooling: FlagScale unified CI, Torch-FL backend routing alignment, TE-FL Hygon interfaces (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagScale #1285 (cross-accelerator CI unification)

FlagScale’s [CICD] Unify dependency-coupled CI across accelerator platforms (#1285, 09-15 09:28) consolidates mutually coupled CI dependencies across accelerator platforms—the training framework side’s response to the cost problem of “one change must run across many platforms.” Torch-FL has two: ci: align backend routing with platform FlagGems availability (#272) aligns backend routing with the actual availability of FlagGems on the platform; fix: restore MUSA FlagGems routing and fall back the ops it cannot run (#275) restores MetaX MUSA’s FlagGems routing and adds a fallback path for operators it cannot run—this patch precisely illustrates that “routing exists but operator coverage is incomplete” is a real constraint of the current multi-chip stack. TransformerEngine-FL has one: support for Hygon’s two multi-tensor scale interfaces (#122). FlagPrism has one: adding profiler and debugger support for Moore Threads and updating the docs (#11).

1.12 FlagGems-sglang: competition outputs merged in batch, 10 PRs into the main repo the same day (09-14)

Date: 2026-09-14 Source: FlagGems-sglang #59 (fused rmsnorm warp2)

Within an hour and a half, from 11:58 to 13:31, 10 PRs were merged into the main repo in sequence, with submitters and branch names clearly pointing to cross-chip operator optimization competition tasks: competition/bmm-chunk, competition/decode-attention, competition/decode-grouped-attention, competition/embedding-lora-a, competition/mamba-layernorm-gated, competition/sgemm-lora-b, competition/fused-rmsnorm-warp2, as well as flagos-task19, flagos-task21-moe-sum-reduce and other task branches. Competition outputs entering the main repo in this batch form means the repo has taken on the role of a “landing channel for competition results,” whereas in the previous window the same repo had only sporadic merges.

1.13 build-infra: Hygon vLLM rc2.post1 dual image tags recorded and re-verified (09-14)

Date: 2026-09-14 Source: build-infra #883 (Hygon re-verification record)

All 5 build-side commits revolve around the Hygon vLLM line: first adding a pending changelog for the hygon plugin’s rc2 rebuild (#879), then consecutively recording two application image tags—2.1.2-0.3.0rc2.post1_gc9bbcf0.d20260914 and 2.1.2-0.2.2rc2.post1_g52949b6.d20260914 (both corresponding to hygon-dtk26.04, and to vllm-plugin-FL’s two version lines respectively), then redefining the upstream_prs field as a list for tracking PRs pending merge (#882), and finally recording the Hygon re-verification completed on the aforementioned rc2.post1 tags (#883). Image tags appearing in pairs with version tags is the direct evidence of whether each backend can enter the 2.2 delivery scope.

1.14 docs and release-info: documentation line merged into CICD overhaul, release artifact site synced (09-15)

Date: 2026-09-15 Source: docs #502 merge

The docs repo had 4 commits between 09-15 09:12 and 09:28, all merges of the new/flagcicd branch with main (#501, #502, and two merges), with the documentation line merged into this round of CICD overhaul. The release-info repo had pushes during the window (09-14 12:42), but neither the default branch nor the branches API showed any new in-window commits, so per existing criteria it is classified as release artifact site/tag synchronization with no verifiable content; its repo description is “a placeholder site used only for publishing information about FlagOS stack release artifacts.”


II. News Coverage and Ecosystem

2.1 Open3D-PIMC Open-Sourced at China Computing Conference: Jointly Developed by Tsingmicro and BAAI, Completing the Compilation Stack for 3D Computing Chips (09-14)

Date: 2026-09-14 Source: Xinhua News, “Hardware as the Bone, Open Source as the Pulse: Open3D-PIMC Completes a Key Piece of China’s Domestic Computing Software Ecosystem”

At the 2026 China Computing Conference held in Langfang, Hebei, Open3D-PIMC — the world’s first programming model and open-source software framework for 3D computing chips — was officially open-sourced. Jointly developed by Beijing Tsingmicro Intelligent Technology and the Beijing Academy of Artificial Intelligence (BAAI), the code is now live in the open-source community, positioned as one of the primary solutions of the domestic AI software stack FlagOS for next-generation 3D computing chips. Men Chunlei, head of AI systems research at BAAI, offered this assessment: key technologies for 3D chips are achieving breakthroughs and mass-produced chips will soon appear in succession; without a companion compiler, this would create a predicament of “advanced hardware, inefficient software,” and the project’s goal is to converge industry forces to form a global standard for 3D chip compilers. The report also cites FlagOS’s overall figures: it has been adapted to more than 30 AI chips from 20 chip vendors, covering GPU, NPU, GPGPU, DSA, RISC-V AI, ARM, and other architectures; on the day Alibaba released Qwen3.8-2.4T-A95B in August 2026, the FlagOS community completed adaptation and validation on nine chips including Huawei Ascend, MetaX, and Tsingmicro. The project team says it will release its latest optimization results for 3D chip computing at a top global technology conference in Q4 this year.

2.2 Eleventh Consecutive Quiet Window for Component-Level Search: Core Hits Concentrated in BAAI-Side Coverage (09-14~09-15)

Date: 2026-09-14 ~ 09-15 Source: Google News search (FlagOS and component names)

Chinese- and English-language searches using FlagOS and component names such as FlagGems, FlagScale, FlagTree, FlagPerf, FlagAttention, FlagCX, KernelGen, FlagQuantum, FlagPrism, FlagBLAS, and FlagOS-Robo as keywords (including when:7d and when:14d windows) yielded only one strongly relevant Chinese report within 24 hours (the Open3D-PIMC item in 2.1), with no hits on the English side; HN Algolia queries for terms such as FlagOS and FlagGems also produced no substantive technical discussion within the window. This is the eleventh consecutive quiet window for component-level news search, and the technical information surface for this window comes almost entirely from code repositories. It is worth noting that the flagship terms remain drowned out by financial market coverage: of the hundreds of Chinese results retrieved using member company names, the vast majority were stock prices, lock-up expirations, IPO first-day gains, and fund holdings, all of which were removed in bulk based on headline semantics.

2.3 Three Industry Conference Threads: The Domestic Computing Narrative at CIFTIS, the Bund Summit, and the China Computing Conference (09-09~09-14)

Date: 2026-09-09 ~ 09-14 Source: Guandian, “Moore Threads CFO Xue Yansong: Nvidia’s Share of China’s AI Chip Market Falls to Under 8%”

The industry’s messaging on domestic computing during the window came from three conferences: CIFTIS in Beijing (Moore Threads CFO Xue Yansong stating that Nvidia’s share of China’s AI chip market has fallen from 95% to under 8%, that domestic accelerator cards have surpassed 60% share, and disclosing that the company worked with a national-level laboratory to complete pretraining of a 236-billion-parameter scientific large model on a cluster of over 12,000 cards, is now tackling trillion-parameter models, and that a Peking University team’s world model based on the S5000 cluster has topped a Stanford track leaderboard for over 47 days); the Bund Summit in Shanghai (09-09 to 09-12, where Tsingmicro demonstrated a tenfold leap in large model inference performance after tuning with Ant Group, along with a strategic partnership with Tencent Cloud and Token Factory plans); and the China Computing Conference in Langfang (the debut of Open3D-PIMC, plus multiple releases by member units in the heterogeneous inference direction). The common narrative across all three conference threads is “from usable to genuinely good,” consistent with the focus of work during the FlagOS-side testing period described in 2.2.


III. Deep Dive on Member Organizations

3.1 Qingwei Intelligent: The Dual Narrative of Open3D-PIMC and 5000P Reconfigurable Compute (09-14)

Date: 2026-09-14 Source: Sina Finance (Huanqiu.com) “As AI Moves from Conversation to Production, Qingwei Intelligent Decodes New Internet Compute Demands with Reconfigurable Compute”

Qingwei had no dedicated commits on the GitHub side this window; activity was concentrated on the news side, advancing on two fronts. The first is Open3D-PIMC from 2.1: the project is jointly developed by Qingwei and BAAI, and Li Bin, Qingwei’s Vice President of Software, summarized the company’s eight-year path in a keynote at the China Computing Conference as “compensating for process with architecture, surpassing process with integration, aggregating compute with systems, and creating an ecosystem with open source,” while stating that the company has deployed and under-construction reconfigurable compute exceeding 5000P nationwide. The second is its exhibition at the Bund Conference: since the second half of 2025 it has been deeply tuning with Ant Group, achieving a tenfold leap in large model inference performance, and in 2026 it reached strategic partnerships with internet customers including Tencent Cloud and is advancing its Token Factory layout. Reading the two together, Qingwei’s current role in the FlagOS ecosystem is not filling in operators but positioning itself toward “software definer of the next-generation chip form” — the 3D chip compilation stack and the reconfigurable compute system are upstream and downstream of the same narrative.

3.2 Hygon: From 17 Operators to rc2.post1 Image Re-verification, the Largest Engineering Effort This Window (09-14)

Date: 2026-09-14 Source: FlagGems-Experimental #568 (Hygon conv_depthwise2d)

Hygon had the largest engineering effort among member organizations this window, advancing on four fronts simultaneously: operators (17 Hygon operators in FlagGems-Experimental merged in batch between 17:51 and 18:54); inference (moe_sum added to FlagGems-vllm); frameworks (TransformerEngine-FL supports two multi-tensor scale interfaces); and build and delivery (build-infra recorded two hygon-dtk26.04 rc2.post1 application image tags and completed re-verification, see 1.13). On the news side, Hygon’s “dual-chip acceleration solution for token operations” — first unveiled at the China Unicom Partner Conference on 09-09 (using CPU as the intelligent scheduling hub and DCU as the accelerated compute engine, closing the business loop from token production to repurchase and scaling, and opening up across four dimensions: compute, interconnect, security, and software stack) — continued to be relayed by industry media during the window, a productized expression of converting compute sales into token operations. The alignment of batch operator completion + image re-verification + version tags indicates that the Hygon line has entered a deliverable state.

3.3 Moore Threads: Attention Backend, Profiler Integration, and CIFTIS Share Figures (09-14)

Date: 2026-09-14 Source: FlagPrism #11 (mthreads profiler/debugger)

Moore Threads’ technical actions during the window spanned four modules: FlagAttention added the mthread attention backend (#61); FlagGems and FlagGems-Experimental each added Moore Threads-specific operators (linear, cholesky_inverse, linalg_ldl_solve, ormqr); FlagGems-vllm added the persistent_topk backend for MTT S5000 (#758); and FlagPrism added profiler and debugger support (#11), which was also integrated into FlagTree (#1106). Taken together, what Moore Threads filled in this round is “observability” and “platform specialization of inference hot operators,” rather than simply expanding operator count. On the industry side are the CIFTIS share figures from 2.3 along with data on the 12,000-card cluster and the S5000 world model, making it the most aggressive in its external messaging among member organizations.

3.4 MetaX: metax CI and vLLM 0.24.0, Xiyun C700 on the Eve of Tape-out (09-14)

Date: 2026-09-14 Source: vllm-plugin-FL #488 (metax CI and vLLM 0.24.0)

Two items on the software side: FlagAttention added the metax attention backend (#60); vllm-plugin-FL enabled metax CI and bumped vLLM to 0.24.0 (#488), in sync with the version line recorded in build-infra; additionally, Torch-FL restored MUSA’s FlagGems routing and added fallbacks for operators that cannot execute (#275) — the value of this patch lies in acknowledging the reality that “operator gaps remain even after routing alignment.” On the hardware side, multiple media outlets on 09-14 reported MetaX’s statements at its semi-annual earnings briefing (09-07): the core design and functional verification of the next-generation flagship GPU Xiyun C700 are nearing completion, with the next steps being tape-out, software adaptation, customer testing, and volume production ramp-up; it adds FP4 and other low-precision support, targets overall performance comparable to NVIDIA’s H100, and has domestic solutions across the supply chain in wafer foundry, packaging, and high-capacity memory; the previous-generation C600 already entered volume production in May 2026, equipped with 144GB HBM3e and approximately 1000TFLOPS of FP8 compute. Institutional forecasts put C700 volume production in the second half of 2027.

3.5 Enflame: enflame Backend Continues to Land, Post-IPO Market Cap and Customer Structure Narrative (09-14)

Date: 2026-09-14 Source: FlagGems #6200 (enflame sync)

Two items on the technical side: FlagAttention added the enflame attention backend (#58); the FlagGems mainline merged a batch sync from the Enflame branch (#6200, enflame to flagos 20260911), folding vendor branch changes back upstream. On the industry side, there was continuous post-IPO coverage: on 09-14, reports said its market cap exceeded 180 billion yuan, its stock surged 179% on its first trading day, it raised $912 million, a single winning lot yielded about 130,000 yuan in profit, and public fund IPO subscriptions had unrealized gains of about 3.559 billion yuan; at the same time, skeptical reports noted that while revenue grew 2.8x, losses widened, and dependence on a single customer, Tencent, reached 83%. What is worth recording for FlagOS is that Enflame in this window was both a party “continuously landing new backends” and the member organization with the densest simultaneous capital narrative and business skepticism.

3.6 Iluvatar CoreX: iluvatar Backend Naming Unified, Capital Markets Enter Pre-Lockup Window (09-14)

Date: 2026-09-14 Source: FlagCX #583 (iluvatar build key unification)

The action within FlagOS was renaming the FlagCX build key from iluvatar_corex to iluvatar (#583), a backend naming convergence that reduces the occurrence of two identifiers for the same vendor in the toolchain. There were no major technical releases on the industry side this window; it mainly appeared in market reports on Hong Kong and A-share semiconductor sectors (e.g., semiconductor stocks rising in the 09-15 morning session, with related names up over 3%), which by convention are not included as technical developments and are recorded here only for reference.

3.7 Kunlunxin and DAMO XuanTie (External Ecosystem Parties): XPU Tags and PPU Dual-Track Integration (09-14/09-15)

Date: 2026-09-14 ~ 09-15 Source: FlagCX #592 (T-Head PPU CI)

Two vendors not on the member organization list deepened their participation simultaneously this window, both moving further from “delivery workflows” toward “communication and operator layers.” Kunlunxin line: FlagTree fixed XPU device query and cuda-graph capture errors (#1168) and added a new 0.7.0rc2+xpu3.6 tag (the first accelerator tag without a Triton suffix); on the operator side, it added mvlgamma backend specialization (FlagGems-Experimental #605). DAMO XuanTie (T-Head) line: FlagCX first added the PPU CI workflow (#592), then added PPU torch plugin support and extended the workflow (#596, 09-15 00:49) — this is the first integration of PPU at the communication library level; FlagGems-vllm added the T-Head variant of moe_sum the same day (#772). Both follow the path of “first connect CI, then fill in operators, finally enter the list,” showing that FlagOS’s new chip onboarding has become process-driven.

3.8 BAAI (Lead Party): RC2 Governance Discipline and Open3D-PIMC’s Rule-Definition-Right Statement (09-14)

Date: 2026-09-14 Source: community 2.2 release schedule

The lead party had only one repository action this window, but it orchestrated the entire release chain: after the RC2 manifest submission at 09:35, 23 components were collectively tagged within ten minutes. Reading the manifest, schedule, and commit notes together reveals several clear constraints of this governance system — the testing period between feature freeze (08-31) and release (09-28) accepts only bug fixes and acceptance tests; GA graduation is defined by “the FEP’s Test Plan passing acceptance on the multi-chip matrix”; and version iteration increments via .postN without changing module version numbers. In external messaging, Xinhua’s report framed the significance of Open3D-PIMC as “actively defining the rules”: whoever first establishes a programming model and open source ecosystem for next-generation hardware holds the initiative in defining the technology roadmap. This is consistent with FlagOS’s long-standing positioning of “develop once, run on multiple chips.”


IV. Summary and Trend Observations

This window is the most information-dense day in the 2.2 release cycle, and can be summarized across three dimensions:

  • Release side: 2.2 enters RC2, and the full-stack snapshot transition is complete. Manifest first, script execution, and 23 components collectively tagged within ten minutes, locking in the fixes accumulated during the testing period into a set of locked versions that can be pulled in bulk; with 14 days remaining until the 09-28 GA, attention should focus on the cadence of rc2.postN increments and the Go/No-Go conclusion.
  • Engineering side: acceptance gaps are being filled one by one, with clear vendor division of labor. Hygon completed 17 operators and passed image re-verification, Moore Threads added profiler and inference hot operators, MetaX connected vLLM 0.24.0 and metax CI, Enflame merged vendor branches back into the mainline, and Kunlunxin and DAMO XuanTie are onboarding along the “CI → operators → manifest” path. The direct evidence for judging the 2.2 delivery scope remains the paired image tags and version labels in build-infra.
  • Ecosystem side: moving from “adapting chips” to “defining hardware form factors.” Open3D-PIMC, jointly open-sourced by Tsingmicro and BAAI, pushes FlagOS’s boundary from a multi-chip software stack to a programming model and compilation stack for 3D compute chips, and explicitly plans to release optimization results at a global top-tier tech conference in Q4. The payoff of such moves lies not in the current quarter’s operator count, but in design reference rights for next-generation chips.
  • Items to watch: First, FlagTree has for the first time introduced an accelerator-specific tag in the +xpu3.6 form — whether this will be generalized into a standard delivery form for all backends. Second, Torch-FL added a fallback path for operators that cannot execute, indicating that “routing aligned but operator coverage incomplete” remains a common constraint of multi-chip stacks; whether the 2.2 acceptance matrix can cover such gaps is worth tracking. Third, PPU (DAMO XuanTie) and XPU (Kunlunxin), two external ecosystem parties, are onboarding noticeably faster than conventional vendors — whether this will translate into formal membership.

Limitations note: All technical facts in this report come from public code repositories and official release materials, with release cadence, manifest structure, and version numbers based on community repository files; industry-side information comes from media reports, including vendors’ unilateral statements at conference settings (such as chip market share, compute deployment scale, and model leaderboard results), which have not been independently verified and are cited with sources noted. Component-level news searches have yielded no valid hits for eleven consecutive windows, so technical conclusions rely on code evidence rather than media reports; among the hundreds of search results related to member organizations, the vast majority are capital-market content and have been semantically filtered out, which may have missed a small number of technical news items drowned out by market coverage.


Complete Source List