Research window: 2026-09-11 10:18 ~ 2026-09-14 10:18 Beijing Time (Monday window, past 3 days, covering the weekend) Sources: GitHub (org: flagos-ai, 53 repos pushed_at + single commit search of 150 records fully verified by committer-date + per-repo commits re-check + commit details and patches + repo tree API), community repo 2.2 release list and schedule (raw direct fetch), Google News RSS (18 sets of Chinese and English query terms, via proxy), HN Algolia, FlagOS official community site (versions and livestream list), Tavily/web search, Cailian Press/Securities Times/ITHome/Sina Finance and other industry reports (see appendix)


Index

  • I. Open-Source Project Progress (GitHub Activity)
    • 1.1 2.2 RC1 manifest enters rc1.postN iteration: FlagGems advances to v5.4.0-rc1.post2 (09-11)
    • 1.2 FlagCX productization: libflagcx moves from “build from source” to “per-backend .deb packaging + per-distro apt repository releases” (09-11/09-12)
    • 1.3 FlagCX kernel surface: net adaptor refactored across 6,601 lines, PAL default device API backend completed (09-11/09-13)
    • 1.4 build-infra sglang line: Kunlunxin xre5.37.1 and Iluvatar CoreX corex4.4.0 application lines land (09-12)
    • 1.5 build-infra vllm line: Tsingmicro opens 0.20.2, Iluvatar CoreX 0.24.0 converges to a unified plugin wheel (09-13)
    • 1.6 FlagTree: benchmarks and backend sync after 0.7.0rc1 — Kunlunxin cluster +17,604 lines, Iluvatar CoreX benchmarks enter CI (09-11/09-13)
    • 1.7 FlagGems: KMCompiler Ascend algebra completed, KernelGen multi-vendor onboarding, Kunlunxin copy family migrated to TLE (09-11/09-14)
    • 1.8 Inference and training plugin lines: Hygon vLLM 0.24.0 workflow opens, four SGLang competition branches merged (09-11/09-12)
    • 1.9 Scientific computing and domain operator libraries: Hygon and MetaX operator batches filled in, FlagBLAS Iluvatar CoreX L2, FlagSparse dual-backend (09-11)
    • 1.10 FlagQuantum: vNext architecture merge + 0.2.0 release artifact verification + QPU digital twin (09-11/09-13)
    • 1.11 Training frameworks and tooling: Megatron-LM-FL strips CUDA hard dependency, FlagScale/FlagPrism/TransformerEngine-FL each advance (09-11/09-14)
  • II. News Coverage and Ecosystem
    • 2.1 Tenth consecutive quiet window for component-level search: all technical information comes from code repositories (09-11~09-14)
    • 2.2 Three industry developments: Mobile Cloud heterogeneous inference system released, Enflame’s first day of trading closes, MetaX lockup expiry window approaches (09-11~09-13)
    • 2.3 Community events and competitions: SGLang cross-chip operator optimization contest and operator bounty contest sharing session, competition outputs begin entering the main repo (09-11~09-12)
  • III. Member Unit Deep Dive
    • 3.1 Enflame: STAR Market bell-ringing completes “domestic GPU four dragons” capitalization, technical side simultaneously fixes enflame device adaptation (09-11)
    • 3.2 Iluvatar CoreX: enters Mobile Cloud heterogeneous inference system, while running five parallel workstreams within FlagOS (09-11~09-13)
    • 3.3 Tsingmicro: vLLM 0.20.2 application line opens, end-to-end validation and image tag closed loop (09-13)
    • 3.4 Hygon: communication library deb backend connected, vLLM workflow opened, operator batch fill-in — three lines in motion (09-11/09-12)
    • 3.5 Moore Threads and MetaX: generative operators and backend operators land in two batches, capital节奏 diverges on the industry side (09-11~09-13)
    • 3.6 Kunlunxin (external ecosystem party): major compiler-side sync, but communication library recorded as “non-deliverable” (09-11/09-12)
    • 3.7 BAAI (lead party): 2.2 test-period discipline and RC1 manifest maintenance, GA set for 09-28 (09-11)
  • IV. Summary
  • Appendix: Complete Source List

I. Open-Source Project Progress (GitHub Activity)

Window Overview: Of the 53 repositories in the org, 21 had pushes during the window; the commit search returned 150 in-window commits (full 5 pages, sorted by committer-date descending), and per-repo commit review added another 23, for a total of 173, distributed across 18 repositories: FlagGems-Experimental 54, build-infra 33, FlagGems 21, FlagQuantum 15, FlagCX 7, FlagGems-vllm 6, Megatron-LM-FL 5, FlagTree 4, FlagBLAS 4, FlagGems-sglang 4, FlagDNN 4, FlagSparse 10 (8 of which came from per-repo review), FlagScale 1, FlagPrism 1, vllm-plugin-FL 1, TransformerEngine-FL 1, FlagOS-Compressor 1, community 1. In addition, three repos — sglang-plugin-FL, docs, and release-info — had pushed_at falling within the window, but neither the default branch nor the branch API review showed any new in-window commits (these were branch or tag pushes). No new repositories and no new GitHub Release entries.

The shape of this window = the dual track of “delivery surface taking form + external backends stepping up” during the RC test period: On the governance side, the 2.2 RC1 manifest advanced to a second round of verification (flaggems bumped to rc1.post2); on the delivery side, the heaviest line was FlagCX’s .deb packaging and apt repository publishing being fully walked through (from “users compiling against vendor SDKs themselves” to apt-get install); on the compiler and operator side, Kunlunxin, Iluvatar, Tsingmicro, and Hygon each pushed their application lines, benchmark lines, and operator surfaces forward simultaneously. The previous window’s prediction of an “rc1.postN increment cadence before 2.2 GA” materialized in this window, and at the same time the first backend explicitly recorded as “undeliverable” appeared (the Kunlunxin communication library) — engineering boundaries are starting to be written clearly, not just pushed forward.

1.1 The 2.2 RC1 manifest enters rc1.postN iteration: FlagGems bumped to v5.4.0-rc1.post2 (09-11)

Source: community #108 (09-11 11:45) (the only in-window change to the manifest file release/2.2/release-2.2-rc1.yaml)

  • Second round of verification landed: The flaggems entry in the manifest was updated from v5.4.0-rc1.post1 to v5.4.0-rc1.post2, with the tag pointing to the rc1 branch head 270bebe1. A cumulative 4 fixes landed after post1 — reverting FlagTune’s Mul cost model, Iluvatar randperm/sort fixes, an index int64 offset overflow fix, and a fix for the missing DSA __init__.py packaging.
  • Current RC1 manifest (25-component snapshot): The L0 layer comprises the three FlagTree lines 0.7.0rc1.post1+triton3.3/3.5/3.6 and FlagCX v0.14.0-rc1.post1; operators and domain libraries are FlagGems v5.4.0-rc1.post2, FlagGems-vllm v0.2.0-rc1.post1, FlagGems-sglang v0.1.0-rc1.post1, FlagAttention v0.4.0-rc1.post1, FlagFFT v0.2.0-rc1.post1, FlagSparse v0.3.0-rc1.post1, and FlagBLAS/FlagDNN/FlagTensor/FlagAudio all at v0.3.0-rc1.post1; framework integrations are vLLM-plugin-FL v0.3.0-rc1.post1 (plus a 0.2 line at v0.2.2-rc1.post1), SGLang-plugin-FL v0.2.0-rc1.post1, Torch-FL v0.2.0-rc1.post1, and TransformerEngine-FL / Megatron-LM-FL v0.3.0-rc1.post1; training and tooling are FlagScale v2.1.0-rc1.post1, KernelGen v2.2.0-rc1.post1, KernelGenBench v0.2.0-rc1.post1, FlagRelease v0.3.0-rc1.post1, and FlagOS-Compressor v0.1.0-rc1.post1.
  • Timeline: The 2.2 cycle runs feature freeze 08-31 → test and stabilization period 09-01 to 09-24 (bug fixes only, no new features) → GA 09-28; the graduation criteria require each FEP’s Test Plan (commands + environment + expected results, covering multiple chips) to pass acceptance during the test period, with any unpassed portions filed as separate acceptance issues.

Interpretation: The 25-entry snapshot and the “rc1.postN increment” mechanism together say one thing — the 2.2 release surface is no longer defined by a feature list, but by “which round of verification passed.” The only manifest change in the window bumped FlagGems to post2, and all 4 fixes were numerical/packaging defects (cost model revert, randperm sort, int64 offset, missing subpackage) — a direct projection of test-period discipline: after the freeze, version numbers are driven only by fixes, not by features.

1.2 FlagCX delivery-ization: libflagcx moves from “build from source yourself” to “build .deb per backend + publish apt repos per distro” (09-11/09-12)

Sources: build-infra #847 (09-11 18:02), #849 (09-11 18:50), #851 (09-11 19:18), #853 (09-11 21:37), #854 (09-11 22:00), #855 (09-11 22:06), #856 (09-12 11:20), #865 (09-12 17:20), #866 (09-12 18:25), #868 (09-12 21:08)

  • Why packaging (#847, +2449 lines): The commit message states that FlagCX is distributed as source, so downstream users must first compile against their own vendor SDK before they can link anything; building the .deb inside the vendor’s own base image fixes the soname, -dev headers, and glibc floor at build time, leaving users with just “install the package.” Implementation-wise, backends.yaml carries the packaging facts (make flags, apt dependencies, asserts proving the vendor SDK exists), deb-config.py and generate_matrix.py --runtime are merged, and --check serves as a drift alarm between the two; the build clones the FlagCX tree at a pinned ref inside base/<backend>, normalizes SONAME and strips rpath, and splits runtime libraries from headers into separate packages; the verify stage relies solely on file installation — proving resolvability and loadability inside the base image, and proving on a clean Ubuntu that Depends won’t drag the vendor SDK onto the user’s machine.
  • Publishing to the apt repository (#851): Publishing is dispatch-optional (publish defaults to false) and lands in a repository named after the Ubuntu version used for the build (flagos-apt-ubuntu24.04 / flagos-apt-ubuntu22.04), with arch not part of the address and left to apt to route; the publish step is placed inside the verify job rather than a downstream publisher, on the rationale that “the file on the path to the user’s apt-get install must be the very file that was just verified”; publish without verify is rejected at the set-matrix stage, and non-tag refs are also rejected (version numbers come from the cloned tags).
  • Six backends enabled at once (#865): iluvatar-corex4.4.0/4.5.0, mthreads-musa4.3.6/5.2.0, sunrise-tangrt1.2.0, and tsingmicro-tsm260610 each got asserts and vendor_lib_dirs determined by in-container probing. This commit also fixed a subtle defect: the Makefile only globs flagcx/adaptor/*.cc, so flagcx_device.cc’s hard reference to devApiBackend was left with an undefined symbol when the vendor .mk didn’t provide PLATFORM_EXTRA_SRCS — and -shared tolerates undefined symbols, with only dlopen(RTLD_NOW) surfacing them; cambricon already had coverage, while iluvatar/sunrise/tsm needed it added.
  • Hygon DTK backend brought online (#866): The probing conclusions landed — DTK does not export CUDA_PATH at all, so du.mk’s DEVICE_HOME ?=/CCL_HOME ?= capture empty values, and both paths were pinned to /opt/dtk/cuda/cuda-12; on DTK, -lnccl actually resolves to RCCL, so the resulting NEEDED entry is librccl.so.1, and the nccl spelling doesn’t match in debian/rules’ shlibs loop; DTK’s libraries aren’t registered with ldconfig and LD_LIBRARY_PATH comes from a profile script that only takes effect in bash, so dpkg-shlibdeps must list both directories explicitly. The same commit fixed the verify ldd gate being affected by the image’s own environment (the base image points BASH_ENV at a hook that re-exports LD_LIBRARY_PATH, and since ldd is a bash script it executed that hook first).
  • First “undeliverable” record (#868): The deb entry for Kunlunxin xre5.37.1 remains disabled, with the reason written into the backends.yaml probing notes — the CCL library is not in the image used for the .deb build (the base Containerfile doesn’t install CCL packages, the XRE 5.37.1.0 install payload doesn’t include xccl/bkcl, and the only libbkcl.so in the stack is a runtime artifact inside the vendor torch wheel), and even if found it couldn’t be bound (the library exports C++ mangled symbols and has no pure-C bkcl_* entry points). DESIGN.md’s triage went from “1 probe pending” to “19 ready + 1 with a recorded reason.”
  • Supporting engineering actions: #853 has the pipeline read back the published apt repository “the way a user would”; #854 still verifies rows that built successfully when one row fails; #855 lets .deb dispatch accept a backend list; #856 reads the .deb with the toolchain that built it; #859 forwards the node proxy into the verify container.

Interpretation: This line is the most product-significant change of the period. Previously, FlagCX’s usability equaled “whether users could compile the communication library themselves” — for a software stack that must cover more than a dozen vendors’ proprietary toolchains, this was the last manual barrier to scaled delivery. Putting .debs into distro-partitioned apt repositories means FlagCX’s installation surface is, for the first time, at the same level as the pip install of the vLLM/SGLang plugins; and the three constraints — “publishing must be the same file that was just verified,” “publish must not bypass verify,” and “non-tag refs may not publish” — show this layer was built with an audit mindset. Kunlunxin being explicitly recorded as undeliverable rather than vaguely shelved also makes “19 ready / 1 with a reason” a delivery surface that can be stated externally.

1.3 FlagCX kernel surface: net adaptor refactored at 6601 lines, PAL default device API backend filled in (09-11/09-13)

Sources: FlagCX #578 (09-13 23:24), #582 (09-13 01:18), #584 (09-13 01:20), #585 (09-13 01:20), #586 (09-13 01:21), #588 (09-12 23:56), #580 (09-11 17:02)

  • Network adaptor refactor (#578, +5309/-1292): The largest single change in the main FlagCX repo during this window, larger than the other six combined. The PAL (platform abstraction layer) network adaptor was rewritten, continuing the main line of “the communication library absorbing platform differences into an abstraction layer.”
  • Device API backend wiring (#582): Links the default device API backend for the remaining platforms — the other side of the same problem described by build-infra #865’s “vendor .mk left blank causing devApiBackend to be undefined”: the main repo adds the link, the delivery side adds the field.
  • Usability and documentation surface: #586 falls back to the system FlagCX library instead of failing outright when FLAGCX_PATH is unset; #584 lists all missing backends in getting_started (making “which platforms are supported” a documentation fact); #585 pins the clang-format version used by the format hook.
  • Supply-chain hardening (#588): Pins the Hygon CI image to the digest prior to SHCA, preventing upstream image changes from silently altering the build environment.
  • Vendor adaptation fix (#580): Corrects the topsDeviceProp type name in the Enflame device adaptor.

Interpretation: FlagCX did three qualitatively different things in this window — kernel-surface refactoring (#578), platform wiring fixes (#582/#580), and delivery and CI hardening (#584/#585/#586/#588). Combined with the packaging line in 1.2, the communication library is turning from “a library that can connect” into “a library that can be installed, located, and reproducibly built.” For a multi-chip software stack, the communication library is precisely the component most prone to accumulating special cases across vendor SDKs (Hygon’s RCCL/ldconfig, Kunlunxin’s missing pure-C entry points, Enflame’s type names); this window’s actions show these special cases being converged one by one into a testable abstraction layer rather than left on vendor branches.

1.4 build-infra sglang line: Kunlunxin xre5.37.1 and Iluvatar corex4.4.0 application lines land (09-12)

Sources: build-infra #857 (09-12 11:35), #860 (09-12 12:27), #861 (09-12 12:28), #862 (09-12 13:00), #863 (09-12 14:42), #864 (09-12 13:14), #867 (09-12 18:57), #869 (09-12 19:05), #870 (09-12 21:09), #871 (09-12 19:28)

  • Kunlunxin sglang line closed loop (#857→#864): First made kunlunxin-xre5.37.1 “buildable and drivable,” then recorded F/T (functional and performance) verification results, landed the sglang0.5.18-kunlunxin-xre5.37.1 changelog, recorded the application image tag 2.1.2-0.1.dev1_g7fb22a0c2, and wrote the Kunlunxin application image into the status matrix. The same batch fixed vllm-plugin-wheel aborting on SIGPIPE while resolving plugin_ref (#860).
  • Iluvatar corex4.4.0 application line (#867→#871): Completed and validated the corex4.4.0 application config, landing the sglang0.5.18-iluvatar-corex4.4.0 changelog and image tag 2.1.2-0.1.dev1_g4d44a24cd.
  • 910C line wrap-up: #848 restored the pending-release changelog entry for the Ascend 910C rebuild, #850 (09-11 19:19) recorded the CANN 8.5.0-910c image tag 2.1.2-0.2.0_gf31b199.d20260911, #835 recorded the 910C end-to-end results and the delivered cann8.5.0 T-path fix, and #852 (09-11 20:55) documented the docker layer cache trap behind “stale application images.”

Interpretation: The sglang-side action pattern has standardized into a five-step loop — make a backend buildable, run F/T verification, land the changelog, record the image tag, enter the status matrix. Kunlunxin and Iluvatar each completed one pass in the same window, showing this loop (validated separately by Ascend and Tsingmicro in the previous window) is now a reproducible routine rather than a one-off effort per backend. For 2.2 GA, what actually determines release scope is the number of “verified” cells in the status matrix; this window added two — Kunlunxin and Iluvatar corex4.4.0.

1.5 build-infra vllm line: Tsingmicro opens 0.20.2, Iluvatar 0.24.0 converges to a unified plugin wheel (09-13)

Sources: build-infra #872 (09-12 21:09), #873 (09-13 10:27), #874 (09-13 20:10), #875 (09-13 22:38), #876, #877 (09-13 22:49), #878 (09-13 22:56)

  • Tsingmicro tsm260610 opens vLLM 0.20.2 (#873): Application matrix membership is determined by each backend’s deps_app key in configs.yaml; Tsingmicro previously had only vllm0.24.0, causing generate_matrix.py --app vllm0.20.2 to output an empty include list and the build job to be skipped before it ran. This change added vllm0.20.2: [] (the 0.20.2 plugin wheel needs no extra vendor packages), along with a pending-release entry, with the image tag recorded as 2.1.2-0.2.1_g90ffdf0.d20260912 (#874) and the plugin pinned to the VPF #489 head that the F/T end-to-end verification targeted (#872 recorded that round’s 0.20.2 F/T results).
  • Iluvatar 0.24.0 convergence (#875→#878): The two corex application images were rebuilt from a single-commit vllm-plugin-FL head (symm_mem stub), converging corex4.4.0 and 4.5.0 onto the same vllm_fl version instead of each pinning VPF #434 and the g07063fd variant wheel; the image tag is recorded as 2.1.2-0.2.1_gc9e2573.d20260913. The same commit also explicitly rejected the plan to “fix the corex4.4.0 T path in the same wheel” — that fix was never merged, the probing patch could only make serve output garbage, and sharing the Iluvatar backend module meant carrying it would put the already-verified corex4.5.0 T path at risk, so it was ultimately recorded in section 14.4 rather than forced through.

Interpretation: Both actions point to the same engineering orientation — the matrix must be “enumerable,” and backends must “converge onto the same plugin.” Tsingmicro’s problem wasn’t that the chip was unusable, but that the application matrix key wasn’t fully written, so the build job was silently skipped; this kind of absence that “looks like nothing happened” is hardest to spot during the RC period, and best illustrates that configs.yaml has become the real manifest of the release surface. The Iluvatar case is a textbook trade-off record: giving up a fix that would contaminate a shared module in exchange for the certainty of the verified corex4.5.0 path — the RC test period’s “fixes only, no features” discipline manifests in a concrete commit as “rather not fix than break a verified cell.”

1.6 FlagTree: benchmarks and backend sync after 0.7.0rc1 — Kunlunxin cluster +17604 lines, Iluvatar benchmarks into CI (09-11/09-13)

Sources: FlagTree #1150 (09-11 19:55), #1153 (09-13 22:28), #1155 (09-11 17:55), #1157 (09-11 17:24)

  • Kunlunxin XPU cluster sync (#1155, +17604/-1775): Synced cluster analysis and passes from internal 5a664566 — introducing Scalar/Tile/Vectorizability analysis and XPU-specific passes (Normalize, AsyncLoadSchedule, TLELegalize, LoopInvariantStaging, LegalizeExternEW), adding stage_sm / load_scalar_indexed operators and GM2SM lowering, preserving the hand-written OffsetAnalysis attribute, and defaulting budget-tiling and loop-invariant-staging to off; it also introduced the P-TLE raw frontend (tle.raw), fixed lowering of size-one make range, and had compiler.py pass through the UnrollControl budget knob (pin_unroll_num=-1 makes vector-add take the existing unroll path). The source is annotated as baidu/xpu/triton 6848085b..5a664566.
  • Iluvatar benchmarks into CI (#1153): Added iluvatar’s vLLM benchmark to the workflow, consistent with the previous window’s PPU line.
  • PPU benchmark refresh (#1157): Updated the PPU vLLM benchmark results on 890P (33 rows of data).
  • Triton upstream sync (#1150): Brought in warp layout broadcast from upstream triton.

Interpretation: After the 0.7.0 version convergence (previous window), the compiler’s center of gravity shifted to “backend benchmarking + cluster sync.” The two benchmark lines (Iluvatar vLLM benchmark into CI, PPU 890P results refreshed) show performance data beginning to be continuously recorded per backend; the Kunlunxin +17604-line cluster sync shows the XPU line still advancing via “internal branch → open-source cluster,” accompanied by conservative settings on “which passes are off by default” — benchmarks and pass toggles are equally managed objects.

1.7 FlagGems: KMCompiler Ascend algebra filled in, KernelGen multi-vendor landing, Kunlunxin copy family migrated to TLE (09-11/09-14)

Sources: #6093 (09-11 18:59), #6136 (09-11 17:48), #6161 (09-11 17:13), #6201 (09-14 09:44), #6203 (09-11 17:16), #6205 (09-11 17:51), #6209 (09-14 09:48), #5671 (09-14 10:12), #6183 (09-11 17:33), #6192 (09-11 18:37), #6213 (09-13 20:15), #5642 (09-14 09:09)

  • KMCompiler’s Ascend algebra continues to fill in: In-window merges include the Ascend backend’s matrix_rank (#6161), igammac (#6136), gru (#6201), adaptive_max_pool3d (#5671, with a Triton kernel shared by NVIDIA and Ascend), the Ascend baseline path skip for linalg_solve_triangular (#6209), and dependency declarations (#6205).
  • KernelGen-generated operators land across vendors: On the NVIDIA side, split_with_sizes (#5590) and convolution_overrideable (#5642, with a Triton kernel); on the Moore Threads side, four specialized operators landed at once — conv_transpose1d (#6174), upsample_linear1d_backward (#6175), fmod_ (#6170), and matmuladd (#6180).
  • Kunlunxin copy family migrated to TLE (#6093): The copy family operators were migrated from the regular path to a TLE (Triton Language Extension) implementation, corresponding to the TLE push on the FlagTree side.
  • Correctness and engineering surface: Three batches of [k]-prefixed category fixes — indexing and sorting (#5971), nn (#5969), and math (#5967); the autotuning cache enabled SQLite WAL mode and busy_timeout to fix “database is locked” (#6203); a CI fix for sort_exports mistakenly deleting imports when there are dual __all__ (#6213); and updated FlagTree CI docker images for Kunlunxin/Iluvatar (#6192).
  • Benchmarks: Added a standalone FP8 matrix multiplication performance test using vLLM as the baseline (#6183); fused_marlin_moe docs now state that “the weight layout is not the vLLM Marlin layout” (#6204).

Interpretation: The structure of FlagGems’ 21 commits this window is clear — Ascend is filled in operator by operator along the KMCompiler route (one commit per operator), NVIDIA and Moore Threads land in batches along the KernelGen route, and Kunlunxin enters TLE-ization. The significance of three parallel paths is that the same operator library simultaneously carries three production capacities — “compiler auto-generation,” “generator batch output,” and “hand-written TLE optimization” — each corresponding to backends of different maturity. On the

II. News Coverage and Ecosystem

2.1 Tenth Consecutive Quiet Window for Component-Level Search: All Technical Information Comes from Code Repositories (09-11~09-14)

Sources: Google News RSS 18 query sets in Chinese and English (via proxy), HN Algolia (four sets: FlagOS / FlagGems / FlagScale / FlagTree), FlagOS official community site version and livestream listings, Tavily/web search

  • Zero hits across all component-level queries: FlagOS, FlagGems, FlagScale, FlagTree, FlagPerf, FlagAttention, FlagCX, KernelGen, FlagOS-Robo, FlagQuantum (when:7d~14d, bilingual), and BAAI open source all returned no valid hits within the window, constituting the tenth consecutive quiet window for component-level search (the previous window was the ninth).
  • No valid entries on HN: The four query sets returned only general technical discussions within the window (a flagged GrapheneOS submission, DietPi v10.7, several Show HN projects), none related to FlagOS components; all were excluded per existing criteria.
  • Member organization keyword hits dominated by capital market content: 燧原 when:2d (46), 沐曦 when:2d (27), 摩尔线程 when:2d (23), 海光 when:2d (11), 智源研究院 开源 when:7d (7), 天数智芯 when:2d (6), 地平线 开源 when:2d (1); capital market and repost content accounted for the vast majority, BAAI community articles (AI conference, open-source model analysis, etc.) had no direct connection to the FlagOS tech stack, and SEO articles containing “sports/login/entry/official site/download/gambling” were excluded in full. Retained industry items after filtering are covered in 2.2.
  • No new versions on the official community site within the window: The latest post on the FlagOS official community site remains 08-28 (GLM-5.3-Flash Day0 adaptation across 9 chips); new additions within the window were livestream/event items only, contrasting with the dense activity on the code repository side.

Interpretation: The tenth consecutive quiet window is itself a piece of information—FlagOS’s external visibility is still primarily constituted by code and release cadence, and industry coverage has yet to form sustained tracking of components. For users, this means the answer to “are the components moving” must be read from the repositories rather than from the news; the gap between this period’s 173 commits and zero news also indicates that ahead of the 2.2 GA (09-28), the community kept its attention on delivery rather than on promotion.

2.2 Three Industry Developments: Mobile Cloud Heterogeneous Inference System Released, Enflame’s Market Debut Closes, MetaX Lockup Expiry Window Approaches (09-11~09-13)

Sources: ITHome (09-13), Cailian Press/STAR Market Daily (09-13), Securities Times (09-11 12:22), Cailian Press (09-11 21:00), Sina Finance (09-13/09-14)

  • China’s first “domestic GPU + neuromorphic chip” heterogeneous hybrid LLM inference system released (reported 09-13, released 09-11~13): At the 2026 China Computing Conference held in Langfang, Hebei, Mobile Cloud, together with CETC Nanhu Research Institute, Beijing Lynxi Technology, Shanghai Tianshu Zhixin, Tsinghua University, and Peking University, released the system. The technical approach applies PD/AF separation to the Transformer—Prefill and Attention are assigned to domestic GPUs, while latency-sensitive modules such as FFN (MoE experts) are assigned to neuromorphic chips, leveraging their compute-in-memory and on-chip large-capacity SRAM bandwidth advantages; the team developed its own model compiler, high-speed interconnect protocol, and unified inference engine to handle task decomposition, coordinated scheduling, and result aggregation, with the solution benchmarked against NVIDIA’s next-generation Vera Rubin with Groq heterogeneous inference architecture. In testing, 3 Tianshu Zhixin GPU servers paired with 3 neuromorphic racks running DeepSeek V4 Flash achieved more than double the inference output and inference energy efficiency versus a pure-GPU cluster of equivalent investment scale, with business operating costs reduced by over 40%; the project has accumulated 15 granted invention patents and 6 software copyrights, and has now entered small-batch trial production.
  • Enflame Technology lists on the STAR Market, closing up 179% on day one (09-11): Issue price of RMB 142.18/share, 43,035,200 shares issued, total raised approximately RMB 6.12 billion; opened at RMB 410 (+188.37% vs. issue price), intraday high of RMB 475, closed at RMB 397 (+179%), full-day market cap of approximately RMB 170.9 billion, turnover rate of 76%; online effective subscription multiple of 6,109x, allotment rate of 0.0246%. The company is not yet profitable and was included in the STAR Growth Tier per regulations; revenue for 2023 to 2025 was RMB 301 million / 722 million / 990 million, with H1 2026 revenue of RMB 1.12 billion, up 279.08% year-over-year; it is the only one of the four domestic GPU companies betting on a DSA-specific architecture and not compatible with the CUDA ecosystem. With Enflame’s listing, the “four little dragons” of domestic GPUs (Moore Threads, MetaX, Enflame, Biren) have all completed capitalization.
  • MetaX lockup expiry window approaches (previewed 09-13, effective 09-17): MetaX’s 13.966 million restricted shares will be unlocked on September 17, corresponding to a market value of approximately RMB 6.829 billion, roughly 75% of the free float; concurrent reports noted its share price fell nearly 30% in September and has dropped over 50% from its peak. Background data (08-30 semi-annual report, outside this window): H1 2026 revenue of RMB 1.324 billion, up 44.67% year-over-year, with net profit attributable to shareholders of RMB 612 million marking a turnaround (non-GAAP still at -RMB 49 million), including Q2 net profit of RMB 711 million; the company announced on 2026-06-12 a planned H-share issuance, launching an A+H dual-platform layout.

Interpretation: The three developments sit at different levels—Mobile Cloud’s heterogeneous inference system is a signal at the technology roadmap level (GPUs are no longer viewed as the sole form of compute; neuromorphic chips enter the inference chain positioned as “bandwidth-sensitive modules,” and Tianshu Zhixin is the GPU provider in this system); Enflame and MetaX are signals at the capital level (with all four little dragons capitalized, market focus shifts from “can they build it” to scaled delivery and profit realization, and MetaX’s lockup expiry window is a direct stress test of this shift). The relevance to FlagOS lies in this: both types of signals are raising the practical value of a “multi-chip unified software stack”—heterogeneous inference systems require compilers and unified inference engines to schedule two types of compute, and as domestic GPUs enter scaled delivery, the migration cost of their software stacks directly determines whether orders can land.

2.3 Community Events and Competitions: SGLang Cross-Chip Operator Optimization Contest and Operator Bounty Contest Sharing Session, Competition Output Begins Entering the Main Repository (09-11~09-12)

Sources: FlagOS official community site (events and livestream listings), merge records in the FlagGems-sglang repository within the window

  • SGLang Cross-Chip Operator Optimization Contest: Co-hosted by Zhongzhi FlagOS and IEEE, with SGLang as co-organizer; Track 1 covers multi-chip performance optimization of SGLang framework operators, featuring 200+ real inference operator problems, with participants free to choose problems released in batches, development using Triton / Triton-TLE, unified validation and unified speedup measurement across multiple chip platforms, and a real-time leaderboard with three award categories: “Full-Domain Conquest Award / Breakthrough Award / Single-Problem Extreme Performance Award.”
  • Operator Bounty Contest Champion Sharing Session: The official community site recorded a 09-10 19:00 livestream by winning participants (journey, key decisions, pitfalls to avoid, practical tips, and interactive Q&A), part of an ongoing competition series.
  • Competition output enters the main repository: Within 40 minutes on 09-12, FlagGems-sglang merged four PRs from competition/ and participants’ personal branches (add-chunk-local-cumsum-vec, add-context-attention, qkv-lora-b, chunk-state-varlen-v2), all targeting chunked attention and LoRA-related operators in the SGLang inference chain.
  • Ecosystem events: The forum “Open AI Computing: Building a New Open-Source Software Ecosystem for Diverse Heterogeneous Hardware,” hosted by the Zhongzhi FlagOS community, was held on 09-07 in Shanghai alongside the KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China (with participation from BAAI, Shanghai AI Laboratory, the PyTorch Foundation, SGLang, and others).

Interpretation: Competition output entering the FlagGems-sglang main repository as PRs is the most easily underestimated item this period. A common problem with community competitions is “lively leaderboards, no changes in the main repository,” yet four competition/ branches were merged into master on the same morning, indicating that the problem sources, validation environment, and main repository are now connected—the competition has in effect become an external supply line for operator capacity, and these operators will enter the 2.2 release list with the next rc1.postN. This also explains the move in 1.8 where “the plugin layer made tests vendor-agnostic”: to unify speedup measurement, tests and benchmarks must first be made cross-vendor comparable.


III. Deep Dive into Member Units

3.1 Enflame: STAR Market Listing Completes Capitalization of the “Four Little Dragons” of Domestic GPUs, Technical Side Simultaneously Patches enflame Device Adaptation (09-11)

Sources: Securities Times (09-11), Cailian Press (09-11), FlagCX #580 (09-11 17:02)

  • Capital side: Listed on the STAR Market on 09-11 (688801), issue price 142.18 yuan/share, opened at 410 yuan, closed at 397 yuan (+179%), market cap approximately 170.9 billion yuan; all four domestic GPU companies have now completed capitalization, with the “final piece of the puzzle” in place. Prospectus and financial report figures show H1 2026 revenue of 1.120 billion yuan, up 279.08% year-over-year, cumulative R&D investment of 3.676 billion yuan from 2023 to 2025, and 643 R&D personnel as of end-2025 (76.73% of employees), undertaking 12 national and local science and technology research projects and participating in the formulation of 58 key national and industry standards for AI chips and intelligent computing systems.
  • Technical side (within window): FlagCX #580 corrects the type name of topsDeviceProp in the Enflame (enflame) device adapter, a vendor adaptation patch at the platform abstraction layer of the communication library.
  • Position within FlagOS: Enflame still exists in the build-infra application matrix via two sglang lines, enflame-tops1.9.10 and enflame-tops1.10.6, and is one of the enabled probe-passing entries in FlagCX’s .deb backend manifest.

Interpretation: Enflame is the only one of the four little dragons pursuing a DSA-specific architecture with a self-developed software stack, which also determines that its relationship with FlagOS leans toward “self-contained software stack + selective integration”—its two enflame-tops application lines have long existed within the FlagOS matrix, but the only technical action in this window is a single type name correction, which is maintenance-oriented rather than expansionary. By contrast, Iluvatar CoreX, Tsingmicro, and Hygon all made substantive advances on backend and application lines in the same window. After capitalization is complete, Enflame’s next hurdle is monetizing R&D and delivering on orders, which will directly reflect in whether it continues to expand its investment in the multi-chip software stack.

3.2 Iluvatar CoreX: Enters Mobile Cloud Heterogeneous Inference System While Running Five Workstreams in Parallel Within FlagOS (09-11~09-13)

Sources: ITHome (09-13), FlagBLAS #115 (09-12 17:31), build-infra #867/#869/#870/#871 (09-12), build-infra #875/#877/#878 (09-13), FlagTree #1153 (09-13 22:28)

  • Industry side: In the “domestic GPU + neuromorphic chip” heterogeneous hybrid inference system led by Mobile Cloud, Iluvatar CoreX is the GPU provider (3 Iluvatar CoreX GPU servers + 3 neuromorphic cabinets, running DeepSeek V4 Flash), handling Prefill and Attention computation.
  • Five workstreams within FlagOS: First, FlagBLAS L2-level support merged into master (+5325 lines); second, the build-infra sglang application line adds sglang0.5.18-iluvatar-corex4.4.0 (config validation + changelog + image tag); third, on the vllm side, corex4.4.0/4.5.0 for 0.24.0 are consolidated into a unified plugin wheel; fourth, FlagTree adds the iluvatar vLLM benchmark to the CI workflow; fifth, in FlagCX’s .deb backend manifest, iluvatar-corex4.4.0/4.5.0 are both enabled.

Interpretation: Iluvatar CoreX is the most frequently appearing member unit this period, and its technical and industry lines do not overlap—on the industry side it is “entering a heterogeneous inference system as a domestic GPU,” while on the FlagOS side it is advancing simultaneously across four layers: “base libraries + application lines + benchmarks + packaging.” These two things are logically complementary: the partitioning of bandwidth-sensitive modules in heterogeneous inference systems ultimately depends on operator libraries, compilers, and communication libraries for scheduling, and FlagOS is precisely the most direct software foundation candidate for such systems. Notably, FlagTree writing benchmarks into CI means that Iluvatar CoreX’s performance data will henceforth exist as continuous records rather than one-off reports.

3.3 Tsingmicro: vLLM 0.20.2 Application Line Opened, End-to-End Validation and Image Tag Closed Loop (09-13)

Sources: build-infra #872 (09-12 21:09), #873 (09-13 10:27), #874 (09-13 20:10)

  • Application matrix key completion: The Tsingmicro backend tsingmicro-tsm260610 previously only carried vllm0.24.0, causing the output of --app vllm0.20.2 to be empty and the build job to be silently skipped; #873 adds vllm0.20.2: [], bringing that row into the build matrix, along with a changelog and pending release entries.
  • Validation closed loop: #872 records the F/T end-to-end results for 0.20.2, #874 records the image tag 2.1.2-0.2.1_g90ffdf0.d20260912, with the plugin pinned to the VPF #489 head corresponding to the validation; tsingmicro-tsm260610 is also an enabled entry in the FlagCX .deb backend manifest.
  • Continuation: The two workflows established in the previous window—tsingmicro3.6 build-and-test and delivery (on the FlagTree side)—have no new additions in this window, remaining in routine delivery-line operation status.

Interpretation: The value of this Tsingmicro line lies not in new capabilities but in exposing a class of systemic risk—missing keys in the application matrix cause builds to be silently skipped. In a matrix covering over a dozen vendors, each with multiple version combinations, “no errors but nothing built” is harder to detect than a build failure. After this row is added, Tsingmicro has both 0.20.2 and 0.24.0 lines on the vLLM side, plus FlagTree’s dual workflows and FlagCX’s deb backend, giving it integration depth comparable to the earliest batch of member units.

3.4 Hygon: Communication Library deb Backend Connected, vLLM Workflow Enabled, Operators Batch-Completed—Three Lines in Motion (09-11/09-12)

Sources: vllm-plugin-FL #436 (09-11 10:54), build-infra #866 (09-12 18:25), FlagCX #588 (09-12 23:56), FlagGems-vllm #756 (09-12 12:08), FlagGems-Experimental #586 (09-11 18:50)

  • Communication library .deb backend connected (#866): The Hygon DTK backend moves from “probe pending” to enabled status, resolving three classes of environment issues along the way—DTK not exporting CUDA_PATH causing DEVICE_HOME/CCL_HOME to be empty, -lnccl actually resolving to RCCL (with the artifact’s NEEDED being librccl.so.1), and the DTK image’s BASH_ENV hook polluting the verify ldd environment.
  • Plugin and CI: #436 enables the Hygon workflow for vLLM 0.24.0; #588 pins the Hygon CI image to the digest before SHCA to avoid upstream image drift.
  • Operator side: FlagGems-vllm adds persistent_topk and integrates it into the fused __init__; FlagGems-Experimental adds over a dozen Hygon backward and special-function operators in one batch (reflection_pad1d_backward, embedding_bag_dense_backward, mse_loss_backward, binary_cross_entropy_backward, lift_fresh, baddbmm_, etc.).

Interpretation: Hygon has the highest action density among member units in this window, with the three lines being complementary in nature—the communication library side addresses “can it be installed” (the three special cases of DTK paths, RCCL naming, and ldconfig are all real friction points in vendor toolchains, and all were written into commit messages rather than left on someone’s machine), the plugin side addresses “can it run,” and the operator side addresses “is coverage sufficient.” One detail worth noting is pinning the CI image to a digest: in multi-vendor pipelines, the determinism of the build environment is often more easily overlooked than the code itself, and the Hygon line proactively hardens this.

3.5 Moore Threads and MetaX: Two Batches of Generated and Backend Operators Landed, Diverging Capital Rhythms on the Industry Side (09-11~09-13)

Sources: FlagGems #6174, #6175, #6170, #6180 (09-11 18:17~18:21), FlagPrism #10 (09-11 21:24), FlagGems-vllm #771 (09-11 18:07), FlagSparse #58 (09-11 07:44), Sina Finance (09-13/09-14)

  • Moore Threads: The KernelGen route batch-lands four specialized operators (conv_transpose1d, upsample_linear1d_backward, fmod_, matmuladd); FlagPrism adds profiler and debugger support for it and updates documentation; FlagDNN adds a Moore Threads README; in FlagCX’s .deb backend manifest, both mthreads-musa4.3.6 and mthreads-musa5.2.0 are enabled; on the industry side, the JD Cloud 100,000-card cluster news continues (original event 09-09, already covered in the previous issue; this window contains reposts and analysis pieces, not counted again).
  • MetaX: FlagGems-vllm optimizes the GDN chunk kernel (#771); FlagGems-Experimental has approximately 35 MetaX-related commits (new support and fix implementations interleaved); FlagSparse adds metax support; the MetaX line in FlagCX’s .deb backend manifest is not enabled in this window.
  • Industry-side divergence: MetaX will see 13.966 million shares unlock on 09-17 (approximately 6.829 billion yuan), with multiple reports in the window focusing on its stock price pullback and unlock pressure; Moore Threads has a new public record of a “MT Lambda” trademark registration application in this window.

Interpretation: Both companies are in the “batch operator completion” phase on the technical side, but their capacity sources differ—Moore Threads’ four operators come from the KernelGen generation route, while MetaX’s batch is concentrated in hand-written implementations and fixes in FlagGems-Experimental. The industry-side divergence is equally clear: MetaX is entering a share unlock and stock price pressure cycle, while Moore Threads is in the ascending phase of a 100,000-card order narrative. For FlagOS, neither company’s technical investment pace has slowed with capital market fluctuations, and each ranks among the top member units in commit volume in this window.

3.6 Kunlunxin (External Ecosystem Partner): Major Compiler-Side Sync, but Communication Library Recorded as “Non-Deliverable” (09-11/09-12)

Sources: FlagTree #1155 (09-11 17:55), FlagGems #6093 (09-11 18:59), build-infra #857/#861/#862/#864 (09-12), build-infra #868 (09-12 21:08)

  • Compiler side: FlagTree’s XPU cluster sync is the largest single commit this period (+17604/-1775), introducing a complete analysis framework with XPU-specific passes, the P-TLE raw frontend, and UnrollControl budget propagation.
  • Operator side: FlagGems migrates Kunlunxin’s copy family operators to TLE implementations (#6093), echoing the TLE push on the compiler side.
  • Application line: build-infra makes kunlunxin-xre5.37.1 buildable and drivable on the sglang line, completing F/T validation, changelog, image tag, and status matrix records.
  • Boundary record: FlagCX explicitly records kunlunxin-xre5.37.1 as “non-deliverable”—the CCL library is not in the image used for .deb builds (the base image does not install CCL packages, the XRE 5.37.1.0 payload does not contain xccl/bkcl, and the only libbkcl.so resides within the vendor torch wheel, exporting C++ mangled symbols with no pure C entry point).

Interpretation: Kunlunxin presents both “advance” and “limitation” in this window. The +17604 lines on the compiler side indicate that synchronization between its internal branch and the open-source cluster is intensifying; the application side has completed the full sglang closed loop; but the communication library layer is judged currently non-deliverable, with the reasons stated very specifically (the library is not in the build image, and even if found, it cannot be bound). This boundary of “can compile and run, but the communication library can’t be packaged” is precisely the state that most needs explicit management in a multi-chip software stack—it determines whether the backend can support single-machine inference or communication-intensive multi-machine scenarios.

3.7 BAAI (Lead Party): 2.2 Testing Period Discipline and RC1 Manifest Maintenance, GA Set for 09-28 (09-11)

Sources: community #108 (09-11 11:45), community repository 2.2 timeline and RC1 manifest (raw direct fetch)

  • Testing period governance: The 2.2 cycle runs feature freeze 08-31 → testing and stabilization period 09-01 to 09-24 (bug fixes only, no new features) → GA 09-28; FEP graduation criteria require Test Plans covering multi-chip scenarios to pass acceptance during the testing period, with any unpassed portions tracked as separate acceptance issues on the milestone.
  • Manifest maintenance: The only manifest change in the window advances FlagGems from v5.4.0-rc1.post1 to v5.4.0-rc1.post2, with four classes of fixes (cost model rollback, randperm sorting, int64 overflow, subpackage packaging) written into the commit message.
  • Version side: FlagTree’s three triton lines are recorded in the manifest as 0.7.0rc1.post1+triton3.x, FlagCX as v0.14.0-rc1.post1, with a full manifest of 25 component snapshots.

Interpretation: The lead party’s actions in this window are entirely governance-oriented—maintaining the manifest, enforcing freeze discipline, and writing acceptance criteria into verifiable Test Plans. Such work is invisible in external reporting, but it determines whether 09-28 can deliver a surface where “every component has a tag, and every tag has passed multi-chip validation.” Viewed together with engineering-side actions (FlagCX packaging, vendor application lines), the core question for 2.2 is no longer “how many chips are supported” but “can these supports be reproduced, installed, and audited.”


IV. Summary

  1. The delivery surface became the main thread of this window, and the communication library is the first component to be thoroughly engineered (most important change): FlagCX has completed the full chain of “building .deb per backend inside vendor base images → publishing to apt repositories split by Ubuntu version → reading back and verifying in the user’s own way,” and enabled six backends at once; the accompanying constraints (published artifacts must be the same copy as the just-verified file, publish must not bypass verify, non-tag refs must not be published) indicate that this layer is being built with audit awareness. Previously, FlagCX’s usability amounted to “whether users can compile it themselves against the vendor SDK”; henceforth it amounts to a single apt-get install.

  2. 2.2 has entered the latter half of the testing period, with version advancement driven entirely by fixes: The RC1 manifest’s 25 component snapshots are in place, and the only manifest change within the window was pushing FlagGems to v5.4.0-rc1.post2 (four fixes: cost model fallback, randperm sorting, int64 offset overflow, missing subpackage packaging); Iluvatar CoreX 0.24.0 explicitly abandoned a T-path fix that would pollute a shared module in order to preserve the already-verified 4.5.0 path—”rather not fix than break a verified cell” is RC discipline in its most direct code form. GA is set for 09-28, with the testing period running from 09-01 to 09-24.

  3. The “five-step closed loop” of the application matrix has become a replicable procedure, and the risk of missing keys has been exposed: Make the backend buildable → run F/T verification → write the changelog → record the image tag → enter the status matrix; in this window, Kunlunxin (sglang xre5.37.1) and Iluvatar CoreX (sglang corex4.4.0) each completed one pass; Tsingmicro’s issue (deps_app not fully written, causing the build to be silently skipped) indicates that this matrix now has a failure mode of “no errors but nothing was built,” and that configuration completeness is itself part of the delivery surface.

  4. External backends continue to ramp up, and an explicitly recorded delivery boundary appears for the first time: Kunlunxin XPU cluster synchronization (+17,604 lines), Iluvatar CoreX L2 libraries and benchmarks entering CI, Tsingmicro opening vLLM 0.20.2, and Hygon DTK deb backend getting through—four lines advancing simultaneously; at the same time, the Kunlunxin communication library was recorded as “non-deliverable” with the reasons stated (the CCL library is not in the build image, and there is no pure C entry point). The value of a software stack is no longer measured only by “how many chips it supports,” but also by the statable boundary of “to what extent each is supported.”

  5. Operator capacity is advancing on three fronts, and competitions are beginning to become an external supply line: Among FlagGems’ 21 commits, Ascend is filling in one by one via KMCompiler, NVIDIA and Moore Threads are batch-merging via KernelGen, and Kunlunxin is undergoing TLE rework; among FlagGems-Experimental’s 54 commits, Hygon added more than a dozen backward and special-function operators in one go, while MetaX had roughly 35 interleaved “additions + fixes.” More significant in the long term, FlagGems-sglang merged four competition/ branch PRs from the cross-chip operator competition within 40 minutes on the morning of 09-12—competition output landing in the release manifest as main-repo commits for the first time.

  6. Two industry-side threads raise the practical weight of the multi-chip software stack: Mobile Cloud, together with Lynxi Technologies, Iluvatar CoreX, and others, released China’s first “domestic GPU + neuromorphic chip” heterogeneous hybrid inference system (PD/AF separation, 3 Iluvatar CoreX GPU servers + 3 neuromorphic cabinets running DeepSeek V4 Flash, with more than double the performance and energy efficiency and more than 40% lower cost), indicating that heterogeneous compute scheduling has entered the engineering stage; Enflame debuted on the STAR Market with a 179% gain on its first day, all four little dragons are now capitalized, and MetaX faces a RMB 6.8 billion lockup expiration window on 09-17—the focus for domestic GPUs is shifting from “can it be built” to scaled delivery and profit realization, and migration cost and cluster stability are precisely the two variables most dependent on the system software stack in this shift.

Forecast: Four points to watch going forward—first, the incremental cadence of rc1.postN before the 09-28 GA and which other manifest entries will go through another iteration; second, whether FlagTree’s manifest record (0.7.0rc1.post1+triton3.x) will advance to a formal release tag, and whether build-infra’s NVIDIA line pin value will catch up to the mainline from 0.6.1 (the version mismatch noted in the previous window persists); third, whether FlagCX’s apt repository sees its first batch of real user installation feedback, and when backends such as MetaX that have not yet been enabled will be filled in; fourth, whether the Kunlunxin communication library’s “non-deliverable” status will be reversed (for example, whether a subsequent XRE version carries CCL and exposes a pure C entry point), and whether scaled trial production of Mobile Cloud’s heterogeneous inference system will surface specific requirements for the FlagOS operator/compiler layer.


Appendix: Complete Source List

Source Verification results within the window
GitHub org repos API (flagos-ai, 53 repos) 21 repos pushed within the window: FlagGems, build-infra, FlagGems-Experimental, FlagQuantum, FlagCX, FlagGems-vllm, Megatron-LM-FL, FlagTree, FlagBLAS, FlagGems-sglang, FlagDNN, FlagSparse, FlagScale, FlagPrism, vllm-plugin-FL, TransformerEngine-FL, FlagOS-Compressor, community, sglang-plugin-FL, docs, release-info; no new repos; no new entries in GitHub Releases within the window
GitHub commit search (org-wide 150 results, 5 pages, sort=committer-date) Substantive merges into default branches across 11 repos: FlagGems-Experimental 54 / build-infra 33 / FlagGems 21 / FlagQuantum 15 / FlagCX 7 / FlagGems-vllm 6 / FlagTree 4 / FlagBLAS 4 / FlagGems-sglang 4 / FlagPrism 1 / FlagScale 1
per-repo commits recheck (active repos not captured by search) 23 additional commits: FlagSparse 10, Megatron-LM-FL 5, FlagDNN 4, vllm-plugin-FL 1, FlagOS-Compressor 1, TransformerEngine-FL 1, community 1; sglang-plugin-FL, docs, and release-info had no new commits on their default branches within the window (branch/tag pushes only)
GitHub commit details and patches #847 (+2449 lines, .deb packaging and verify criteria); #851 (+85/-25, apt repository publishing and publish/verify constraints); #865 (+147/-75, six-backend enablement and undefined devApiBackend symbol); #866 (+32/-19, DTK paths/RCCL/BASH_ENV); #868 (+30/-11, reasons Kunlunxin is undeliverable); #873 (+46, missing deps_app keys and 0.20.2 opened); #875 (+20/-6, unified plugin wheel and abandonment of T-path fix); #1155 (+17604/-1775, XPU cluster sync and internal source 5a664566); #156 (+103/-54, platform-aware device API); FlagCX #578 (+5309/-1292, net adaptor refactor)
community repo 2.2 release materials (fetched raw) release/2.2/release-2.2-rc1.yaml contains 25 component snapshots; release/2.2/schedule_CN.md specifies feature freeze 08-31, testing and stabilization period 09-01~09-24, GA 09-28, FEP graduation criteria, and the [URGENT] exception channel; the only change to the manifest file within the window is #108 (FlagGems → v5.4.0-rc1.post2, tag pointing to rc1 branch head 270bebe1, including four fixes)
Google News RSS (18 query sets in Chinese and English, via proxy) Component-level queries (FlagOS/FlagGems/FlagScale/FlagTree/FlagPerf/FlagAttention/FlagCX/KernelGen/FlagOS-Robo/FlagQuantum/BAAI open source) all returned zero hits, the tenth consecutive quiet window; among member-organization keyword hits, retained three categories of items: China Mobile Cloud heterogeneous inference system, Enflame IPO, and MetaX lock-up expiry; routine BAAI community articles and gambling-related SEO pieces were removed in bulk
HN Algolia (FlagOS / FlagGems / FlagScale / FlagTree) Results returned within the window were general technical discussions (GrapheneOS, DietPi v10.7, Show HN series), unrelated to FlagOS components, all removed
FlagOS official community site (version and event listings) The latest article/post remains 08-28 (GLM-5.3-Flash Day0 adaptation for 9 chips), with no new versions within the window; new additions within the window are event entries: SGLang cross-chip operator optimization competition (hosted by Zhongzhi FlagOS × IEEE), 09-10 operator bounty challenge champion sharing session, 09-07 KubeCon Shanghai “Open AI Computing” forum
Tavily/web search (original-site date verification) China Mobile Cloud heterogeneous inference system: ITHome 09-13, Cailian Press/STAR Market Daily 09-13 (3 Tianshu Zhixin GPU servers + 3 neuromorphic cabinets, DeepSeek V4 Flash, performance and energy efficiency more than doubled, cost reduced by over 40%, 15 patents + 6 software copyrights, small-batch trial production); Enflame IPO: Securities Times 09-11, Cailian Press 09-11 (issue price 142.18 yuan, opened at 410 yuan, closed at 397 yuan, market cap about 170.9 billion yuan, turnover 76%, allotment rate 0.0246%, 2026H1 revenue 1.120 billion yuan +279.08%); MetaX lock-up expiry: Sina Finance 09-13/09-14 (13.966 million shares, 6.829 billion yuan, about 75% of the free float)
Items verified but not counted again JD Cloud and Moore Threads 100,000-card cluster (original event 09-09, already covered in the previous issue; within the window only reprints and analysis pieces); MetaX 2026 semi-annual report (disclosed 08-30, used only as background data)

Limitations note: The technical information in this issue is still based primarily on code repositories, with industry coverage used only for cross-reference on member organizations’ capital and product developments; Google News RSS was accessed via proxy, and hit volume and coverage are affected by the aggregation sources’ indexing cadence, with component-level queries now having returned zero hits for ten consecutive windows. Article/post updates on the FlagOS official community site lag behind repository activity (latest article/post 08-28), and community event information is taken from its on-site event listings. Corporate financial and market data come from public reports and have not been further cross-validated.