FlagOS Weekly Report (2026-09-21 ~ 2026-09-25)
Reporting period: Monday, September 21, 2026 to Friday, September 25, 2026, 5 working days in total Sources: This publication’s daily FlagOS updates (five issues: 09-21, 09-22, 09-23, 09-24, 09-25; windows connect day by day with no gaps); all items come from verified sources for the relevant period, with additional GitHub cross-checks and supplements during compilation Editorial note: This week is the last full working week before FlagOS 2.2 GA (09-28), with the main thread being release closeout—the three-layer validation of “tags, manifests, scope” was completed in succession; it also covers four dimensions: component and code co-building, onboarding of new chip partners, member organizations, and the ecosystem. Across the five daily report windows, a total of 916 submissions were matched (163 / 224 / 143 / 211 / 175), with 22 / 19 / 24 / 23 / 24 active repositories in each period, covering about 30 repositories within the org in total
I. Weekly Highlights
- 2.2 Release Freeze: Ten components’ official version tags batch-committed (09-24): In the early hours of 09-24 between 01:49 and 01:50, official version tags (bare version numbers, no rc/post suffix) for ten L1 modules were pushed all at once — flaggems v5.4.0, flagfft v0.2.0, flagsparse v0.3.0, flagdnn v0.3.0, flagblas v0.3.0, flagtensor v0.3.0, flagaudio v0.3.0, flagattention v0.4.0, flaggems-vllm v0.2.0, flaggems-sglang v0.1.0; each tag was placed on the HEAD of the corresponding module’s rc2 branch and verified by read-back one by one. A companion PR (community #114) adds the “official tag ← rc2 baseline” mapping to the official manifest; still under review as of this writing.
- rc2 channel closed + scope audit launched (09-24): community #113 uniformly bumped the 13 modules in the rc2 manifest to their latest verified post tags (flaggems rc2.post4, etc.). The rc2 stabilization period ended on 09-24 per schedule, after which only hotfixes are accepted. On the same day, community #115 opened the “2.2 FEP scope audit” — 18 Implemented FEPs retained, with FEP 0019 / 0069 / 0081 / 0084 / 0093 / 0100 deferred to 2.3; retained items are limited to “implemented interfaces + verified configurations.” Release validation has advanced from “code freeze” and “artifact freeze” to “scope and claims freeze.”
- Release Info portal rebuilt, platform matrix rapidly densified (09-22 ~ 09-25): The Release Info portal was rebuilt on 09-23 (20 platform rows for base / runtime + two vLLM matrices + Megatron + sglang application images); the Megatron application image
2.2.0-0.2.3covers 14 platform rows, the megatron_rl matrix is unblocked and opens all 18 backends; the sglang application page shipped seven releases in a single night (Kunlunxin / Ascend / Enflame / Hygon / Moore Threads / Iluvatar / generic — seven platforms), and the DAMO XuanTie thead-ppu2.1.0 application page entered the portal the same day. - New chip partners onboarded in quick succession (09-24 ~ 09-25): Biren’s SUPA backend came online the same day with a two-layer stack of “FlagGems operator library + vllm-plugin-FL inference plugin” (#6367 / #530); Sunrise ported the PTPU 0.2.1 patch set to the vLLM 0.24 mainline (#555); Tsingmicro’s FlagTree platform tags are complete — the three-layer isomorphic integration of “operator library + inference plugin + compiler” has become the standard path for external chips entering FlagOS.
- FlagFFT completes Hygon backend, 102 commits across three workstreams in a single window (09-21 ~ 09-22): The Hygon BW1000 (HCU) adapter entered dev, and the README backend list expanded to seven items: CUDA / MUSA / PPU / IX / MACA / NPU / HCU; in the same window, MACA (MetaX) performance closure ran in parallel with NPU (Ascend) Stockham batching, and the packaging trio (#12 / #18 / #22) was merged simultaneously — new backend + performance closure + distribution packaging completed in the same window.
- DAMO XuanTie advances on two fronts: Zhenwu V900 launch + PPU enters FlagOS matrix (09-22 ~ 09-25): At the Yunqi Conference (09-22), the new-generation training-inference integrated AI chip Zhenwu V900 was launched (216GB memory, 1200GB/s inter-die interconnect, native FP8 / FP4, single-chip performance three times that of the previous generation M890, mass production in Q1 2027); on the software side, PPU entered the build and application matrices in its own right — FlagTree
0.7.0+ppu3.6tag, sglang thead-ppu application page, TE-FL PPU CI; the SAIL software stack announced expanded open-source efforts, with the GitHub development branch to become the main R&D line. - 910C reinforcement learning pipeline ready end-to-end (09-24 ~ 09-25): Two Ascend 910C configurations (CANN 9.0.0 + torch 2.10 / CANN 8.5.0 + torch 2.9) completed single-step-install megatron-core[rl] GRPO end-to-end validation (exit code 0), with three RL fixes including packed_seq gating built into the wheel; the accompanying fix for NPU invisibility caused by Ascend-Docker-Runtime was also resolved — providing a reproducible evidence chain for the 2.2 Ascend claims in RL scenarios.
- Dense industry and ecosystem events (09-21 ~ 09-25): AICC2026 held in Beijing (two FlagOS talks, “Reliable AI” evaluation system launched by 20+ organizations); MetaX MXMACA merged into the LMCache mainline on an “upstream-first” basis and officially announced Day 0 support for 39 mainstream models; Moore Threads’ RLinf official CI integrated MTT S5000, Protenix-v2 full-pipeline inference sped up by ~26%; Loongson released the first version of its accelerated computing platform; Open3D-PIMC coverage matrix expanded to Sina Finance.
II. Versions and Release Activity
RC Cadence and GA Milestones
- Timeline: The rc2 stabilization window closes on 09-24 (only hotfixes accepted thereafter); GA Go/No-Go is set for 2026-09-28; the release tracking issue (community #47) and the milestone both have a due date of 09-28.
- Final full alignment of the rc2 channel (09-24 00:42, community #113): 13 modules bumped to the latest verified tags on the rc2 branch HEAD — flaggems
rc2.post4; flagfft / flagdnn / flagblas / flagtensor / flagaudiorc2.post2(covering Debian / RPM packaging and the Nexus release chain, plus the FlagAudio gain identity fix); flagtree across three lines at0.7.0rc2.post2; vllm-plugin-fl across two lines atpost3; transformerengine-flpost2; flagscalepost2. - Official manifest landed in batches: On 09-22, community #111 added the first entry, sglang-plugin-fl v0.2.0 (rc2 baseline HEAD
1c84db7, with the file header explicitly stating the version convention that “official tags are always bare version numbers, with FlagTree as the exception”); on 09-24, #114 added the L1 ten-component mapping (flaggems v5.4.0 ← post4, flagattention v0.4.0 ← post1, etc.). As of writing, this PR is still under review. - Scope reduction and audit: FEP-0085 moved GEMM+ReduceScatter out of the 2.2 scope (community #112, 09-24, landed via the “test evidence + documentation revision” chain); scope audit PR #115 established the delivery boundary of “18 retained / 6 deferred to 2.3” and required the removal of “unsupported completion / performance claims.”
- Development mainline promoted: The day after the official tags landed, FlagSparse
mainwas bumped to 0.4.0, the FlagAudio version line was corrected to 0.4.0, and FlagAttention underwent a synchronized version update — the 2.2 release artifacts (rc2 baseline + official tags) and the development mainline (0.4.0 series) have formally diverged, and the 2.3 development cycle has effectively begun.
Component Tags and Release Manifest (Within the Window)
- Official tags (ten L1 components): flaggems v5.4.0, flagfft v0.2.0, flagsparse v0.3.0, flagdnn v0.3.0, flagblas v0.3.0, flagtensor v0.3.0, flagaudio v0.3.0, flagattention v0.4.0, flaggems-vllm v0.2.0, flaggems-sglang v0.1.0 (09-24).
- FlagTree: The 0.7.0 series was produced in bulk and continues to expand — three mainline lines at
0.7.0+triton3.6 / +triton3.5 / +triton3.3; platform variants including+mthreads3.6,+ppu3.6,+xpu3.6,+tsingmicro3.6 / +tsingmicro3.3,+ascend3.5,+iluvatar3.6,+metax3.6,+tileir3.6,+hcu3.6, and more than ten others (09-23 to 09-25) — platform-specific tags have expanded from “mainline” to “vendor-specialized” and are formally incorporated into the version system. - FlagCX: v0.14.0-rc2.post2 released (09-21 22:11) — fixes the
torch_txdalibrary directory resolution for Tsingmicro TSM, resolving TSM runtime image linking failures. - Other tag lines: FlagScale
v2.1.0-rc2(cherry-picked from main on 09-23) andv2.1.0-rc2.post2; vllm-plugin-FL0.3.0-rc2branch (including the 0.24.0 adaptation for tsingmicro txda); sglang-plugin-FL v0.2.0 (first entry in the official manifest).
Image and Artifact Status
- Application image matrix: The Megatron application image
2.2.0-0.2.3is registered across 14 platform rows (nvidia-cuda12.8 / 13.3, ascend-cann8.5.0 / 9.0.0, cambricon-neuware4.4.3 / 4.7.2, enflame-tops1.9.10 / 1.10.6, hygon-dtk26.04, iluvatar-corex4.4.0 / 4.5.0, metax-maca3.8.1.3, mthreads-musa4.3.6 / 5.2.0); the megatron_rl application row was unblocked and opened to all 18 backends (09-23, build-infra #1057). - vLLM line: The 0.20.2 line completed a 20/20 backend on-machine validation snapshot; the reasons three backends cannot run FlagTree 0.7.0rc2 were made public (MetaX
tl.dotICE, Enflame’s old toolchain rejectingenable_i64, and cuda13.3 depending on libcudart.so.12, each with 0.6.x comparisons and pinning records); the 0.24.0 line’s2.2.0-0.3.0rc2.post2records were written back. - sglang line: Seven platform application pages were published in a single night (kunlunxin-xre5.37.1, ascend-cann8.5.0, enflame-tops1.10.6, hygon-dtk26.04 dual pages, mthreads-musa4.3.5, iluvatar-corex4.5.0, generic-c13.0), and thead-ppu2.1.0 was updated to the official component tags (09-25); application-layer artifacts were simultaneously renamed to vendor-neutral naming.
- Release portal and build integration: The Release Info portal was rebuilt (09-23, covering base / runtime for 20 platform rows and the vLLM / Megatron / sglang application matrices); build-infra bumped the FlagGems version in the build matrix to 5.4.0 (09-24, linked to the official tags, merged approximately 7 hours later — a standard example of the “tag → build” sequence).
- Multi-track distribution channels: The FlagCX wheel channel expanded from pilot to all rows and landed on Kunlunxin (precompiled wheels carrying device bitcode); the unified noarch deb / rpm release channel covers every packaged component (FEP-0019 Wave 1); the third-party channel flagos-packaging released v2026.09.20 (apt / dnf dual formats, six-distribution repos, GPG signing).
III. Components and Code Co-building
This week, the org’s five-window commit counts were 163 / 224 / 143 / 211 / 175, totaling approximately 916, aggregated by module below.
Operators and Domain Libraries (FlagGems and various domain libraries)
- FlagGems (mainline, this week’s first line): Five-window commit counts of 34 / 29 / 28 / 59 / 46. Three production lines running in parallel: KernelGen bulk ingestion (approximately 60 Nvidia long-tail operators injected in one week, covering
inverse,_gather_sparse_backward,msort,linalg_inv_ex, the_ctc_lossfamily, the_cudnnfamily,histogramdd,_scaled_grouped_mm_v2, etc.); Kunlunxin preparation batch (31-operator optimization batch ingested in one pass, 27/31 reaching the 0.8x performance line, plus nearly 20 XPU fix batches covering the FlashAttention fast path, scaled_mm host fast path, convolution family, and special functions, with test machines migrated to klx-p800); TLE multi-backend rollout (Ascend glu usingtle.dsa.extract_slice, Iluvatar enabling TLE in vendor descriptor, Hygon W8A8 pipeline rewritten to segmented tl.range due to whitelist restrictions, TLE entering four backends within 48 hours). - Quantization and tuning: Hygon
fused_inv_rope_fp8_quant(DeepSeek-V4 attention path “inverse RoPE + FP8 group quantization” fusion), Hygon INT8 RMSNorm registration; MetaXtopk_w8a16_fp8; Moore Threads index_add / polar optimization; six-backend FP8 test completion (NVIDIA / Ascend / Hygon / XuanTie / MetaX / Iluvatar); FlagTune cost model extended to Hopper, MetaX MM, and DAMO XuanTie ZW810E (v1.1.0 under review), FP8 MM local tuning enabled. - Platform fix batches: Ascend (apply_rotary_pos_emb cos/sin duplication, batched linalg_solve_triangular incorrect results, log_ grid padding); Iluvatar (linalg_solve_triangular TLE path disabled,
index_reduce_); Hygon (complex scalar multiplication); MetaX (masked_fill out-of-bounds); Kunlunxin (special_bessel_y1 / softplus / weight_norm batch). - Engineering quality: Operator registration path failure fix, benchmark reference-only baseline validation, candidate pre-flight checks and in-process profiling hooks, weekly backend container image updates; continued SiliconFlow contributions (
_scaled_grouped_mmcross-backend fixes and tuning, adaptive_avg_pool2d optimization, index_reduce_ three-platform fixes). - FlagFFT: Hygon HCU adapter (see Top Stories); MACA line “1d single” measurement strategy switched to default, FP64 register swap experiment tried and reverted same day; packaging trio (Debian + RPM, Nexus release, 0.2.0-rc2 port); JIT artifact cross-process isolation and multi-process regression testing.
- FlagSparse: Collaboration with NCIC-AlphaSparse entering high-frequency phase (six batches merged in one week: MACA SpMV CSR baseline, DCU SpMM COO optimization, ASCEND numerical debugging, biv algorithm addition, MUSA delivery fixes); main branch bumped to 0.4.0 after v0.3.0 official tag; rpm packaging and pip compatibility fixes.
- FlagBLAS: Ascend L2 performance optimization in two batches (SPR / SPR2 / HPR / HPR2, banded packed and triangular, SYMV startup path simplification); Moore Threads L2 support merged; external contributor packaging line landed.
- FlagDNN / FlagTensor / FlagAudio / FlagAttention: FlagDNN restructured Iluvatar and Hygon platform implementation architecture, Moore Threads reusing compiled artifacts, packaging merge; FlagTensor 0.3.0-rc2 packaging line (deb / rpm, CMake finding Torch, Nexus upload); FlagAudio first three PRs merged and moving to implementation (spectrogram Triton operators, gain identity fix, openEuler installation fallback); FlagAttention v0.4.0 tag, test and benchmark updates.
- FlagTrain (new direction): First implementations landed after repo creation on 09-16 — Triton lamb training operators and test methods submitted by KMCompiler merged, positioned as “reusable training operators for mainstream training frameworks implemented in Triton.”
Compilers and Communication Libraries (FlagTree / FlagCX / libtriton_jit)
- FlagTree: MetaX backend specialization refactoring advanced overnight with ten commits (three-layer source specialization for TritonIR / TritonToTritonGPU / TritonGPUToLLVM, unified CMake entry point, Gluon layout interface); Iluvatar synced to Triton v3.6.x and extended TLE distributed and TLE_RAW (30-file large PR, distributed primitives supported by FlagCX); NVIDIA TLE added SMEM sub-slicing and multi-writer TMA pipeline; two AMD compiler fixes (exposed by W7900 precision testing); CI coverage expansion (910B workflow, mthreads3.6, hcu3.6, MetaX vLLM benchmark, PPU pid validation); packaging fixes (triton library name unified to
_C/libtriton.so, rpm packaging clang, deb PEP 440 compliant version numbers). - FlagCX: rc2.post2 fixes Qingwei TSM plugin; heterogeneous transport hardening and alignment with accelerator CI, p2p engine rpc serve extension, barex transport upgrade; rpm Requires excludes unversioned CANN sonames; CRL heterogeneous gather aggregation scheduling moved after transport; PAL supports SHCA verbs ABI and 17-bit LID; wheel distribution line (all lines + Kunlunxin).
- libtriton_jit: Multi-architecture script directory fixes, multi-backend packaging, nlohmann-json compatibility fixes, packaging trigger conditions narrowed — Triton JIT runtime converging toward multi-backend distribution form.
Inference and Training Frameworks (FlagGems-vllm / FlagGems-sglang / Torch-FL / plugin line / FlagScale / TE-FL / Megatron-LM-FL)
- FlagGems-vllm (four platforms in same window): Ascend XCS operator batch advancement (slot_mapping optimization measured full service round 1050.9 → 81.2 ms, kpool state compression, lightning indexer, kv_rmsnorm_rope_cache, etc.); Moore Threads shared INT4 / FP8 fused Marlin MoE and W8A8 variable-length attention; Hygon INT8 block-wise BMM and einsum contraction; MetaX fused inverse RoPE FP8 quantization and persistent_topk MC550 backend; DAMO XuanTie TLE version topk; Gemma RMSNorm completed across three platforms in one pass; “same operator multi-platform comparative implementation” became the standard working mode.
- FlagGems-sglang: Fused MoE router three-way merge (Kunlunxin and MUSA integration, Ascend two tensorcore variants, Enflame tensorcore dot-acc); Enflame GCU vendor registration (device query, TLE enablement, heuristic configuration).
- Torch-FL (engineering closure): DCU out-of-bounds indexing directly rejected, Ascend kernels changed to execute on the device where operands reside, PPU restores CUDA metadata for CPU torch wheel; build matrix converged to single table, GEMS_VENDOR unrecognized directly errors, removed global
Tensor.__getitem__patch and MetaX retired code; platform detection unified with capability matrix documentation added, version pins centralized to single file, import-time side effects converged into ordered phase pipeline. - vllm-plugin-FL: vLLM 0.28 upgrade line started (NVIDIA side merged, Ascend 910C adaptation #487 under review); Biren SUPA backend merged, XiWang PTPU ported to 0.24 mainline, operator-level profile tool merged; Hygon vLLM 0.24 workflow and MUSA CI aligned.
- FlagScale / TransformerEngine-FL / Megatron-LM-FL: FlagScale fixed metax parameters and advanced v2.1.0-rc2.post2, queue includes MemRift training compression, NPU profiler, DAMO XuanTie 810E training CI, Hygon DCU health sampling; TE-FL added NPU backend LayerNorm and multi-tensor Adam, supports MetaX vanilla TE package layout (and cherry-picked into rc2.post2), PPU unit and integration tests under review; Megatron-LM-FL stabilized vendor backend compatibility and CI, MetaX build jit_fuser kept no-op to prevent memory leak, XiWang PT-PU platform support.
Release Engineering and Governance (build-infra / community / docs)
- build-infra: Beyond version and image lines, this week also completed FlagCX wheel channel (including Kunlunxin), noarch deb / rpm unified channel, packaging pipeline split (release path build, native-rpm per-suite download, release version stamping), release CI five-fold hardening (pinned-version build, verify validation, symbol check, per-package enumeration, target repository validation), 910C RL E2E and NPU visibility fixes, sglang application page and portal maintenance.
- community: All three lines of 2.2 release governance actions (tags / manifests / scope) landed (#111 ~ #115); FEP documentation continuously revised per “release boundaries and evidence” criteria.
- docs and wiki: Four model list updates (ModelScope / HuggingFace, etc.); FlagTree wiki multi-platform user manual updated multiple times before release.
Quantum, Sparse, and Emerging Directions
- FlagQuantum (fastest-scaling direction this week): Two consecutive windows with 30+ commits (38 / 35, other windows 19 / 13, totaling 100+). Five lines in parallel: cloud QPU (Azure Quantum scaffolding, Quafu task API), cross-framework execution bridge matrix approaching completeness (Cirq, Qiskit Aer, PennyLane Lightning, Braket, CUDA-Q), simulation fusion acceleration (state-vector gates, dual-line fusion, shallow Clifford routing), validation gates (holdout set validation, drift alignment, evidence-constrained simulator selection advisor), CI reuse and parallelization.
- Open3D-PIMC: Media coverage matrix continued to expand after code landed (Xinhua Net 09-14 → CNR 09-18 → Guangming Net 09-23 → Sina Finance 09-24), core facts consistent with no new technical information; directly related to FlagOS’s 3D compute chip direction.
- FlagOS-Compressor: Linear INT8 export PR for DeepSeek V4.1 and MiMo V2.5 merged then reverted same day (commit message did not state reason), follow-up actions listed as observation item for this publication.
IV. Member Units and Vendor Adaptation
- BAAI (lead): Led 2.2 release governance (manifest / labels / scope audit), RC2 closure, and FEP release boundary definition; two FlagOS talks at AICC2026 (Lin Yonghua’s “From Architectural Innovation to Usable Compute” and Zhao Shuai’s special session) and co-building of the “Reliable AI” evaluation system; Open3D-PIMC joint development.
- Hygon: First HCU (BW1000) backend adapter for FlagFFT; FlagGems W8A8 TLE pipeline replacement, fused_inv_rope_fp8_quant, INT8 RMSNorm, mv optimization, complex scalar fixes; FlagSparse DCU checklist and spv precision; FlagScale DCU hardware health sampling (under review); FlagTree added hcu3.6 FlagGems baseline and test workflow; sglang hygon-dtk26.04 application page registered.
- MetaX: MXMACA merged into LMCache mainline on an “upstream-first” basis with a CI verification mechanism established; officially announced Day 0 adaptation for 39 flagship models since December 2025 (including Qwen-Image-2.1); on the code side, ten proprietary refactoring commits for FlagTree, plus multiple fixes beyond Iluvatar entering rc2.post2 content (TE-FL layout, jit_fuser, persistent_topk), topk_w8a16_fp8, MetaX vLLM benchmarks into CI,
0.7.0+metax3.6tag. - Moore Threads: RLinf official CI integrated MTT S5000 starting from v0.3 (training curve correlation coefficient 0.976, full-step time reduced 26%, idle rate 18.5% → 3.1%); Protenix-v2 full-pipeline inference averaged ~26% faster than mainstream international GPUs, completing the closed loop for three major AI4S scenarios; W8A8 variable-length attention and shared INT4 / FP8 fused Marlin MoE; AICC2026 talk and participation in co-building “Reliable AI”.
- Ascend: XCS operator batch (slot_mapping / kpool / indexer series) and 910B FlagGems baseline workflow; 910C dual-configuration RL end-to-end passed, NPU visibility fix; grouped_matmul added to operator manifest, rrelu_with_noise, cat/concat dispatch improvements; vLLM 0.28 adaptation under review; TE-FL NPU backend completed; multi-line image registration and sglang application page.
- Enflame: FlagGems-sglang registered GCU vendor and tensorcore dot-acc fused router; megatron application image line registered; plugin 0.24 line follow-up.
- Iluvatar CoreX: Iluvatar backend synchronized with Triton v3.6.x and enabled TLE distributed; wheel rebuilt on Ubuntu 22.04 and builder released; solve_triangular TLE path disabled; corex4.4.0 / 4.5.0 two image lines registered.
- Kunlunxin: 31-operator optimization batch (27/31 meeting targets) and nearly 20 XPU fix batch; test machine migrated to klx-p800; sglang kunlunxin-xre5.37.1 application page; FlagCX wheel landed on Kunlunxin.
- Tsingmicro: KernelGen 2.2.0 included as newly added hardware; FlagCX rc2.post2 fixed TSM torch plugin; FlagTree
+tsingmicro3.6 / +tsingmicro3.3tags; AICC2026 “Open3D-PIMC” talk and portal line update. - DAMO XuanTie (T-Head): Zhenwu V900 released (new-generation PPU); PPU entered the FlagOS build and application matrix (
0.7.0+ppu3.6, sglang thead-ppu application page, TE-FL PPU CI, FlagTree PPU pid validation); ZW810E FlagTune cost model v1.1.0 under review; SAIL software stack expanded open source, GitHub development branch to become the main R&D line. - Biren and Xiwang (new partners): Biren SUPA backend dual-layer landed on the same day (operator library + inference plugin); Xiwang PTPU patch set ported to vLLM 0.24 mainline (Qwen3.6 27B / 35B-A3B running normally).
- Cambricon: megatron application image
2.2.0-0.2.3landed on neuware4.4.3 / 4.7.2 two lines.
V. Ecosystem and Community
Conferences and Industry Events
- AICC2026 (09-21, Beijing): The 7th Artificial Intelligence Computing Conference — Lin Yonghua, Vice President and Chief Engineer of BAAI, delivered a report titled “From Architectural Innovation to Usable Compute: FlagOS Driving a New Paradigm of Software-Hardware Co-design for AI Computing”; BAAI researcher Zhao Shuai delivered a report titled “Zhongzhi FlagOS: Unlocking the Value of Diverse AI Compute with an Open Software Stack”; at the conference, the “Reliable AI” evaluation working group was launched and an evaluation framework was released (jointly by the Fifth Electronics Research Institute of MIIT, BAAI, Inspur, Moore Threads, MetaX, and more than 20 other organizations); IDC projects China’s intelligent compute scale to reach 2576.5 EFLOPS in 2026 (+87.9% YoY).
- Yunqi Conference (09-22, Hangzhou): DAMO XuanTie released Zhenwu V900 (see Top Stories); previewed Zhenwu J900 (March 2028) and Yitian 720 / 730 server CPUs (2027).
- Domestic compute software stack developments: Loongson Technology released the first software version of the “Loongson Accelerated Computing Platform” (based on the LG200 series GPU, covering drivers, compilers, runtime, operator libraries, and inference engines, supporting OpenCL 3.0 and CUDA migration); the 5th Global Digital Trade Expo opened on 09-23 in Hangzhou (featuring for the first time an integrated circuit application ecosystem section, with MetaX and others exhibiting).
- SAIL open-source progress (09-23): Lu Shenghua, Senior Director of DAMO XuanTie Software Ecosystem, disclosed — 39 low-precision quantized models have been open-sourced (over 348,000 downloads), SAIL has adapted to more than 260 frameworks, average adaptation time for mainstream inference frameworks is under 7 days, and the Zhenwu series has served over 650 enterprise customers.
Open-Source Releases and Coverage Matrix
- Open3D-PIMC: The open-source coverage matrix jointly developed by Tsingmicro and BAAI continues to expand (see Section III), described as “targeting over 30 domestic chips, exploring a global standard for 3D chip compilers.”
- flagos-packaging v2026.09.20: The third-party maintained FlagOS native package channel released its weekly update (APT / YUM dual format, six-distribution repos, GPG signing), advancing in parallel with the build-infra deb / rpm pipeline.
Competitions and Community
- FlagOS × MiniCPM Model Inference Throughput Performance Optimization Challenge: Development and submission phase opened on 09-21 (until 11-20, judging in December), with the challenge covering two evaluation scenarios: 4k / 16k.
- SGLang Cross-Chip Operator Optimization Competition: Competition results continue to enter the main repository as PRs (this week’s merged three-way MoE router merge is an extension of this).
- 2026 Artificial Intelligence Open Computing Conference & Zhongzhi FlagOS Technical Conference: To be held October 17–18 in Beijing, registration ongoing; workshop topics (KernelGen, edge-side, TLE Attention, FlagScale-Agent) correspond one-to-one with this week’s engineering observations.
News Coverage and External Channels
- Component-level news search: sixteenth through twentieth consecutive quiet windows: Across five full weekly windows, Chinese and English searches for component names including FlagOS / FlagGems / FlagScale / FlagTree / FlagPerf / FlagCX / KernelGen yielded zero valid hits; external information增量 was concentrated on member organizations and conference milestones (AICC2026, Yunqi, Digital Trade Expo).
- Inclusion boundary note: Stock market reports, gambling SEO, and general think-tank content are not counted; pure capital-market developments are recorded as background only.
VI. Trend Observations
- Release discipline completes a “three-layer freeze” evolution: Within a single week, the freeze sequence progressed through artifact freeze after code freeze (batch tag commits, rc2 manifest closure), then scope and declaration freeze (FEP scope audit, with retained items limited to verified configurations); shifting from “humans reading manifests” to “state derivable, evidence attachable.” No new features are pursued in the 3 days before GA; risk is concentrated on “installs correctly, runs correctly, describes accurately.”
- Multi-chip matrix continues to expand before GA, with onboarding paths now standardized: Biren and Xiwang appeared on the code side for the first time this week, forming an isomorphic path with Tsingmicro, DAMO XuanTie, Iluvatar CoreX, and others — “operator library + inference plugin + compiler platform tag” three layers landing on the same day; FlagTree 0.7.0 platform variants expanded to over ten, with platform tags formally incorporated into the version system. “Vendor specialization” is no longer outside the mainline but part of the release artifacts.
- Distribution formats complete multi-track evolution: Beyond container images, precompiled wheels (carrying device bitcode), noarch deb / rpm, and apt / dnf native packages now run in parallel (FEP-0019 Wave 1 landing); the “verifiability” of release artifacts (portal, snapshots, line-by-line records, pinned-version notes, vendor-neutral naming) has become a delivery surface as important as functionality.
- Operator libraries shift from “expansion” to “acceptance and tuning”: KernelGen long-tail batch (approximately 60 operators) and Kunlunxin 31-operator optimization batch (27/31 meeting targets) proceed in parallel; “same-operator multi-platform comparative implementations” and “same-platform batch fix batches” have become standard working modes; test infrastructure is simultaneously platformized (klx-p800 migration, mthreads3.6 / hcu3.6 workflows, six-backend FP8 comparative testing). Post-feature-freeze performance actions are released in a concentrated burst before GA.
- Non-mainline components establish their own cadence, with the 2.3 cycle starting early: FlagQuantum kicked off with 30+ commits in a two-week window, FlagSparse merged at high frequency with external collaborators, FlagFFT completed a new backend + packaging, and FlagTrain landed its first batch of operators; the day after the official tag was committed, the development mainline was bumped to 0.4.0 — the de facto fork between 2.2 delivery and 2.3 development has already occurred (new directions such as quantum and 3D compute are formally cut off from the 2.2 boundary).
- Industry narrative and code cadence are in sync: The conference season’s “unified software stack / diverse compute” narrative is being released at high density (AICC2026, Yunqi, SAIL, Loongson, Open3D-PIMC), mutually corroborating the closure actions on the code side; member organizations are simultaneously placing bets on both training (RLinf × S5000) and inference (Protenix-v2, Day 0 adaptation). The next reporting peak is expected at the 09-28 GA release window.
Appendix: Sources and Verification
- Sources (5 issues total, windows contiguous day by day with no overlap): 09-21 issue (window 09-20 10:18 to 09-21 10:18, 163 commit hits), 09-22 issue (09-21 10:18 to 09-22 10:00, 224 hits), 09-23 issue (09-22 10:00 to 09-23 10:00, 143 hits), 09-24 issue (09-23 10:00 to 09-24 10:00, 211 hits), 09-25 issue (09-24 10:00 to 09-25 10:00, 175 hits).
- Verification method: Full pushed_at verification across all GitHub org flagos-ai repositories (54 repos, with 22 / 19 / 24 / 23 / 24 active in each issue); commit search verified item by item by committer time (916 commits across the five issues); tags / releases checked repo by repo for official tags and post tags; community and build-infra manifests and PR texts compared item by item; Release Info portal checked page by page; news side covered by 24 / 25 sets of Google News RSS queries in Chinese and English plus community channel sweeps.
- Supplementary review at compilation (09-26): community #114 (L1 manifest) and #115 (scope audit) were still under review at compilation time; build-infra had, from the evening of 09-25 to 09-26, opened the first batch of build registrations for the Megatron 0.18.2 application line (application image
2.2.0-0.3.0landed for nvidia-cuda12.8, hygon-dtk26.04, metax-maca3.8.1.3 and other rows), and landed the FlagCX bitcode recipe for cuda13.3 and the metax training verification record (these are follow-on actions after this cycle, for tracking in the next issue). - Scope note: This is the 16th to 20th consecutive quiet window for component-level news; technical conclusions are based on code repositories and official releases; industry-side information includes vendors’ one-sided statements at conference venues (such as 39 Day 0 items and performance improvement ratios), all of which have been attributed in the daily reports; existing topics (Open3D-PIMC, etc.) are included based on incremental additions.
- Main references (searchable and locatable in the corresponding repositories by name):
- 2.2 release schedule — https://github.com/flagos-ai/community/blob/main/release/2.2/schedule.md
- 2.2 official manifest — https://github.com/flagos-ai/community/blob/main/release/2.2/release-2.2.yaml · rc2 manifest — https://github.com/flagos-ai/community/blob/main/release/2.2/release-2.2-rc2.yaml
- community governance PRs — https://github.com/flagos-ai/community/pull/111 · https://github.com/flagos-ai/community/pull/112 · https://github.com/flagos-ai/community/pull/113 · https://github.com/flagos-ai/community/pull/114 · https://github.com/flagos-ai/community/pull/115
- Release Info portal — https://flagos-ai.github.io/release-info/
- build-infra key PRs — https://github.com/flagos-ai/build-infra/pull/1057 · https://github.com/flagos-ai/build-infra/pull/1080 · https://github.com/flagos-ai/build-infra/pull/1081 · https://github.com/flagos-ai/build-infra/pull/1105 · https://github.com/flagos-ai/build-infra/pull/1107
- FlagTree tags — https://github.com/flagos-ai/FlagTree/tags · Iluvatar TLE — https://github.com/flagos-ai/FlagTree/pull/1184
- FlagGems Kunlunxin optimization batch — https://github.com/flagos-ai/FlagGems/pull/6412 · TLE — https://github.com/flagos-ai/FlagGems/pull/6522
- FlagCX rc2.post2 — https://github.com/flagos-ai/FlagCX/releases/tag/v0.14.0-rc2.post2
- FlagSparse version bump — https://github.com/flagos-ai/FlagSparse/pull/90 · FlagFFT — https://github.com/flagos-ai/FlagFFT/commits/dev
- FlagTrain first batch of operators — https://github.com/flagos-ai/FlagTrain/pull/1
- Biren dual-layer integration — https://github.com/flagos-ai/FlagGems/pull/6367 · https://github.com/flagos-ai/vllm-plugin-FL/pull/530 · Xiwang — https://github.com/flagos-ai/vllm-plugin-FL/pull/555
- AICC2026 — https://tech.gmw.cn/2026-09/22/content_39015304.htm · https://finance.eastmoney.com/a/202609233882244525.html
- Zhenwu V900 — https://www.guandian.cn/article/20260922/605518.html
- SAIL open-source progress — https://www.zhidx.com/p/596928.html
- RLinf × MTT S5000 — https://www.stcn.com/article/detail/4195458.html
- Protenix-v2 — https://tech.ifeng.com/c/8wgezlRjldv · https://www.yicai.com/news/103377730.html
- MetaX LMCache — https://www.metax-tech.com/ndetail/12657.html · LMCache PR — https://github.com/LMCache/LMCache/pull/4606
- MetaX 39 Day 0 items — https://www.jiemian.com/article/15120797.html
- Loongson accelerated computing platform — https://www.ithome.com/1/005/595.htm
- flagos-packaging — https://github.com/shiptux/flagos-packaging/releases/tag/v2026.09.20