FlagOS Daily Intelligence Report (2026-09-28)
Research window: 2026-09-25 10:00 ~ 2026-09-28 10:00 (approx. 72 hours; Monday window covering the weekend; continues from the 09-25 daily report window with no gaps) Sources: GitHub (org: flagos-ai, full pushed_at verification across 54 repos, 15 repos active within the window; 173 commits found across 11 repos via commit search, plus item-by-item review of 6 branch lines including FlagGems P1 test batches, FlagFFT dev, Megatron-LM-FL 0.3.0-rc2, two FlagTree feature branches, and FlagScale 2.1.0-rc2; item-by-item comparison of community #114/#115 and build-infra PR originals; item-by-item verification of release-info portal deployment records), Google News RSS (26 sets of Chinese and English query terms, via proxy), reports from Changsha Evening News / VOC / BAAI Community / ifeng / QbitAI / Sina Tech, etc. (see appendix source list for details)
Index
- Today’s Focus: 2.2 GA Day (09-28) — Final Preparations Within the Window: Megatron 0.18.2 Upgrade Chain Closure and v0.3.0 Wheel Relocation Re-verification (09-27)
- I. Open-Source Project Progress (GitHub Activity)
- 1.1 Release Engineering: GA Day Arrives — community Manifest Pending Merge; release-info Portal Deployed Automatically 5 Times Within the Window (09-25 ~ 09-26)
- 1.2 build-infra: Megatron 0.18.2 Upgrade Chain — RL E2E Three-Platform Verification and v0.3.0 Wheel Relocation (09-26 ~ 09-27)
- 1.3 build-infra (Part 2): Skills Framework Landed — Claude Code skills Carry Backend Integration and Change Gates (09-27)
- 1.4 FlagTree: iluvatar C++ Specialization Refactor Branch Continues; vLLM Performance Fix Branch Opened (09-25 ~ 09-28)
- 1.5 FlagGems: Five Consecutive P0/P1 Test Batches; Multi-Platform Optimization PR Queue (09-25 ~ 09-28)
- 1.6 FlagGems-vllm: chunk_kda Released Across Three Platforms; moe_sum and fused_add_rms_norm (09-28)
- 1.7 Torch-FL: Cross-Platform Fix Batch of 24 — DCU / Ascend / Enflame / MUSA Full Matrix (09-25 ~ 09-28)
- 1.8 FlagFFT: 18 Commits on dev Branch — Hygon / Iluvatar / MetaX Three-Platform FFT Optimization (09-25 ~ 09-27)
- 1.9 FlagSparse / FlagBLAS: Sparse and BLAS Multi-Platform Pipelines (09-25 ~ 09-27)
- 1.10 Other Activity: FlagCX / FlagQuantum / FlagScale / FlagDNN / docs (09-25 ~ 09-28)
- II. News Coverage and Ecosystem
- 2.1 Component-Level Search: Twenty-First Consecutive Quiet Window (09-25 ~ 09-28)
- 2.2 Yuelu Conference “AI+OpenCL”: Domestic OpenCL Operator Library Co-Construction Signing (09-24/09-25)
- 2.3 DAMO XuanTie: SAIL Second Batch Open-Source Preview — Late This Month to Early October (09-26)
- 2.4 Shishi Technology Meta-Infer: PCIe GPU DeepSeek Inference Nearly 7x Improvement (09-24, Previous Window Backfill)
- III. Member Unit Deep Dive
- 3.1 Ascend: Torch-FL Fix Batch and FlagGems Optimization Queue; 0.18.2 Image Line Pending Migration (09-25 ~ 09-28)
- 3.2 Hygon: DCU Fix Batch + HCU FFT Optimization + hygon RL E2E Verification (09-25 ~ 09-27)
- 3.3 MetaX: MACA FFT Optimization Batch + Metax RL E2E + Image Tag (09-26 ~ 09-27)
- 3.4 Moore Threads: chunk_kda Integration, FP64 L2 Support, and bmm Family Optimization Queue (09-25 ~ 09-27)
- 3.5 Iluvatar / Enflame / Kunlunxin: Iluvatar Three Lines in Parallel; Enflame Kernel Guard; Kunlunxin Optimization Queue (09-25 ~ 09-27)
- 3.6 Tsingmicro: “Space String” Computing Constellation Launched — Reconfigurable Chips Join the Space Industry Chain (09-25/09-26)
- IV. Summary and Trend Observations
- Appendix: Source Verification Table
- Complete Source List
Today’s Focus: 2.2 GA Day (09-28) — Final Prep Within the Window: Megatron 0.18.2 Upgrade Chain Closed Out and v0.3.0 Wheel Relocation Re-Verified
Date: 2026-09-27 Sources: community #115, build-infra #1181, #1145, #1144, release-info portal
The official schedule shows FlagOS 2.2’s GA date as 09-28 (the date of this report). As of this report’s window cutoff (10:00), two community-side release artifacts awaiting merge — the L1 operator library official manifest (#114) and the 2.2 FEP scope audit (#115) — remain in a pending-merge state (both opened on 09-24, with no new activity within the window), leaving the GA announcement ready for release at any time. The main build-side activity within the window was the Megatron 0.18.2 upgrade chain close-out:
- RL E2E validation across three platforms: #1144 records megatron rl 0.18.2 end-to-end passing under two nvidia configurations (cuda12.8 / cuda13.3); #1145 records passing via both the hygon-dtk26.04 and metax-maca3.8.1.3 paths — the 0.18.2 RL training chain has secured validation records across multiple platforms.
- v0.3.0 wheel relocation: Megatron-LM-FL merged three fixes on the 0.3.0-rc2 branch (#192 restoring the NullTokenizer bos and packed_seq guards for RL, #193 restoring v0.3.0’s paged-attention capability dispatch, #194 adding the flag_gems paging fallback for CUDA-oriented builds) and was repackaged; build-infra recorded the rebuild entry (#1167) and completed no-shim RL E2E re-verification on the “relocated v0.3.0 wheel” (#1181, closed out 09-27 23:02).
- Artifact refresh: The application image tag
2.2.0-0.3.0was recorded in bulk across four platforms — nvidia-cuda12.8 / nvidia-cuda13.3 / hygon-dtk26.04 / metax-maca3.8.1.3 (over 20 tag records submitted within the window); the release-info portal auto-deployed 5 times within the window (09-25 10:17 to 09-26 15:15), with five entries in the 0.18.2 series for megatron_rl / megatron_training listed.
Assessment: The final window before GA contained no new feature announcements; all activity centered on “validation close-out + artifact rebuild + record-keeping” — consistent with the tone of the earlier scope audit (#115): external claims retain only “implemented + validated,” and 2.2’s public delivery will be grounded in an evidence chain reproducible from the artifacts.
1. Open-Source Project Progress (GitHub Activity)
1.1 Release Engineering: GA Day Arrives — community Manifest Pending Merge; release-info Portal Auto-Deploys 5 Times Within Window (09-25 ~ 09-26)
Date: 2026-09-25 ~ 2026-09-26 Source: community #114, #115, release-info portal, release-info deployment log
- community status: The L1 operator library official manifest (#114) and the 2.2 FEP scope audit (#115) saw no new activity within the window and remain pending merge; the release branch
release/2.2-official-l1is on record (opened 09-24) — merging the GA official manifest is the next visible milestone. - Portal refresh: release-info (the build infrastructure portal) completed 5 auto-deployments within the window (09-25 10:17, 09-26 09:50 / 10:06 / 13:06 / 15:15), with content continuously tracking the build-infra mainline; the portal now includes 0.18.2-series entries for megatron_rl / megatron_training (cambricon-neuware4.7.2, generic-12.8, generic-13.3, hygon-dtk26.04, metax-maca3.8.1.3) and application pages for both the vllm 0.20.2 / 0.24.0 series.
- Implication: Release engineering has entered a state where “the portal is the mirror manual” — documentation, image manifests, and verification records scroll and refresh in sync, preparing for external delivery after GA.
1.2 build-infra: Megatron 0.18.2 Upgrade Chain — RL E2E Three-Platform Verification and v0.3.0 wheel Relocation (09-26 ~ 09-27)
Date: 2026-09-26 ~ 2026-09-27 Source: build-infra #1144, #1145, #1146, #1167, #1181, Megatron-LM-FL #194 (99 commits in total, see appendix)
- RL E2E verification chain: #1144 (nvidia cuda12.8 / cuda13.3 dual config) and #1145 (hygon-dtk26.04 / metax-maca3.8.1.3 dual path) successively recorded megatron rl 0.18.2 end-to-end passing; the day before (09-26) the training matrix closed out first — #1131 marked metax training verification and recorded the cambricon root cause, #1132 completed the cuda13.3 flagcx bitcode compilation recipe (F/T training verification).
- Verification workbook: The 0.18.2 tracking document (starting from #1119) and the verification workbook (#1126 / #1128 / #1130 series) record each platform’s matrix status as a “single source of truth” — release verification now has a unified ledger.
- v0.3.0 wheel relocation: After the three fixes on the Megatron-LM-FL 0.3.0-rc2 branch (#192 ~ #194) were merged, the wheel was rebuilt; build-infra recorded the rebuild entry (#1167 “pending entries for the #194 wheel rebuild”), and #1181 subsequently completed no-shim RL E2E re-verification on the relocated artifact — the “fix - repackage - re-verify” loop was completed within a single window.
- flagcx wheel build chain: #1133 pinned the builder clang to the flagtree built-in LLVM snapshot; #1134 deprecated the FLAGCX_BITCODE_PATH transitional approach for cuda13.3; #1135 changed the validation script to first uninstall the pre-installed flagcx; #1136 made the cuda12.8 runtime image carry the flagcx wheel — engineering details of distributing the communication component with the image are being closed out item by item.
- Implication: The platform migration from 0.17.1 to 0.18.2 (training + RL) completed main-body verification before GA, and is the engineering main axis of this release cycle’s closing phase.
1.3 build-infra (Part 2): Skills Framework Lands — Claude Code skills Carry Backend Onboarding and Change Gates (09-27)
Date: 2026-09-27 Source: build-infra #1155, #1156, #1166, #1176
- skills framework foundation: #1155 “land the Claude Code skills framework foundation” merged — release engineering processes begin to be distilled into a skill set executable by AI coding assistants.
- Onboarding process turned into skills: #1156 added the add-backend and verify-app-backend process skills, subsequently refactored into the update-backend flow (with scope modeled per the “layer propagation chain”) — “new chip backend onboarding + application image verification” becomes a reusable standard process for agents.
- Gate refinement: #1166 scoped the changelog gate to application image push scenarios; #1176 decoupled
--no-cachefrom push logic — automation processes are calibrated item by item through real iterations. - Observation: This echoes the Skills / SkillHub narrative of the FlagOS 2.0 release both internally and externally — agent skills are not only aimed at external developers, but are also beginning to feed back into FlagOS’s own development infrastructure.
1.4 FlagTree: iluvatar C++ Specialization Refactor Branch Continues; vLLM Performance Fix Branch Opened (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: FlagTree refactor/iluvatar-cpp-specialization-v2, fix_perf_vllm
- iluvatar C++ specialization refactor (v2): Three commits migrating and refactoring the Iluvatar platform C++ specialization source — “move specialized C++ source”, “restore core CMake source wiring”, “reuse main Triton rewriter definitions” (09-25) — a continuation of the previous branch of the same name, organizing the specialization engineering path for the iluvatar backend.
- fix_perf_vllm: A new branch opened on the morning of 09-28 — two commits, “[CI] Fix perf.sh” and “fix pre-commit”, pointing to fixes for the vLLM performance benchmark script (09:37 / 09:47).
- Main branch silent: No commits on the main branch within the window (the latest being the 09-24 MetaX vLLM benchmark workflow), and no new 0.7.0-series tags — work during the GA freeze period has shifted to feature branches.
- Observation: Main branch frozen while refactors and fixes proceed on branches — another example of “release window and development window separation”.
1.5 FlagGems: Five Consecutive P0/P1 Test Batches; Multi-Platform Optimization PR Queue (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: FlagGems #6687, #6689, #6691, #6694, #6709 (30+ PRs in total, see appendix)
- P0/P1 test campaign: The test batch work started on 09-24 concentrated upstream within the window — five P1 test PRs opened together (#6687 sparse reduction / softmax / cuDNN, #6688 sparse log-softmax / sum / addmm, #6689 sparse matmul / fused convolution / Cholesky / QR / Bessel, #6691 FFT / psi / linear algebra, #6694 ISTFT / sparse reduction / cuDNN CTC / semi-structured apply), all with replayable benchmarks; the earlier P0 suite was withdrawn from review (#6670) and switched to batched PRs.
- Main branch: Only 1 merge within the window (#6709 fixing the pre-commit failure in the shifted Chebyshev test) — clear traces of the release freeze.
- Multi-platform PR queue (active within the window): Ascend addmm / bmm / baddbmm / mv / mul optimizations and index_put single-element indexing fix (#6697 ~ #6703 range); Moore Threads bmm / addmm / baddbmm (#6705 ~ #6707); MetaX MM dispatch and tuned kernels (#6663); Hygon addmm FlagTune and conv2d (#6696 / #6700); Kunlunxin batch two (#6659 / #6662); Sophgo backend optimization (#4615 / #4632); KernelGen Nvidia operator migration queue (#6608 ~ #6635, six commits: dim, crow/ccol_indices_copy, _dim_arange, _dimV, _values, etc.).
- Assessment: The rhythm of “add tests before release, take on performance after release” is clearly visible at the PR level — the test batch and platform optimization queues advance in parallel.
1.6 FlagGems-vllm: chunk_kda Launched Across Three Platforms; moe_sum and fused_add_rms_norm (09-28)
Date: 2026-09-28 Source: FlagGems-vllm #856, #859, #860, #855, #850, #846
- chunk_kda launched across three platforms (morning of 09-28): Moore Threads (#856), Hygon TLE version (#859), and Iluvatar (#860) successively merged the chunk_kda operator — a linear attention inference operator completed across three platforms in a single morning.
- Other operators: Iluvatar moe_sum (#855); Ascend fused_add_rms_norm performance optimization (#850); Hygon TLE-optimized topk_softplus_sqrt (#846); mhc_post new layout (#854).
- Background: FlagGems-vllm is the FlagGems operator plugin for vLLM (v0.2.0 is a component of the 2.2 official L1 manifest) — it maintains a daily operator completion pace even during the GA freeze, continuously densifying platform coverage on the inference side.
1.7 Torch-FL: 24 Cross-Platform Fixes — Full DCU / Ascend / Enflame / MUSA Matrix (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: Torch-FL #450, #453, #447, #449, #437, #448 (24 commits in total, see appendix)
- Release artifact alignment: #447 publishes and validates the wheel compatibility manifest; #420 CI uniformly installs the released FlagOS wheel; #440 CI venv constraints — release artifacts and CI environment aligned first.
- DCU (Hygon): #453 “pins” scatter_add_ / argmin / mean series to CUDA boxing; #437 FlagGems device names no longer align only with CUDA.
- Ascend: #450 fixes the Qwen-Image path and adds a verified regression profile; #436 SDPA realShift supports finite additive attention bias; #426 opens the device-side profiler contract.
- Enflame (GCU): #438 rejects 64-bit Triton kernels before the vendor compiler aborts; #429 int64 escapes to vendor kernels; #423 new_ones provided by vendor kernels — a guard-style engineering paradigm of “reject early / escape”.
- Distributed and general: #435 gloo requests answered by the flagos backend, collective operations via host memory; #434 MSPTI external correlation stack popped with real addresses; #421 flagos backend supports DataParallel; #449 flex_attention on flagos devices; #410 compressed sparse (CSR / CSC / BSR / BSC) tensor support; #416 CUPTI runtime cbid table auto-generation.
- Implication: The 24 commits cover three lines — “release artifact alignment + platform fixes + general capabilities” — Torch-FL, as the unified plugin on the torch side, is adding platform-level evidence for each capability declared in 2.2.
1.8 FlagFFT: 18 Commits on dev Branch — Hygon / Iluvatar / MetaX Three-Platform FFT Optimization (09-25 ~ 09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: FlagFFT dev branch commits
- Hygon (HCU): Five-commit CT leaf optimization batch — radix-16 leaf comparison, short CT leaf optimization and launch overhead cleanup, leaf scheduling restricted to 1D requests, tuning not spilling over to decomposition child nodes (09-25).
- Iluvatar (IX): Ten-plus multi-wave optimizations — real CT factor matching, N210 packing, 1024-point dual-warp scheduling, batched CT scan, hardware-table R2C twiddle, fused R2C leaf enabled by default after measured validation (09-26 ~ 09-27).
- MetaX (MACA): Single-transform C2R packed real path and measured gains documented (09-27).
- Implication: FFT is a representative operator domain for the scientific computing direction (the library’s README already lists seven backend types: CUDA / MUSA / PPU / IX / MACA / NPU / HCU) — three domestic platforms advancing within the same window, echoing the 2.2 domain expansion “from large models to scientific computing”.
1.9 FlagSparse / FlagBLAS: Sparse and BLAS Multi-Platform Pipelines (09-25 ~ 09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: FlagSparse #101, #100, #98, FlagBLAS #135, #134, #133
- FlagSparse (NCIC-AlphaSparse collaboration line): Multiple rounds of merges across #98 ~ #101 — DCU (Hygon) spmv CSR precision and bandwidth, COO-version SpMV / SpMM, MACA (MetaX) sddmm baseline, XPU (Kunlunxin) gather fp16, BIV algorithm documentation restoration and CI refinement; the 0.4.0 development cycle continues to advance.
- FlagBLAS: #135 Moore Threads FP64 L2 support; #134 / #133 DAMO XuanTie L2 equivalent-performance extension and GEMV / L2 baseline improvement; #132 Ascend L2 optimization (removing platform-specific skip items) — the L2 layer (matrix-vector class) has become the common performance frontline for all three platforms.
1.10 Other Activity: FlagCX / FlagQuantum / FlagScale / FlagDNN / docs (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: FlagCX #629, #628, FlagQuantum #221, FlagScale #1303, docs #523
- FlagCX: #629 CRL runner correctness baseline; #628 CRL unified transport dispatch and P2P path negotiation fix; #627 PAL network adapter shared transport foundation; #626 PAL IBUC reliability hardening and CUDA adaptor CI — foundational engineering progress of the communication library on the two experimental lines CRL / PAL.
- FlagQuantum: #220 batched product-state Clifford layer; #221 deferred product-state SWAP materialization — the quantum component continues its CPU-side optimization batch.
- FlagScale: #1303 fixes flagcx non-heterogeneous initialization (main, 09-28 09:21), and cherry-picks to the 2.1.0-rc2 branch (#1304).
- FlagDNN: Hygon compatibility adjustments (09-25).
- docs: FlagPrism documentation aligned with upstream updates (#523, 09-27) — documentation for the new component is completed before the 2.2 release.
II. News Coverage and Ecosystem
2.1 Component-Level Search: Twenty-First Consecutive Quiet Window (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: Google News RSS (26 Chinese and English query terms, via proxy), HN Algolia
- Component-name queries (FlagOS / FlagGems / FlagScale / FlagTree / FlagPerf / FlagCX / KernelGen / FlagOS 2.2) returned zero hits within the 72-hour window — the twenty-first consecutive quiet window on the news side.
- 18 ecosystem-side hits within the window: follow-up coverage of the Yuelu Conference OpenCL and DAMO XuanTie SAIL was retained by topic (see 2.2 / 2.3); the remainder were stock-market pieces (Enflame IPO-related, semiconductor sector), think-tank content, and unrelated reprints, filtered out by topic.
- HN Algolia returned no relevant hits for FlagOS / FlagGems / FlagScale / BAAI.
- The quiet window has now become the norm across three release cycles (2.0 / 2.1 / 2.2): external information increments cluster around release milestones and conference milestones, while the day-to-day rhythm is driven by the code side and the community side.
2.2 Yuelu Conference “AI+OpenCL”: Joint-Build Agreement for Domestic OpenCL Operator Library (09-24/09-25)
Date: 2026-09-24 ~ 2026-09-25 Source: Changsha Evening News (Zhangshang Changsha), Voc.com.cn (reprinted by ifeng)
- At the “AI+OpenCL Heterogeneous Computing” themed forum of the 2026 Internet Yuelu Conference (09-24, Changsha): Changsha Quanying Semiconductor, the Hunan Integrated Circuit Industry Technology Innovation Strategic Alliance, CSDN, Jingjia Micro, and Shouzhao Technology signed a strategic cooperation on an OpenCL operator library — jointly building a unified domestic OpenCL operator library to directly address the “strong hardware, weak software” pain point of domestic chips, avoid “reinventing the wheel,” and lower developers’ cross-platform porting costs.
- Forum lineup: Khronos explained the global OpenCL standards system (over 450 devices certified, with domestic chips including Jingjia Micro and Chengying GPGPU already adapted); scholars from Tsinghua University and Fudan University shared open-source GPGPU software stacks and edge-side deployment; Sophgo, Lisuan Tech, and others shared practices including financial quantitative operator libraries and AI-assisted integrated circuit design.
- Relation to this stack: A multi-track parallel effort to “fill in the software stack” for domestic compute — the OpenCL open-standard track (cross-vendor, cross-hardware compatibility) and the FlagOS “unified plugin stack” track serve as ecosystem counterparts to each other; this also continues the recent “software foundation” narrative (previous windows already saw similar moves such as Loongson’s full-stack platform and Weijing).
2.3 DAMO XuanTie: SAIL Second Batch Open-Source Preview — Late This Month to Early October (09-26)
Date: 2026-09-26 Source: BAAI Hub, ifeng (Dafeng Hao)
- Follow-up: After the 09-23 Yunqi Conference announcement of SAIL open-source progress (core facts already covered in the 09-24 issue of this daily), the window saw continued dissemination via BAAI Hub and other channels, along with incremental information — the second batch of SAIL open-source code will be released late this month to early October: coverage includes TensorFlow / JAX framework adaptation, the FLA / acext acceleration libraries, ppu-gdb / ppu-dcgm debugging and cluster operations tools, and more communication libraries.
- Already open-sourced list (media recap): PyTorch-for-sail framework adaptation, the sailify source migration tool, the Triton-for-sail operator development tool, DeepGEMM-for-sail, FlashAttention-for-sail; “the GitHub-based development branch will serve as the main R&D line.”
- Relation to this stack: SAIL targets Zhenwu PPU (i.e., the PPU support platform of FlagOS); the two “open-source software stack” tracks — vendor-proprietary stacks and the community unified stack — are advancing in the same period — PPU’s build entries on the FlagOS side (FlagTree ppu3.6 tag, sglang thead-ppu application page) can be tracked and compared on an ongoing basis.
2.4 METASTONE Meta-Infer: Nearly 7x DeepSeek Inference Improvement on PCIe GPUs (09-24, backfilled from previous window)
Date: 2026-09-24 Source: QbitAI, Sina Tech
- METASTONE released its self-developed Meta-Infer inference engine: pure software optimization (kernel completion, operator optimization and communication refactoring, parallelism and capacity tuning, cache reuse) unlocks the potential of PCIe-Only GPUs — in an 8-GPU environment, DeepSeek-V4.1-Flash input throughput went from 1932 → 13274 tok/s (about 6.87x), with support for 1M context; DeepSeek-V4-Flash improved 1.55x; the company also claims the same methodology has been validated on domestic GPUs (vendor-reported data, not independently verified by third parties).
- Methodology highlights: frameworks such as vLLM / SGLang silently degrade on non-certified hardware (mismatched metadata definitions, page sizes, and communication fast-path thresholds); fixing the adaptation alone unlocks several times the performance — “first confirm whether the high-performance kernels are actually executing, then talk about scaling.”
- Relation to this stack: A third-party practice addressing the same proposition as FlagOS — “a unified software stack unlocks the potential of diverse hardware”; its judgment that “long-tail hardware adaptation is an underrated ecological niche” serves as a reference point alongside this stack’s multi-chip track. Backfill note: this report first appeared on the evening of 09-24, after the previous daily’s window had closed.
III. Deep Dive into Member Units
3.1 Ascend: Torch-FL fix batch and FlagGems optimization queue; 0.18.2 image line pending migration (09-25 ~ 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: Torch-FL #450, #436, #426, FlagGems #6697, FlagGems-vllm #850, release-info portal
- Operator layer: Six active PRs in the FlagGems Ascend optimization queue during the window (addmm / bmm / baddbmm / mv / mul optimizations + index_put single-element indexing fix, in the #6697 ~ #6703 range); FlagGems-vllm fused add rms norm performance optimization (#850).
- Framework layer: Torch-FL fixes the Ascend Qwen-Image path and adds a verified regression profile (#450), SDPA realShift supports finite additive bias (#436), device-side profiler contract (#426) — Ascend’s torch-side gap-filling continues.
- Image line: The megatron 0.18.2 series in the portal does not yet include Ascend entries (still the 0.17.1-ascend-cann8.5.0 / 9.0.0 series) — migration of 0.18.2 to the Ascend platform is the next observation point.
3.2 Hygon: DCU fix batch + HCU FFT optimization + hygon RL E2E validation (09-25 ~ 09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: Torch-FL #453, FlagGems-vllm #859, FlagSparse #101, build-infra #1145, FlagGems #6696, FlagFFT dev commits
- Operators and framework: FlagSparse DCU spmv CSR precision and bandwidth, COO family (#101 and other batches); FlagGems-vllm chunk_kda TLE version and topk_softplus_sqrt (#859 / #846); Torch-FL DCU weight “pinning” CUDA boxing (#453) and device name alignment (#437); FlagGems addmm FlagTune and conv2d optimizations enter the PR queue (#6696 / #6700).
- Scientific computing: FlagFFT HCU CT leaf optimization batch (see 1.8).
- Training chain: megatron rl 0.18.2 hygon RL E2E dual-path validation (build-infra #1145) + hygon-dtk26.04 image tag
2.2.0-0.3.0— Hygon is one of the first platforms validated on 0.18.2. - Implication: Hygon DCU / HCU spans four layers in this window — operator library, FFT, framework, and training chain — making it the platform with the deepest reach in the window.
3.3 MetaX: MACA FFT optimization batch + Metax RL E2E + image tag (09-26 ~ 09-27)
Date: 2026-09-26 ~ 2026-09-27 Source: FlagFFT dev commits, build-infra #1145, Torch-FL #427, FlagGems #6663
- FlagFFT MACA single-transform C2R packed real path and benefit documentation (09-27); megatron rl 0.18.2 metax RL E2E dual-path (#1145) + metax-maca3.8.1.3 image tag
2.2.0-0.3.0; Torch-FL fixes the fp32 matmul precision pin for the MetaX integration task (#427); FlagGems MetaX MM dispatch and tuned kernel PR (#6663); FlagSparse maca sddmm baseline merged. - Observation: MetaX is synchronized across multiple lines — “training chain + FFT + operators” — placing it alongside Hygon as one of the two most active platforms in the window.
3.4 Moore Threads: chunk_kda integration, FP64 L2 support, and bmm family optimization queue (09-25 ~ 09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: FlagGems-vllm #856, FlagBLAS #135, FlagGems #6705, #6706, #6707
- FlagGems-vllm chunk_kda Moore Threads version merged (#856); FlagBLAS adds FP64 L2 support (#135); FlagGems PR queue has three bmm / addmm / baddbmm optimizations (#6705 ~ #6707, Opus-family automated flow
codex/mthreads-*branches). - Observation: The MUSA platform continues its dual-line push into “inference operators + BLAS foundation library,” and FP64 support fills the double-precision gap at the L2 layer.
3.5 Iluvatar / Enflame / Kunlunxin: Iluvatar runs three lines in parallel; Enflame kernel guards; Kunlunxin optimization queue (09-25 ~ 09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: FlagTree refactor branch, FlagGems-vllm #860, Torch-FL #438, FlagSparse commits, FlagGems #6662
- Iluvatar (IX): FlagTree C++ specialization refactor v2 branch (see 1.4); FlagGems-vllm chunk_kda and moe_sum (#860 / #855); FlagFFT IX optimization batch (see 1.8) — compilation, inference operators, and FFT running in parallel.
- Enflame (GCU): Three guard-style fixes in Torch-FL (early rejection of 64-bit Triton kernels #438, int64 escape #429, new_ones vendor kernel #423) — an engineering paradigm of “intercepting before the vendor compiler aborts.”
- Kunlunxin (XPU): FlagGems batch-two optimization PRs (#6662 / #6659) + FlagSparse XPU gather fp16 — entering optimization iterations after the large fix batch on 09-24.
3.6 Tsingmicro: “Space Strings” computing constellation launched — reconfigurable chips join the space industry chain (09-25/09-26)
Date: 2026-09-25 ~ 2026-09-26 Source: Zhejiang Daily (Zhejiang Online), Ifeng reprint
- The 5th Global Digital Trade Expo · Oriental Space Computing Industry Conference (09-25, Hangzhou): Geespace II and Oriental Star Chain jointly announced the “Space Strings” computing constellation — planning a network of over a thousand AI satellites to explore space-based data processing, model deployment, and computing resource coordination; Tsingmicro participates deeply as an industry chain partner (Zhejiang spaceborne AI chip industry chain).
- Background: Tsingmicro and Oriental Star Chain established a strategic partnership at the end of 2025; the “reconfigurable chip space agent” was announced at WAIC 2026, with the first on-orbit test planned for completion by the end of 2026 — the “reconfigurable chips go to space” path continues to advance.
- Relation to this stack: Tsingmicro is a member unit of this stack (FlagTree already has tsingmicro3.3 / 3.6 platform variants, and Open3D-PIMC is also a joint result with BAAI) — space computing is an extension of its reconfigurable chip scenario line, echoing the stack’s long-term direction of “cloud-to-edge full-scenario” coverage.
IV. Summary and Trend Observations
- GA-day wrap-up, release discipline throughout: No new feature announcements within the window; all activity centered on validation wrap-up, artifact rebuilds, and record-keeping; the two community checklists (#114 / #115) are about to be merged — the public delivery of 2.2 will be drafted on a foundation of “evidenced completeness.”
- Megatron 0.18.2 upgrade chain takes shape: RL E2E validated on nvidia×2 / hygon / metax + no-shim re-verification after v0.3.0 wheel relocation + batch recording of four-platform image tags — version advancement for training and RL scenarios is the engineering main axis of the wrap-up phase, and a prelude to the next release cycle (the 0.3.0 line).
- Release engineering becomes “skill-ified”: The Claude Code skills framework lands in build-infra, with backend integration (add-backend / update-backend) and change gates distilled into executable skills — agent capabilities begin to flow back into FlagOS’s own development infrastructure, echoing the community Skills narrative both internally and externally.
- Test infrastructure blitz: Five consecutive FlagGems P0/P1 test batches (sparse, convolution, FFT, linear algebra, cuDNN CTC, etc., all with replayable benchmarks) — quality is front-loaded into the freeze period, laying the foundation for 2.3’s multi-platform matrix.
- Multi-platform development lines in full bloom: FlagSparse (DCU / MUSA / XPU), FlagBLAS (DAMO XuanTie / Ascend / Moore Threads), FlagGems-vllm (Iluvatar / Hygon / Moore Threads), FlagFFT (Hygon / Iluvatar / MetaX), Torch-FL (full matrix) — the dual-track state of “release freeze and development in parallel” appears simultaneously across five components.
- Ecosystem-side contrast: The 21st quiet window on the news front contrasts with off-stage activity — OpenCL operator library co-building, the “SAIL second batch” teaser, and the Meta-Infer “nearly 7x” case all point to the same proposition: the value realization of diverse hardware increasingly depends on software stack capability. The window dividend of a unified software stack moves in step with the competition.
Appendix: Source Verification Table
| Category | Source | Verification Method | Result |
|---|---|---|---|
| GitHub | org: flagos-ai repos API | Full verification of pushed_at across 54 repos | 15 repos active within window |
| GitHub | Commit search + per-repo cross-check | Item-by-item verification within window (incl. branch lines) | Commit search: 173 commits across 11 repos; plus 6 branch lines verified (FlagGems P1 batch / FlagFFT dev / Megatron-LM-FL 0.3.0-rc2 / FlagTree ×2 / FlagScale 2.1.0-rc2) |
| Release engineering | community #114 / #115 original text | Item-by-item comparison of status and branches | Both still pending merge; release/2.2-official-l1 branch on record |
| Release engineering | build-infra PR and release-info deployment | Original text retrieval and cross-check | Megatron 0.18.2 chain closed out; 5 automatic deployments within portal window |
| Tags | FlagTree / FlagGems / FlagScale / FlagBLAS etc. tags atom | Item-by-item verification | No new tags within window (work moved to branches and PR queue) |
| News | Google News RSS (26 query sets in Chinese and English, via proxy) | 72-hour window filter + item-by-item exclusion | Zero hits on component terms (21st quiet window); 3 items retained on ecosystem side |
| Media | Changsha Evening News / VOC / BAAI Hub / ifeng / QbitAI / Sina Tech | Original text retrieval and review | OpenCL operator library / SAIL follow-up / Meta-Infer (supplementary) |
| Community | HN Algolia | Search review | No relevant hits |
| Historical comparison | FlagOS daily reports from previous three days | Deduplication verification | SAIL is a follow-up (increment clear); no additions to Open3D-PIMC; stock market coverage excluded |
Complete Source List
- [1] Today’s Highlights: GA Day and Megatron 0.18.2 upgrade chain — https://github.com/flagos-ai/community/pull/115 · https://github.com/flagos-ai/community/pull/114 · https://github.com/flagos-ai/build-infra/pull/1181 · https://github.com/flagos-ai/build-infra/pull/1144 · https://github.com/flagos-ai/build-infra/pull/1145
- [2] 1.1 Release Engineering — https://github.com/flagos-ai/release-info/commits/gh-pages · Portal https://flagos-ai.github.io/release-info/
- [3] 1.2 build-infra Megatron 0.18.2 — https://github.com/flagos-ai/build-infra/pull/1146 · https://github.com/flagos-ai/build-infra/pull/1167 · https://github.com/flagos-ai/build-infra/pull/1131 · https://github.com/flagos-ai/build-infra/pull/1132 · https://github.com/flagos-ai/build-infra/pull/1133 · https://github.com/flagos-ai/build-infra/pull/1134 · https://github.com/flagos-ai/build-infra/pull/1135 · https://github.com/flagos-ai/build-infra/pull/1136 · https://github.com/flagos-ai/Megatron-LM-FL/pull/192 · https://github.com/flagos-ai/Megatron-LM-FL/pull/193 · https://github.com/flagos-ai/Megatron-LM-FL/pull/194
- [4] 1.3 Skill Framework — https://github.com/flagos-ai/build-infra/pull/1155 · https://github.com/flagos-ai/build-infra/pull/1156 · https://github.com/flagos-ai/build-infra/pull/1166 · https://github.com/flagos-ai/build-infra/pull/1176
- [5] 1.4 FlagTree Branches — https://github.com/flagos-ai/FlagTree/tree/refactor/iluvatar-cpp-specialization-v2 · https://github.com/flagos-ai/FlagTree/tree/fix_perf_vllm
- [6] 1.5 FlagGems — https://github.com/flagos-ai/FlagGems/pull/6687 · https://github.com/flagos-ai/FlagGems/pull/6688 · https://github.com/flagos-ai/FlagGems/pull/6689 · https://github.com/flagos-ai/FlagGems/pull/6691 · https://github.com/flagos-ai/FlagGems/pull/6694 · https://github.com/flagos-ai/FlagGems/pull/6709 · https://github.com/flagos-ai/FlagGems/pull/6670 · https://github.com/flagos-ai/FlagGems/pull/6671 · https://github.com/flagos-ai/FlagGems/pull/6663 · https://github.com/flagos-ai/FlagGems/pull/6696 · https://github.com/flagos-ai/FlagGems/pull/6700 · https://github.com/flagos-ai/FlagGems/pull/6662 · https://github.com/flagos-ai/FlagGems/pull/6659 · https://github.com/flagos-ai/FlagGems/pull/6697 ~ https://github.com/flagos-ai/FlagGems/pull/6703 · https://github.com/flagos-ai/FlagGems/pull/6705 · https://github.com/flagos-ai/FlagGems/pull/6706 · https://github.com/flagos-ai/FlagGems/pull/6707 · https://github.com/flagos-ai/FlagGems/pull/4615 · https://github.com/flagos-ai/FlagGems/pull/4632
- [7] 1.6 FlagGems-vllm — https://github.com/flagos-ai/FlagGems-vllm/pull/854 · https://github.com/flagos-ai/FlagGems-vllm/pull/855 · https://github.com/flagos-ai/FlagGems-vllm/pull/856 · https://github.com/flagos-ai/FlagGems-vllm/pull/859 · https://github.com/flagos-ai/FlagGems-vllm/pull/860 · https://github.com/flagos-ai/FlagGems-vllm/pull/850 · https://github.com/flagos-ai/FlagGems-vllm/pull/846
- [8] 1.7 Torch-FL — https://github.com/flagos-ai/Torch-FL/pull/450 · https://github.com/flagos-ai/Torch-FL/pull/453 · https://github.com/flagos-ai/Torch-FL/pull/447 · https://github.com/flagos-ai/Torch-FL/pull/440 · https://github.com/flagos-ai/Torch-FL/pull/449 · https://github.com/flagos-ai/Torch-FL/pull/437 · https://github.com/flagos-ai/Torch-FL/pull/448 · https://github.com/flagos-ai/Torch-FL/pull/421 · https://github.com/flagos-ai/Torch-FL/pull/438 · https://github.com/flagos-ai/Torch-FL/pull/436 · https://github.com/flagos-ai/Torch-FL/pull/435 · https://github.com/flagos-ai/Torch-FL/pull/434 · https://github.com/flagos-ai/Torch-FL/pull/410 · https://github.com/flagos-ai/Torch-FL/pull/416 · https://github.com/flagos-ai/Torch-FL/pull/420
- [9] 1.8 FlagFFT — https://github.com/flagos-ai/FlagFFT/commits/dev
- [10] 1.9 FlagSparse / FlagBLAS — https://github.com/flagos-ai/FlagSparse/pull/98 · https://github.com/flagos-ai/FlagSparse/pull/99 · https://github.com/flagos-ai/FlagSparse/pull/100 · https://github.com/flagos-ai/FlagSparse/pull/101 · https://github.com/flagos-ai/FlagBLAS/pull/132 · https://github.com/flagos-ai/FlagBLAS/pull/133 · https://github.com/flagos-ai/FlagBLAS/pull/134 · https://github.com/flagos-ai/FlagBLAS/pull/135
- [11] 1.10 Other Updates — https://github.com/flagos-ai/FlagCX/pull/626 · https://github.com/flagos-ai/FlagCX/pull/627 · https://github.com/flagos-ai/FlagCX/pull/628 · https://github.com/flagos-ai/FlagCX/pull/629 · https://github.com/flagos-ai/FlagQuantum/pull/220 · https://github.com/flagos-ai/FlagQuantum/pull/221 · https://github.com/flagos-ai/FlagScale/pull/1303 · https://github.com/flagos-ai/FlagScale/pull/1304 · https://github.com/flagos-ai/docs/pull/523
- [12] 2.1 Quiet Window — Google News RSS 26 query sets in Chinese and English (via proxy) · HN Algolia https://hn.algolia.com/
- [13] 2.2 OpenCL Operator Library — Changsha Evening News https://www.icswb.com/h/168/20260925/997659.html · VOC (republished by ifeng) https://feng.ifeng.com/c/8whcnBpI2aO
- [14] 2.3 SAIL Second Batch — BAAI Hub https://hub.baai.ac.cn/view/58291 · ifeng https://feng.ifeng.com/c/8weyV6bssOX
- [15] 2.4 Meta-Infer — QbitAI https://www.qbitai.com/2026/09/496925.html · Sina Tech https://finance.sina.com.cn/tech/csj/2026-09-24/doc-inisxnuz4997131.shtml
- [16] 3.6 Tsingmicro Space — Zhejiang Online https://zjnews.zjol.com.cn/zjnews/202609/t20260926_31934786.shtml · ifeng repost https://feng.ifeng.com/c/8wiLgiTy3MZ