FlagOS Daily Intelligence Report (2026-09-30)
Research window: 2026-09-29 10:00 ~ 2026-09-30 10:00 (24 hours; Wednesday window; continues from the 09-29 report window with no gap) Sources: GitHub (org: flagos-ai, full check of pushed_at across 54 repos; 16 repos active in-window, 12 repos with in-window commits; full item-by-item verification of 168 commit-search results; also checked tags and Release atoms for Torch-FL / FlagScale / FlagGems-vllm / FlagGems-sglang / FlagTree, community #124 ~ #125, ~30 merge records in build-infra, the docs documentation batch, 7 deployments of the release-info portal, and branch-level activity in FlagFFT), Google News RSS (24 sets of Chinese and English query terms, via proxy), reports from ifeng / Sina Finance, etc. (see the appendix source checklist for details)
Index for This Issue
- Today’s Highlights: Torch-FL v2.10.0 release lands in the repo—the last item on the FlagOS 2.2 checklist is completed (09-29)
- I. Open Source Project Progress (GitHub Activity)
- 1.1 Torch-FL: v2.10.0 officially released—the last item of 2.2 completed (09-29)
- 1.2 Torch-FL (II): Release pipeline iteration—from candidate tags to vendor wheels (09-29 ~ 09-30)
- 1.3 community: Two wrap-up correction commits and the eight-repo cutoff policy (09-29)
- 1.4 Release artifact refresh: FlagScale tag rebuild and Releases for the two FlagGems-family repos (09-29)
- 1.5 build-infra: Wrap-up of three-line image records—mthreads / ascend / enflame (09-29 ~ 09-30)
- 1.6 FlagGems: KernelGen operator batch and Ascend performance batch (09-29 ~ 09-30)
- 1.7 FlagGems-vllm / FlagGems-Experimental: KMCompiler kernels and the 5.5 alpha line (09-29 ~ 09-30)
- 1.8 FlagTree: metax post-patch tags and the all-backend scheduling workflow (09-29)
- 1.9 FlagQuantum: Kernel directory routing and performance fusion batch (09-29 ~ 09-30)
- 1.10 Other activity: FlagCX / FlagFFT / FlagAttention / docs / vllm-plugin-FL (09-29 ~ 09-30)
- II. News Coverage and Ecosystem
- 2.1 Component-level search: the twenty-third consecutive quiet window (09-29 ~ 09-30)
- 2.2 Moore Threads: Wuxi (Huishan) domestic intelligent computing center in operation for three months—six major agents deployed (09-29)
- III. Deep Dive on Member Organizations
- 3.1 Moore Threads: musa image batch, MUSA 3D optimization branch, and Wuxi industry (09-29)
- 3.2 Ascend: megatron 0.18.2 four variants and the FlagGems performance batch (09-29)
- 3.3 Hygon / MetaX: Performance operator batch and metax post-patch (09-29)
- 3.4 Enflame / Iluvatar: Image batch and iluvatar fused kernels (09-29 ~ 09-30)
- 3.5 Kunlunxin / Cambricon / DAMO XuanTie / Horizon Robotics: Other key points (09-29 ~ 09-30)
- IV. Summary and Trend Observations
- Appendix: Source Verification Table
- Complete Source List
Today’s Highlight: Torch-FL v2.10.0 Released and Landed — Final Item on the FlagOS 2.2 Checklist Completed (09-29)
Date: 2026-09-29 Source: Torch-FL v2.10.0 Release, community #124, community #125
Torch-FL — the “sole remaining blank” on the 2.2 checklist (24 items) noted in the 09-28 daily — landed in this window. Two candidate tags (v2.10.0rc1 / rc2) at 12:34 / 17:01 first rehearsed the release pipeline, followed by Release creation at 18:17 and official release at 18:20:
- Release contents: torch-fl is FlagOS’s PyTorch device plugin — it unifies per-operator routing (native vendor kernels / FlagGems portable kernels / CUDA-compatible boxing / CPU fallback) under a single
flagosdevice, letting the same PyTorch program run unchanged across nine classes of accelerator platforms: NVIDIA (A100 / H800), Ascend (910C), Hygon (BW1000), MetaX (C550), DAMO XuanTie (ZW810E), Moore Threads (S5000), Enflame (S60), Kunlunxin (P800), and Horizon-series BPU (s600). - Versioning and packaging: Starting with this version, the version number follows the bound PyTorch minor version line (this time
>=2.10,<2.11); wheels are namedtorch_fl-2.10.0+<sdk>(e.g. +cuda13.3 / +dtk2604 / +cann9.0.0), “one tag, one artifact per platform”; wheels are self-describing (compatibility.json records platform, kernel set, build ABI, and pinned versions of FlagTree / FlagGems / FlagCX), and ship with atorch-fl-preflightvalidation CLI (checks the environment without importing torch_fl). - Validation scope: DDP and FSDP2 pass on all mainstream cards; torch.compile registers flagos as a first-class inductor GPU device (FlagTree backend covers six vendors, Enflame uses vendor Triton, BPU uses the onboard compiler); profiler aligns with CUPTI (including vendor-specific tracers); in cross-platform benchmarks (Qwen-Image-2.1, single-card 1024², 40 steps), H100 native achieves 6.52 s/image versus FlagOS H100 10.67, Moore Threads S5000 14.87, Hygon BW1000 23.59, Ascend 910C 24.94, MetaX C550 31.61, DAMO XuanTie ZW810E 37.44, Enflame S60 58.04 (image quality aligned with H100, CLIP-Score deviation within ±1.6%); the pure CPU Arm path (Mac M5 Pro) achieves a 1.74× speedup over PyTorch CPU.
- Three new capabilities: a skill for onboarding new chip backends (T-Head ZW810E and Kunlunxin P800 were introduced via this path in this cycle), FP8 / FP4 software emulation on the boxing path (for when no FP8 hardware is available), and operator-level automatic routing tuning.
- Boundary notes (included with the release): torch.compile backends remain experimental; Kunlunxin P800’s profiler / RNG and compile are still on the roadmap; flex_attention is unsupported in this version; BPU s600 serves only as a torch.compile graph device (eager falls back to CPU).
- Volume: 269 cumulative commits since v0.1.0 (62 features / 101 fixes / 11 performance / 25 CI / 18 docs / 16 tests / 10 refactors / 4 builds); 100+ HuggingFace models validated on boards across the matrix.
Assessment: With this, all 24 items on the official 2.2 checklist now have corresponding released artifacts landed (the review snapshot for the Torch-FL entry on the checklist remains pending, as expected, to be refreshed in the next correction batch); the four-artifact chain of “tag → Release → wheel → image” closed within this window.
I. Open-Source Project Progress (GitHub Activity)
1.1 Torch-FL: v2.10.0 Officially Released — The Last Item of 2.2 Completed (09-29)
Date: 2026-09-29 Source: Torch-FL v2.10.0, Torch-FL tags list, compare v0.1.0…main
- Timeline: rc1 tag at 12:34 (landing on the “make candidate tags releasable” commit) → rc2 tag at 17:01 → Release object created at 18:17, published at 18:20; the official tag v2.10.0 points to the same commit as rc2 (
#481vendor runner release line). - Platform details (release artifact matrix): Ascend 910C with 354 routed operators, FlagGems routing 221 → 225, Qwen3-0.6B training measured at 0.82× torch_npu; Hygon BW1000 with 2033 operators routed via CUDA boxing, FlagGems routing 388 → 470; MetaX C550 cu-bridge with 2033 operators, routing 377 → 443; DAMO XuanTie ZW810E routing 390 → 435; Moore Threads S5000 with 105 native mudnn kernels, routing 391 → 468; Enflame S60 with 224 native libtopsaten kernels, Qwen-Image-2512 optimized to 7.1 seconds/step; Kunlunxin P800 with 2022 operators, 380 FlagGems-Python routes; on the NVIDIA side, the FlagGems-C++ dispatch path is in place (18 routes).
- Distributed: collectives run per platform via FlagCX+MCCL (MUSA), FlagCX (GCU), RCCL on DTK, HCCL fallback and NCCL-shaped routing; gloo requests are answered on the flagos backend.
- Implication: The “overnight release window” approach — rc candidate line first, official tag the next day — matches how the 2.2 major release was organized: validate the artifact pipeline first, then land the official object.
1.2 Torch-FL (Part 2): Release Pipeline Iteration — From Candidate Tags to Vendor Wheels (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: #473, #474, #475, #476, #480, #483
- Packaging spec: #473 wheels are named after the SDK they were built with, with a fixed interpreter per platform (3.12 for CUDA / GCU / MetaX / PPU, 3.10 for DCU / MUSA, 3.11 for Ascend).
- Tag-driven: #474 builds and publishes wheels from tags; #475 makes release candidate tags releasable (rc1 / rc2 went through via this mechanism).
- Upload reliability: #476 concurrent uploads + timeouts and retries; #480 / #481 build and publish each platform’s wheel via that vendor’s own runner to the vendor PyPI lane.
- Ordering fix: #483 / #484 (09-30 00:55) corrected to “publish first, then upload artifact” — the final pipeline patch of the window.
- Implication: Release engineering has advanced from “tags usable” to “per-platform artifacts reproducible and verifiable” — nine days ago this was still a pending item, and now even the release pipeline itself has gone from 0 to 1.
1.3 community: Two Final Fixes and the Eight-Library Cutoff Policy (09-29)
Date: 2026-09-29 Source: community #124, #125, #114, #115
- #124 (11:54) cutoff alignment: FlagGems-vllm and FlagGems-sglang switched to “the last commit on main before 2026-09-28 09:00” — the two libraries’ previous official tags were deleted and rebuilt on the verified cutoff commit (version names unchanged); the cutoff policy thus extends to all eight operator libraries (the original six plus these two).
- #125 (12:28) FlagScale fix included: FlagScale’s RC2 initialization fix (including #1304, a non-heterogeneous initialization issue) is included in 2.2 — the manifest’s FlagScale entry is changed to
release-source: rc2-head, target commit 4371924. - Status notes: #114 (adding L1 operator libraries to the manifest) was closed unmerged, its scope restructured by the #116 / #117 / #123 line; #115 (FEP scope delivered in 2.2 and items deferred to 2.3) remains open with no action during the window.
1.4 Release Artifact Refresh: FlagScale Tag Rebuild and Two FlagGems-Family Releases (09-29)
Date: 2026-09-29 Source: FlagScale v2.1.0, FlagGems-vllm v0.2.0, FlagGems-sglang v0.1.0
- FlagScale v2.1.0 (12:27): The tag was rebuilt on “4371924f on 2.1.0-rc2” (including the #1304 fix), and the Release object was refreshed in sync — the scope is updated to 33 PRs plus 2 additional commits (the earlier count of 32 PRs is corrected accordingly).
- FlagGems-vllm v0.2.0 (11:52): The tag was rebuilt on cutoff commit
b5c2c224, and the Release landed at 11:52 ~ 11:55 (body notes Part of FlagOS 2.2). - FlagGems-sglang v0.1.0 (11:53): The tag was rebuilt on cutoff commit
575d5bc3, and the Release landed in sync (first version: 80 PRs and 14 additional commits). - Window linkage: The three libraries’ actions clustered between 11:52 and 12:28, linking at minute-level granularity with #124 / #125 — “fix records → rebuild tags → refresh Releases” closed the loop within a single midday window.
1.5 build-infra: Three Image-Record Lines Wrapped Up — mthreads / ascend / enflame (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: build-infra image record batch, #1292, #1296 ~ #1303, #1312, #1314
- Ascend four variants (#1296 ~ #1303, 16:28 ~ 17:04): Four records — cann8.5.0 / cann9.0.0 / cann8.5.0-910c / cann9.0.0-910c — added to the megatron 0.18.2 training matrix, plus the 0.18.2 changelog; the verify container’s proxy passthrough was also fixed (#1297).
- Moore Threads batch (#1294 ~ #1313, 16:19 ~ 00:19): Application tags for vllm (0.3.0 / 0.2.2), megatron (0.2.3) and sglang (0.1.dev1) recorded in bulk across the musa4.3.6 / musa5.2.0 lines; mthreads apps authorized to rebuild on the new flagtree wheel (#1292); sglang image tags restored to the 2.2.0 stack prefix; mthreads sglang docs patched in the morning (#1314, 08:50).
- Enflame batch (#1284 ~ #1291, 10:39 ~ 10:41): vllm (0.2.2), megatron (0.2.3) and sglang records for the tops1.9.10 / 1.10.6 lines landed in bulk.
- Build infrastructure: EXTRA_PYPI fallback switched to the Tencent PyPI mirror (#1312); Nexus large-file upload health probe workflow (#1293) opened for review.
- Volume: 51 commits and roughly 30 merged PRs in build-infra during the window; the release-info portal auto-deployed 7 times alongside the record batches.
1.6 FlagGems: KernelGen Operator Batch and Ascend Performance Batch (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: FlagGems #6610, #6770, #6699, #6553, #6633, #6795, #6741, #6799
- KernelGen operator batch (16 items): New on the NVIDIA side — dim, kl_div, linalg_pinv, weight_int4pack_mm, _spsolve, triplet_margin_loss, max_pool1d_with_indices, conv_tbc, quantized_batch_norm, quantize_per_channel, put, special_hermite_polynomial_he, histogram, fused SDPA backward, batch_norm_stats, bartlett_window — KernelGen-generated operators continue to enter the main library densely.
- Ascend performance batch (#6697 ~ #6703, #6770): baddbmm bias broadcast and tile selection, bmm tile and masking, addmm for large BF16 workloads, mv GEMV dispatch, mul broadcast dispatch continuously tuned; also fixed a global leak of the TRITON_ALL_BLOCKS_PARALLEL environment variable.
- Vendor optimizations: batch optimization of Hygon performance-weak operators (#6553); KMCompiler linear covering four backends — NVIDIA / Hygon / MetaX / Thead (#6633); SiliconFlow line’s geometric family and argsort implemented per backend (#6575 / #6692).
- CI and fixes: Cambricon ops-test image update (#6795), Moore Threads tests migrated into flagcicd, changed operators routed to marked tests (#6741), runner volume cleanup; tle import guard (#6777), SQL cold-start WAL race (#6730), FlashAttention non-contiguous tensor materialization (#6727); benchmark baselines switched to native aten operators (#6195).
- New submissions: Five sparse / metadata test suites (#6799) submitted for review on the morning of 09-30.
1.7 FlagGems-vllm / FlagGems-Experimental: KMCompiler Kernels and the 5.5 Alpha Line (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: FlagGems-vllm #873, #875, #876, #877, FlagGems-Experimental #721
- FlagGems-vllm: #873 adds backend-specialized dequantize_and_gather_k_cache (Ascend / Moore Threads); #875 / #876 optimize Iluvatar’s fused q_kv_rmsnorm and fused add rms norm; #877 bumps the test image to flagtree v0.7.0.
- FlagGems-Experimental: The same batch of KernelGen operators (slow_conv_transpose3d, combinations, sparse_resize_, sparse_coo_tensor, _unpack_dual, can_cast, _fw_primal, flatten_dense_tensors, ccol_indices, etc.) continues to enter the experimental library; the alpha version number advances from 5.4 to 5.5 (#721) — the development track for the next operator-library version line has started.
1.8 FlagTree: metax Post-Patch Tag and All-Backend Scheduling Workflow (09-29)
Date: 2026-09-29 Source: FlagTree tags, #1310, #1307, #1212, #1303
- Post-patch tag:
0.7.0.post2+metax3.6was pushed by flagtree-bot at 18:27 (“pypi 0.7.0.post2+metax3.6”) — the target commit is “Metax bf16 sum reduction promoted to fp32” (#1310), making metax the first backend to receive a post2 fix variant. - Commits: #1307 adds a scheduled workflow for all backends; #1212 adds decode-phase TLE-fused AllReduce + RMSNorm; #1303 has CI validate the pid file path before killing processes.
- Implication: Beyond the 2.2 freeze line, a channel for “incremental post-release patching by variant” is now open — variant granularity (+metax3.6, etc.) supports an independent patch cadence.
1.9 FlagQuantum: Kernel Catalog Routing and Performance Fusion Batch (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: FlagQuantum #240, #236, #241, #259
- Kernel catalog series (#240 ~ #259, 25 commits): After the semantic catalog (#240), evidence manifest (#243) and capability matcher (#244), paths including local 1q / CNOT / CNOT sequences / fused transpose / controlled-subspace transfer / adjoint VJP / invertible VJP / sharded adjoint VJP were progressively wired into catalog routing — the scheduling layer is moving from hardcoded to registry-driven.
- Performance fusion: adjoint boundary fusion (Euler triples, QAOA forward / backward boundaries, observable rotation boundaries), CX gathers fusion, rotation tile widening, final-state recovery skipping, forward-invariant caching — the CPU-native path is being consolidated continuously.
- New features: continuous-time Lindblad evolution (#236), provider-state VQE ground-state solving (#241), Triton compiler provenance parsing (#233).
- Implication: The engineering foundation for the quantum-intelligence fusion direction (a dedicated session at the 10-17 conference) — components are entering a high-frequency iteration phase three weeks before the conference.
1.10 Other Activity: FlagCX / FlagFFT / FlagAttention / docs / vllm-plugin-FL (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: FlagCX #632, FlagFFT activity, FlagAttention #71, docs #521, docs #530, vllm-plugin-FL #563
- FlagCX: #632 integrates shared transport ordering (CRL) and retryable teardown — the cross-backend transport consistency foundation continues to be hardened.
- FlagFFT: Two new optimization branches added —
codex/musa-3d-opt(Moore Threads) andcodex/maca-3d-opt(MetaX) — with multiple pushes from the afternoon into the evening of 09-29; multi-platform optimization lines for 3D FFT are now running in parallel. - FlagAttention: Three PRs — moba attention (#71), Hy3 attention (#56), parallax (#57) — merged consecutively on the morning of 09-30.
- docs batch (23 commits): #521 “2.2 documentation draft” merged as a whole batch (09-30 08:44) — full Chinese and English documentation newly created for FlagAttention, release notes updated for FlagAudio / FlagBLAS / FlagDNN / FlagFFT / FlagGems-vllm, FlagQuantum docs overhauled (architecture / capabilities / algorithms / user guide), deployment workflows expanded; #530 creates Chinese and English documentation pages for FlagGems-sglang (installation / requirements / operator list / release notes); #514 chip adaptation guide; the model list batch continues (ModelScope series, aihuanxin, HuggingFace open PRs).
- vllm-plugin-FL: #563 merged after review; the open PR line is active — MiniCPM-V 4.7 adaptation (#567), Sunrise PTPU static graph (#568), Iluvatar CoreX BI-V150 adaptation to vLLM 0.28.0 (#566), Kunlunxin P800 upgraded to vLLM 0.24.0 (#561), Ascend generic full-decode graph (#409), plus three PRs for H100 support of Qwen3.8-Flash / GLM5.3-Flash / HY4 (#455 / #454 / #453) — still under intensive review on the morning of 09-30.
II. News Coverage and Ecosystem
2.1 Component-Level Search: 23rd Consecutive Quiet Window (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: Google News RSS (24 Chinese and English query terms, via proxy), HN Algolia
- Component and release-line queries (FlagOS / FlagGems / FlagScale / FlagTree / FlagCX / KernelGen / FlagPerf / conference agenda) produced no new public announcements within the 24-hour window — the 23rd consecutive quiet window; Torch-FL’s release actions all occurred on the code and artifact side (see Chapters I and III).
- 11 ecosystem-side hits in the window: one directly relevant item on the Moore Threads industry side (see 2.2); the rest were stock-market pieces (Enflame profitability expectations, GPU “four little dragons” lockup expirations), BAAI-related social news (an interview on AI safety investment) and irrelevant reprints, filtered out by topic.
- HN Algolia returned no relevant hits for FlagOS / FlagGems / FlagScale / BAAI (fuzzy matches were all unrelated topics).
2.2 Moore Threads: Wuxi (Huishan) Domestic Intelligent Computing Center Three Months Into Operation — Six Major Agents Deployed (09-29)
Date: 2026-09-29 Source: ifeng (reprinted with the same headline by Sina Finance and others)
- Milestone: Phase I of the Wuxi (Huishan) domestic intelligent computing center has been in operation for three full months — it approached full capacity just two months after launch, becoming a benchmark for commercial deployment of domestic compute; all 2000 Moore Threads MTT S5000 intelligent computing cards in Phase I are now in operation (part of the Huishan node of the “East Data, West Computing” Yangtze River Delta hub).
- Six major agents: government affairs (the Huishan investment-promotion agent, integrating data-compute-application), healthcare (Huishan District People’s Hospital × Nanjing University of Posts and Telecommunications Edge Intelligence Research Institute smart hospital system, with nearly thirty models deployed; Zhizhen Technology’s WiseAgent health agent), education (East China Normal University’s Qichuang InnoAgent platform, integrating 300+ education agents and 1000+ education skills, now open-sourced), chip design (X-Epic X-IVS / X-IDS, with a Formal Agent case completing end-to-end in 23.8 hours) and other high-value tracks.
- Relation to this stack: MTT S5000 is a FlagOS member platform — the Moore Threads-side musa application image batch in this window, the S5000 data in Torch-FL’s release artifacts (105 native kernels, FlagGems routing 391 → 468), and the “domestic chips training domestic models” industry narrative mirror each other; the “Huishan model” is next to be replicated in more regions.
III. Member Unit Deep Dive
3.1 Moore Threads: musa Image Batch, MUSA 3D Optimization Branch, and Wuxi Industry (09-29)
Date: 2026-09-29 ~ 2026-09-30 Source: build-infra #1294 ~ #1313, FlagGems-vllm #873, FlagFFT activity
- Application images: musa4.3.6 / musa5.2.0 dual-line vllm (0.3.0 / 0.2.2), megatron (0.2.3), sglang tags batch-committed; applications granted rebuild authorization on the new flagtree wheel (#1292).
- Operator layer: FlagGems-vllm dequantize_and_gather_k_cache mthreads specialization (#873); FlagTree bf16 sum reduction fp32 promotion post2 patch is based on this backend scenario (#1310, see 1.8).
- Optimization branch: FlagFFT
codex/musa-3d-optbranch pushed multiple rounds on 09-29 afternoon (13:53 ~ 14:30). - Industry side: Wuxi (Huishan) intelligent computing center operational for three months, six major agents deployed (see 2.2).
3.2 Ascend: megatron 0.18.2 Four Variants and FlagGems Performance Batch (09-29)
Date: 2026-09-29 Source: build-infra #1296 ~ #1303, FlagGems #6699, #6770, FlagGems-vllm #873
- Training framework: megatron 0.18.2 four backend lines (cann8.5.0 / cann9.0.0 / 910c dual variants) records and changelogs committed — Ascend training side completes the 0.18.2 matrix on the first workday after release week.
- Operator performance: FlagGems six Ascend optimizations (baddbmm / bmm / addmm / mv / mul / environment variable leak fix).
- Inference kernels: FlagGems-vllm KMCompiler dequantize_and_gather_k_cache Ascend specialization (#873).
- Open PRs: vllm-plugin-FL Ascend general full-decode graph (#409) under continued review.
3.3 Hygon / MetaX: Performance Operator Batch and metax Post Patches (09-29)
Date: 2026-09-29 Source: FlagGems #6553, #6633, FlagTree #1310, FlagFFT activity
- Hygon: FlagGems batch optimization of performance-deficient operators (#6553); KMCompiler linear Hygon backend folded into four-backend unified refactor (#6633); in Torch-FL release artifacts, BW1000 uses CUDA boxing (2033 operators, 466 FlagGems routes) and runs on upstream PyTorch rather than the DTK branch.
- MetaX: FlagTree
0.7.0.post2+metax3.6post patch tag (bf16 sum → fp32); FlagFFTcodex/maca-3d-optbranch pushed continuously from 09-29 afternoon to evening (15:31 ~ 18:45); in Torch-FL release artifacts, C550 routes 2033 operators via cu-bridge, FSDP2 and Qwen3 training aligned.
3.4 Enflame / Iluvatar: Image Batch and iluvatar Fused Kernels (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: build-infra #1284 ~ #1291, FlagGems-vllm #875, #876
- Enflame: tops1.9.10 / tops1.10.6 dual-line vllm (0.2.2), megatron, sglang application records batch-committed (#1284 ~ #1291); in Torch-FL release artifacts, S60 with native libtopsaten backend (224 kernels) optimizes Qwen-Image-2512 to 7.1 seconds/step.
- Iluvatar: FlagGems-vllm two fused kernels — fused q_kv_rmsnorm and fused add rms norm optimizations (#875 / #876).
3.5 Kunlunxin / Cambricon / DAMO XuanTie / Horizon: Other Highlights (09-29 ~ 09-30)
Date: 2026-09-29 ~ 2026-09-30 Source: vllm-plugin-FL #561, FlagGems #6795, #6633, Torch-FL v2.10.0
- Kunlunxin: vllm-plugin-FL PR upgrading P800 to vLLM 0.24.0 (#561) pending review; in Torch-FL release artifacts, P800 profiler / RNG and compile listed as roadmap items.
- Cambricon: FlagGems ops-test image update (#6795).
- DAMO XuanTie: KMCompiler linear Thead backend folded into four-backend refactor (#6633); in Torch-FL release artifacts, ZW810E routes 390 → 435, AMP integration.
- Horizon: BPU s600 enters Torch-FL platform matrix as edge target (torch.compile graph path, onboard compiler).
- Qingwei: no new activity in this window (FlagScale txda backend already covered in previous window).
IV. Summary and Trend Observations
- 2.2 closure complete: all 24 checklist items committed — Torch-FL v2.10.0 released within the window, completing the final item; the “fix → rebuild → re-release” discipline triple (#123 six-repo source fix → #124 two-repo cutoff alignment → #125 FlagScale fix inclusion) closed the loop within 12 hours.
- Release engineering becomes a visible capability — Torch-FL went from rc1 → rc2 → official release within one day, with multiple iterations on the release pipeline (tag-driven builds, vendor-owned runners, concurrent upload retries, release order fixes); wheel self-description + preflight validation — release artifacts moving from “installable” to “verifiable, traceable.”
- Development lines fully transition to the next cycle — FlagGems-Experimental alpha raised to 5.5, FlagGems-vllm 0.3.0-dev1 (previous window), vllm-plugin-FL open PRs for new models (Qwen3.8-Flash / GLM5.3-Flash / MiniCPM-V 4.7) and new hardware (CoreX BI-V150, P800 upgrade) under intensive review — after release freeze ends, the adaptation line re-accelerates.
- Three main threads in sync within the window — KernelGen operator coverage (16 merged into main repo), vendor platform optimizations (Ascend / Hygon / MetaX / Enflame each with dedicated batches, FlagFFT dual optimization branches), quantum system stack (FlagQuantum 25 items) — aligned with the AI systems / quantum-intelligence fusion track themes of the 10-17 conference.
- The 23rd quiet window on the news side coexists with industry-side moves — the pattern of code-driven development and conference-node narratives continues; the Moore Threads Wuxi case (six agents + full-load benchmark) is a new sample of the “domestic compute commercialization” narrative, corroborating the S5000 data in Torch-FL release artifacts.
- Next observation points: community checklist snapshot refresh for Torch-FL entries, 2.3 development line convergence (FlagGems 5.5, FlagGems-vllm 0.3.0, vLLM 0.28 adaptation), and external releases at the 10-17 conference.
Appendix: Source Verification Table
| Category | Source | Verification Method | Result |
|---|---|---|---|
| GitHub | org: flagos-ai repos API | Full verification of pushed_at across 54 repos | 16 repos active within window (12 of which have commits within window) |
| GitHub | Commit search + per-repo cross-check | Line-by-line verification of 168 commits from commit search (full total_count) | 168 commits across 12 repos (build-infra 51 / FlagGems 38 / FlagQuantum 25); no new repos within org |
| Releases | Torch-FL releases / tags API | Line-by-line verification of tags, Releases, and release timestamps | v2.10.0 (created 18:17 / published 18:20); rc1 / rc2 candidate tags preceded |
| Releases | FlagScale / FlagGems-vllm / FlagGems-sglang / FlagTree atoms | Line-by-line comparison of tag objects (tagger time and target commit) | Three repos’ tags rebuilt (11:52 ~ 12:27); metax post2 tag (18:27) |
| Release engineering | community #124 ~ #125 original text | Verification of titles, bodies, and file diffs | Both merged; cutoff policy extended to eight repos |
| Release engineering | release-2.2.yaml manifest | Original text retrieval and verification | Eight-repo cutoff + FlagScale rc2-head; Torch-FL snapshot pending refresh |
| Release engineering | build-infra PR batch and release-info portal | Line-by-line comparison | ~30 merged; portal deployed 7 times |
| Docs | docs PR #521 / #530 file lists | Per-file verification | Full 2.2 docs batch (including full FlagAttention set) and sglang docs page |
| News | Google News RSS (24 sets of Chinese and English query terms, via proxy) | 24-hour window filtering + line-by-line exclusion | Zero additions on component line (23rd quiet window); 1 item retained on ecosystem side |
| Media | ifeng.com (same-topic reprints on Sina Finance etc.) | Original text retrieval and review (publish time 09-29 18:44) | Wuxi (Huishan) Intelligent Computing Center three-month milestone |
| Community | HN Algolia | Search review | No relevant hits |
| Historical comparison | Previous two days’ FlagOS daily reports | Deduplication verification | Torch-FL pending → release as core window increment; Wuxi case included for the first time; AICC2026 already included on 09-23 / 09-24 |
Complete Source List
- [1] Today’s Highlight: Torch-FL release landed — https://github.com/flagos-ai/Torch-FL/releases/tag/v2.10.0 · https://github.com/flagos-ai/community/pull/124 · https://github.com/flagos-ai/community/pull/125
- [2] 1.1 Torch-FL release — https://github.com/flagos-ai/Torch-FL/releases/tag/v2.10.0 · https://github.com/flagos-ai/Torch-FL/tags · https://github.com/flagos-ai/PyTorch-Plugin-FL/compare/v0.1.0...main
- [3] 1.2 Release pipeline — https://github.com/flagos-ai/Torch-FL/pull/473 · https://github.com/flagos-ai/Torch-FL/pull/474 · https://github.com/flagos-ai/Torch-FL/pull/475 · https://github.com/flagos-ai/Torch-FL/pull/476 · https://github.com/flagos-ai/Torch-FL/pull/480 · https://github.com/flagos-ai/Torch-FL/pull/483
- [4] 1.3 community — https://github.com/flagos-ai/community/pull/124 · https://github.com/flagos-ai/community/pull/125 · https://github.com/flagos-ai/community/pull/114 · https://github.com/flagos-ai/community/pull/115
- [5] 1.4 Release artifact refresh — https://github.com/flagos-ai/FlagScale/releases/tag/v2.1.0 · https://github.com/flagos-ai/FlagGems-vllm/releases/tag/v0.2.0 · https://github.com/flagos-ai/FlagGems-sglang/releases/tag/v0.1.0
- [6] 1.5 build-infra — https://github.com/flagos-ai/build-infra/pulls?q=is%3Apr+app-image-tag · https://github.com/flagos-ai/build-infra/pull/1292 · https://github.com/flagos-ai/build-infra/pull/1297 · https://github.com/flagos-ai/build-infra/pull/1312 · https://github.com/flagos-ai/build-infra/pull/1314
- [7] 1.6 FlagGems — https://github.com/flagos-ai/FlagGems/pull/6610 · https://github.com/flagos-ai/FlagGems/pull/6770 · https://github.com/flagos-ai/FlagGems/pull/6699 · https://github.com/flagos-ai/FlagGems/pull/6553 · https://github.com/flagos-ai/FlagGems/pull/6633 · https://github.com/flagos-ai/FlagGems/pull/6795 · https://github.com/flagos-ai/FlagGems/pull/6741 · https://github.com/flagos-ai/FlagGems/pull/6799
- [8] 1.7 FlagGems-vllm / Experimental — https://github.com/flagos-ai/FlagGems-vllm/pull/873 · https://github.com/flagos-ai/FlagGems-vllm/pull/875 · https://github.com/flagos-ai/FlagGems-vllm/pull/876 · https://github.com/flagos-ai/FlagGems-vllm/pull/877 · https://github.com/flagos-ai/FlagGems-Experimental/pull/721
- [9] 1.8 FlagTree — https://github.com/flagos-ai/FlagTree/tags · https://github.com/flagos-ai/FlagTree/pull/1310 · https://github.com/flagos-ai/FlagTree/pull/1307 · https://github.com/flagos-ai/FlagTree/pull/1212 · https://github.com/flagos-ai/FlagTree/pull/1303
- [10] 1.9 FlagQuantum — https://github.com/flagos-ai/FlagQuantum/pull/240 · https://github.com/flagos-ai/FlagQuantum/pull/236 · https://github.com/flagos-ai/FlagQuantum/pull/241 · https://github.com/flagos-ai/FlagQuantum/pull/259
- [11] 1.10 Other updates — https://github.com/flagos-ai/FlagCX/pull/632 · https://github.com/flagos-ai/FlagFFT · https://github.com/flagos-ai/FlagAttention/pull/71 · https://github.com/flagos-ai/docs/pull/521 · https://github.com/flagos-ai/docs/pull/530 · https://github.com/flagos-ai/vllm-plugin-FL/pull/563
- [12] 2.1 Quiet window — Google News RSS 24 sets of Chinese and English query terms (via proxy) · https://hn.algolia.com/
- [13] 2.2 Wuxi Intelligent Computing Center — ifeng.com https://feng.ifeng.com/c/8worZesI4HF
- [14] 3.x Member unit mirrors and records — https://github.com/flagos-ai/build-infra/pulls?q=is%3Apr+mthreads · https://github.com/flagos-ai/build-infra/pulls?q=is%3Apr+megatron · https://github.com/flagos-ai/build-infra/pulls?q=is%3Apr+enflame · https://flagos-ai.github.io/release-info/
- [15] 3.5 Other key points — https://github.com/flagos-ai/vllm-plugin-FL/pull/561 · https://github.com/flagos-ai/Torch-FL/releases/tag/v2.10.0