Research window: The past 24 hours (2026-10-08 07:00 – 2026-10-09 07:00, Beijing time). First regular daily report after the National Day holiday: the previous regular report was 09-30, and holiday activity was covered by the National Day special combined issue (09-30 – 10-07, with verification extending into the afternoon of 10-08). This window follows seamlessly from that issue; items already recorded in the combined issue are not repeated. Sources: GitHub (push verification across 29 repos in the tile-ai organization; 8 repos had pushes within the window: tilelang, TileOPs, TileOPs-nightly, tilelang-ascend, tilelang-hygon, tilelang-mlir-ascend, tilelang.github.io, TileOPs.github.io; 21 merges and 21 new opens in the organization during the window; 12 merges and 11 new opens in the main repo, 5 merges in the Ascend repo, 4 merges in TileOPs, open PR counts and nightly snapshots parsed individually), new-repo verification for the PTO-ISA organization, Google News RSS multi-language multi-query searches (via proxy, zero hits), Hacker News, arXiv, documentation site and release page verification


Issue Index

  • Today’s Highlights: Post-holiday merge wave — main repo closes the day with 12 merges, Ascend “trilogy” restoration complete (10-08/10-09)
  • I. Core Project Progress
    • 1.1 Main repo late-night batch: Profiler device sync, nested LDG predicate preservation, JIT test cache isolation — three commits landed (10-09)
    • 1.2 Main repo: K-pool Hadamard barrier reduction merged — decode verification consolidated into a single device-to-host sync (10-08)
    • 1.3 Main repo: math codegen dispatch split refactor and AscendC line instruction support merged (10-08)
    • 1.4 Main repo fix batch: zero-fill matching, interleave_weight masking, uint8 sparse GEMM, packed comparison, ROCm rounding (10-08)
    • 1.5 Main repo engineering line: Windows wheel build speedup merged; open PR count at 174 (10-08)
  • II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)
    • 2.1 Ascend: three bot-authored PRs merged in the adaptation repo — mask reuse, sync lifecycle, FP32 Reduce2D (10-08)
    • 2.2 Ascend: CI advances on two fronts — CANN 9.3 self-built image and 950 a5 runner migration (10-08)
    • 2.3 Ascend: MLIR path image released and argmax kernel under review (10-08)
    • 2.4 Hygon: register pipeline and warp-diverge opened — example_gemm benchmarked against rocblas (10-08)
    • 2.5 MetaX / Moore Threads: official repos silent (10-08 verification)
  • III. Ecosystem and Adopters
    • 3.1 PTO-ISA: new quantization teaching repo — three-tier 22-variant ladder across torch/ASC/PTO (10-08)
    • 3.2 TileOPs: FP8 prefill 128-row wave and log_softmax fix merged, four under review (10-08)
    • 3.3 Nightly pipeline: new snapshot with 1,027 correctness items and 1,939 benchmark items, zero failures (10-09)
    • 3.4 Docs and releases: TVM-FFI documentation page updated; TileFoundry under review; no pushes from adopters (10-08)
  • IV. Community, Tutorials, and Events
    • 4.1 Community new repos: H200 source-selection debugger lands; BM1690 pipeline updated in small steps (10-08)
    • 4.2 Media and academia: zero Google News hits, no new arXiv preprints (10-09)
  • V. Trend Observations
    • 5.1 Ascend 950: from “reverted” to “split-commit restoration complete”
    • 5.2 Review backlog vs. test expansion: 174 open PRs against 1,939 nightly benchmarks
    • 5.3 Gaps and risk points
  • Appendix: Materials and verification notes

Today’s Highlights: Post-Holiday Merge Wave — Main Repo Closes Out 12 Merges in a Single Day, Ascend “Trilogy” Restoration Complete (10-08/10-09)

Dates: 2026-10-08, 2026-10-09 Sources: Main repo PR list / FP32 Reduce2D (#1852) / K-pool barrier reduction (#3419)

On the first full working day after the holiday, the tile-ai organization merged and opened 21 PRs each; among them, the main repo tilelang merged 12 in a single day (including three late-night fixes spanning midnight, 11 of which are newly recorded by this newsletter), TileOPs merged 4, and the Ascend adaptation repo merged 5 — the densest merge day since late September. Two main threads deserve attention:

  • Main repo cadence fully restored: Merges span five categories — correctness fixes (Profiler synchronization, nested LDG predicates, JIT cache, quantization masks, sparse GEMM), performance (K-pool barrier reduction), refactoring (math codegen dispatch split), CI (Windows wheel speedup), and Ascend code generation (row instruction support); the latest three landed around midnight on 10-09, forming a continuous “across-midnight” merge band.
  • Ascend “trilogy” restoration complete: The adaptation repo’s three bot-authored PRs — Vector mask reuse (#1850) / synchronized dependency lifecycle (#1851) / FP32 Reduce2D and workspace contract (#1852) — merged in succession during the afternoon; together with the previously recorded C/V attribution (#1848) and buffer domain validation (#1786), all five PRs from the original #1829 “Part 2” split have landed — this 950 feature line, which was reverted wholesale in late September, has completed its restoration; on the same day, the CI side opened two PRs for the CANN 9.3 self-built image and a5 runner migration (see 2.2).
  • Speedup side note: K-pool barrier reduction (#3419) consolidates eleven metadata validation predicates in the decode chain into a single device-to-host synchronization, and covers five levels of Hadamard-128 stages with width-32 warp shuffle — decode head latency has become the new optimization target after operator coverage (see 1.2 for details).

I. Core Project Progress

1.1 Main Repo Late-Night Batch: Profiler Device Sync, Nested LDG Predicate Preservation, JIT Test Cache Isolation — Three Commits Land (10-09)

Date: 2026-10-09 Source: Profiler validation sync (#3457) / Nested LDG predicates (#3390) / JIT cache isolation (#3450)

  • Profiler validation device sync (#3457): Fixes the issue where profiler validation tensors unconditionally sync CUDA (bug #3410) — validation tensors now sync on their own device, no longer triggering a full CUDA sync; the PR was opened at 22:58 that night and merged about an hour later.
  • Nested LDG predicate preservation (#3390): Fixes the issue where predicated LDG ignores outer store predicates (bug #3381) — nested reads now preserve outer predicate constraints.
  • JIT test cache isolation (#3450): Isolates test caches and adds test coverage for the disk cache hit path, eliminating cross-test cache interference.
  • All three landed in succession around midnight on 10-09, joining the previously late-night-merged #3447 to form the window’s tail-end “cross-midnight” merge band.

1.2 Main Repo: K-Pool Hadamard Barrier Reduction Merged — Decode Validation Consolidated into a Single Device-to-Host Sync (10-08)

Date: 2026-10-08 Source: K-pool barrier reduction (#3419)

  • Goal: Reduce latency of the full K-pool decode call. Eleven decode metadata predicates are changed to batch evaluation on fixed-shape tensors, with boolean results returned to the host in one shot (consolidated into a single device-to-host sync); the five-stage Hadamard-128 phases in the compression and decode chain now use width-32 warp shuffle implementations.
  • Engineering boundaries: Sorting and adjacent-value checks are used to detect duplicate tail blocks and cache destinations; padding and pool completion checks maintain their original error priorities; the contracts for cache interfaces, error messages and priorities, BF16 rounding boundaries, FP8 scaling, and ordered tail updates remain unchanged.
  • Context: Follows the already-merged GLM-5.3 independent pipeline (#3254) — for the K-pool path, “sync count and barrier overhead” is now on par with operator-level optimization as a performance lever.

1.3 Main Repo: math codegen Dispatch Split Refactor and AscendC Line Directive Support Merged (10-08)

Date: 2026-10-08 Source: math codegen refactor (#3449) / AscendC line directives (#3447)

  • math codegen dispatch split (#3449): Refactors the dispatch logic for CUDA-side math function code generation — reduces single-point coupling and paves the way for future extensions (merged at 20:30).
  • AscendC line directive support (#3447): Main repo now supports tl.emit_line_directives in AscendC code generation — device source is emitted with line-number directives, enabling source-line alignment for Ascend-side debugging and profiling (merged at 23:48).
  • One infrastructure refactor and one Ascend code generation capability, landing in the evening and late night respectively.

1.4 Main Repo Fix Batch: Zero-Fill Matching, interleave_weight Mask, uint8 Sparse GEMM, Packed Comparison, ROCm Rounding (10-08)

Date: 2026-10-08 Source: -0.0 zero-fill (#3343) / interleave_weight mask (#3140) / uint8 sparse GEMM (#3297) / Packed vector comparison (#3303) / ROCm rounding (#3299)

  • Async copy zero-fill matching (#3343): The zero-fill matcher no longer includes -0.0, eliminating the edge case where negative zero participates in matching (merged at 17:20).
  • interleave_weight mask (#3140): The special branch of interleave_weight in the quantization path now uses a pure integer mask (merged at 15:08).
  • uint8 sparse GEMM (#3297): CUDA sparse GEMM now supports uint8 operands (merged at 15:04).
  • Packed vector comparison (#3303): Emits packed FP16/BF16 vector comparisons (merged at 14:35, co-authored with #3315 of the same day).
  • ROCm rounding (#3299): Fixes round-to-nearest-even in T.round (merged at 13:08) — already covered in the National Day combined issue (first commit after the holiday); recorded here for window completeness.

1.5 Main Repo Engineering Line: Windows wheel Build Speedup Merged; 174 Open PRs (10-08)

Date: 2026-10-08 Source: Windows wheel speedup (#3315) / Main repo PR list / T.copy semantics RFC (#3459)

  • Windows wheel build (#3315): Speeds up Windows wheel builds with a stable uv environment (merged at 17:50) — continues the September build-line remediation streak (ccache hit regression, uv environment migration).
  • Queue temperature: 11 new PRs opened within the window (#3445, #3447, #3449, #3450, #3452 through #3458); main repo open PRs total 174 (116 at the pre-holiday combined issue; a net increase of 58 during the holiday and return-to-work period).
  • Notable among new opens: Two Ascend CI PRs (#3454, #3458, see 2.2), SimdVF multi-buffer fill (#3455), warp specialization stage release (#3452), scalar read index coefficients (#3453), positive integer divisor safety guard flattening (#3445), CodeRabbit review suggestions (#3456); additionally, an RFC clarifying T.copy extent semantics and memory scope roles (#3459) was opened in the early morning.

II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)

2.1 Ascend: Three Bot-Authored PRs Merged in Adaptation Repo — Mask Reuse, Sync Lifetime, FP32 Reduce2D (10-08)

Date: 2026-10-08 Source: Vector mask reuse (#1850) / Sync dependency lifetime (#1851) / FP32 Reduce2D and workspace contract (#1852)

  • Vector mask state (#1850): AscendC Vector mask state is now selected and reused by the compiler — mask switching and restoration move from manual scripts to compile-time management (merged 14:33).
  • Sync dependency lifetime (#1851): Ascend sync dependencies are preserved across access and loop lifetimes, eliminating a class of missing syncs caused by premature release (merged 14:34).
  • FP32 Reduce2D and workspace contract (#1852): Static FP32 row reduction is handed to the C++ compile-time implementation — row count M, elements per row N, and adjacent row start distance S must all be determined at compile time (M>0, N>0, S>=N; multi-row input requires row stride to be a multiple of 32 bytes, single-row is exempt), the compilation chain gains a dedicated Reduce2D pass and a unified workspace contract, and compile-time errors are raised when shape or space requirements are not met (merged 15:35, 42 files).
  • These three are PRs with auto-labeled titles and programmatically generated bot signatures from the pipeline — their merge completes the full re-landing of the split from the original #1829 “Part 2”; this 950 feature line, which was reverted wholesale in late September and whose true cause was identified via a one-line TVM patch, is now fully restored.
  • New and updated batch in the same repo: xAttention example cleanup and Expert v2/v3 flash attention examples (#1863, #1864), cross-core sync configuration validation (#1854), etc.

2.2 Ascend: CI Advances on Two Fronts — CANN 9.3 Self-Built Image and 950 a5 Runner Migration (10-08)

Date: 2026-10-08 Source: CANN 9.3 image validation (#3458) / a5 runner migration (#3454)

  • CANN 9.3 self-built image (#3458): Ascend CI jobs now run validation on a self-built CANN 9.3 image (opened in the early hours of 10-09).
  • a5 runner migration (#3454): Ascend 950 CI migrates to a reconfigured a5 runner.
  • Both are under review — infrastructure migration launched the same day the feature line (#1850/#1851/#1852) merged, keeping 950’s “regression-ready, scalable” buildout in step with feature progress.

2.3 Ascend: MLIR Path Image Release and argmax Kernel Under Review (10-08)

Date: 2026-10-08 Source: 950 image release (#196) / Report and argmax (#193)

  • The tilelang-mlir-ascend repo opened a WIP: adding an Ascend 950 image release for TileLang v0.1.15 (#196); another PR updates the report and adds an argmax kernel (#193) — the MLIR compilation path’s 950 support follows the main repo’s release line; no merges landed on the repo’s default branch within the window, with pushes concentrated on feature branches.

2.4 Hygon: Register Pipeline and warp-diverge Opened — example_gemm Benchmarked Against rocblas (10-08)

Date: 2026-10-08 Source: Register pipeline (#18) / warp size fix (#17)

  • Register pipeline (#18): Adds register pipeline support to the Hygon backend — coordinating with the existing shared memory pipeline, inserting corresponding commit and barrier instructions per pipeline stage; warp-diverge is enabled when warp count exceeds 4. The PR self-reports example_gemm performance reaching levels comparable to rocblas; carried on the gcy_merge_main branch (22 commits, 53 files, +3975/-336, including an upstream main merge).
  • warp size implementation fix (#17): Updates the get warp size related implementation (+6/-4, 2 commits).
  • Both are under review; Hygon is the only domestic backend besides Ascend with code progress in this window.

2.5 MetaX / Moore Threads: Official Repos Quiet (10-08 Check)

Date: 2026-10-08 Source: MetaX repo / Moore Threads repo

  • MetaX tilelang-metax’s most recent push was 09-24, and Moore Threads tilelang-musa’s was 09-30 (the v0.1.15+musa.1 release date); neither saw new activity within the window.

III. Ecosystem and Adoption

3.1 PTO-ISA: A New Quantization Teaching Repo — Three Tiers of torch/ASC/PTO with 22 Variants in a Ladder (10-08)

Date: 2026-10-07, 2026-10-08 Source: tilelang-puzzles-ascend repo / README

  • The PTO-ISA organization has set up a new teaching repo, tilelang-puzzles-ascend (created on the evening of 10-07, with 9 commits densely pushed on 10-08): around four NPU quantization kernels (cast_back, per_token_cast, per_block_cast, per_channel_cast), it provides three isomorphic tiers of implementation — the torch tier (CPU reference, prerequisite knowledge), the asc tier (Ascend SIMD, T.simd), and the pto tier (PTO VMI, T.vmi); each tier has 22 variants, numbered in one-to-one correspondence, directly diffable, with each variant adding only one production configuration.
  • Unified entry point python -m harness.check (including a puzzle mode: blanked-out implementations with hints left in place); all kernels run on device and are cross-checked against the torch tier; a devcontainer is provided (cloning a pinned TileLang version when building the image).
  • It forms a companion set with the main repo v0.1.15’s PTO backend (#3310) and the under-review PTO execution routing (#3316) — the PTO second code path has, for the first time, a systematic hands-on textbook, with teaching brought down to the most commonly used quantization kernel scenarios.

3.2 TileOPs: FP8 Prefill 128-Row Wave and log_softmax Fix Merged, Four Under Review (10-08)

Date: 2026-10-08 Source: FP8 prefill row wave (#2451) / log_softmax fix (#2481)

  • FP8 prefill 128-row wave (#2451, merged at 17:39): 1D2D FP8 prefill tiling is issued as a 128-row wave via a unified WGMMA descriptor — the foundry-prefixed GEMM supply line advances.
  • log_softmax single-tile plan (#2481, merged at 20:36): under the single-tile plan, space is reserved for reduction staging, eliminating contention between staging and data rows.
  • Four under review: linear attention decode (#2487) (DeltaNet / GDN / KDA inference decode decoupled from SM90), GroupNorm wide group (#2486), log_softmax single-tile row (#2485), read-only paged GQA RoPE (#2484).
  • Within the window, TileOPs merged 4 and opened 4, with 6 PRs currently open; among these, the two post-holiday morning PRs (#2482, #2483) have already been covered in the National Day combined issue.

3.3 Nightly Pipeline: New Snapshot with 1027 Correctness Items, 1939 Benchmark Items, Zero Failures (10-09)

Date: 2026-10-09 Source: snapshot commit (ac55e381d6) / snapshot environment metadata

  • One new snapshot within the window (run 37820951969, released in the early hours of 10-09; corresponding to TileOPs merge point 86c1590783, i.e. #2481): 1027 correctness items with zero failures, zero errors, zero skips; 1939 benchmark items (52 suites) with zero failures — benchmark scale remains at the high level reached after the holiday expansion.
  • Runtime environment: H200, CUDA 13.2, torch 2.13.0+cu132; the recent cadence is 1 to 2 per day (the last six: two each on 10-05 and 10-07, one each on 10-06 and 10-08).

3.4 Docs and Releases: TVM-FFI Docs Page Updated; TileFoundry Under Review; No Pushes from Adopters (10-08)

Date: 2026-10-08, 2026-10-09 Source: main docs site commit (0d7563ee) / TileFoundry PR (#213)

  • Docs site: tilelang.github.io saw two consecutive “Update docs” commits (in the early hours of 10-09), the latest updating the API reference page and search index for the TVM-FFI JIT adapter; the TileOPs docs site had 1 automatic gh-pages deployment.
  • TileFoundry: #213 under review — block-scaled FP8 GEMM via SM90 WGMMA going from HIR to TIR — the foundry compilation supply for block-scaled GEMM continues.
  • Adopter check: TileKernels (deepseek-ai) and FlashQLA (QwenLM) had no pushes within the window (their most recent pushes were both on 09-30).
  • Release cadence: the main repo v0.1.15 has been out for a full 9 days with no new version; the latest tilelang-ascend release remains v0.1.2.000-release from 09-09.

IV. Community, Tutorials, and Events

4.1 New Community Repo: H200 Source-Selection Debugger Lands; BM1690 Pipeline Gets a Small Update (10-08)

Date: 2026-10-08 Source: tilelang-debugger / BM1690 pipeline repo

  • Source-selection debugger (tilelang-debugger): created on 10-08 with 4 commits the same day — TileLang source-selection capture targeting NVIDIA H200: establishing the H200 debugging scope and initial design, implementing capture and synchronization validation, adding capability and debug-artifact documentation, and implementing reference analysis while retaining numerical-mismatch samples. This is the first dedicated tool workspace on the community side specifically aimed at TileLang kernel debugging.
  • BM1690 pipeline: the Sophgo BM1690 community pipeline repo had 1 update (a Mul latency test recorded as a manual handoff item), continuing the small-step manual verification cadence.

4.2 Media and Academia: Zero Google News Hits, No New arXiv Preprints (10-09)

Date: 2026-10-09 Source: Google News / arXiv

  • Multiple Google News queries in Chinese and English (TileLang 1-day / 3-day, Ascend / MetaX / Moore Threads combination terms, GPU kernel compilation, etc.) returned zero hits within the window; the 10-02 ~ 10-03 wave of “DeepSeek open-sources Ascend components” coverage predates this window and falls within the combined issue’s scope. Hacker News had no entries on this topic.
  • The latest related arXiv entries remain TileSight from 07-24 and the earlier TileLang paper (2504.17577, v2); there were no new preprints within the window.

V. Trend Observations

5.1 Ascend 950: From “Reverted” to “Split-PR Restoration Complete”

  • In late September, the 950 feature line was once reverted as a whole (the true cause was an out-of-bounds TVM read, identified before the holiday and put back in place with a one-line patch); during the holiday and the return-to-work period, it shifted to a “split-PR” approach, with five items — C/V attribution, buffer domain validation, mask reuse, synchronization lifecycle, and Reduce2D — all merged on the afternoon of 10-08. On the same day, the CI side opened a self-built CANN 9.3 image and a5 runner migration, and the MLIR side opened image publishing — 950’s “features, infrastructure, and compilation path” advancing in sync on three fronts. The restoration is not a return to the starting point, but a reinstallation with CI and debugging capabilities.
  • Cautious note: the 950 backend has been publicly released with v0.1.15 for only 9 days, with no patch release yet; the PTO second path (#3316 execution routing, #3330 DeepGEMM code generation) is still under review. The migration from “adaptation repo” to “main-repo-native + PTO dual path” is still in progress.

5.2 Review Backlog and Test Expansion: 174 Open PRs Against 1939 Nightly Benchmarks

  • Main repo open PRs rose from 116 before the holiday to 174 (+58), with 11 opened and 12 merged within the window — “more opened than merged” continues, and review bandwidth is the key variable for the next version’s cadence; today’s 12 merges can be seen as an initial signal of accelerated review after the return to work.
  • At the other end, the nightly benchmark expanded from 1355 items before the holiday to 1939 and has held at that high level with zero failures and zero skips — test assets and case governance are expanding in sync, providing a pre-merge regression foundation for the large pending review queue.

5.3 Gaps and Risk Points

  • MetaX and Moore Threads official repos are silent; adopters (TileKernels, FlashQLA) had no pushes within the window — ecosystem increments are concentrated within the two organizations tile-ai and PTO-ISA.
  • The four TileOPs PRs under review are concentrated in attention and normalization (linear attention decode, GroupNorm, paged GQA RoPE) — characteristic of a patch-and-adapt period after the holiday-scale sprint.
  • On the media side (Chinese and English), coverage remains at zero; the public narrative for this topic is still driven mainly by GitHub as the source of facts.

Appendix: Sources and Verification Notes

Source Verification Result
tile-ai organization (29 repos) 8 repos had pushes within the window: tilelang (12 merged, 11 newly opened), TileOPs (4 merged, 4 newly opened), TileOPs-nightly (1 snapshot), tilelang-ascend (3 merged, 2 newly opened), tilelang-hygon (2 newly opened), tilelang-mlir-ascend (1 newly opened), tilelang.github.io (2 commits), TileOPs.github.io (1 automated deployment)
Main repo tilelang 12 merges: #3299, #3303, #3297, #3140, #3343, #3315, #3449, #3419, #3447, #3450, #3390, #3457 (#3299 already covered in a previous issue); 11 newly opened (#3445 ~ #3458 range); 174 open PRs; no new releases or tags within the window
Ascend repos (ascend / mlir-ascend) 5 merges (#1848, #1786, #1850, #1851, #1852; the first two already covered in a previous issue); in-review and updated batches include #1854, #1863, #1864, #1795, #1822, #1827, #1828, #1701, #196, #193, etc.; pushes distributed across feature branches
Hygon / MetaX / Moore Threads Hygon: 2 in review (#17, #18) and new gcy_merge_main branch created; MetaX (after 09-24) and Moore Threads (after 09-30) had no activity
TileOPs 4 merges (#2451, #2481 newly recorded; #2482, #2483 already covered in a previous issue); 4 newly opened; 6 open PRs currently
TileOPs-nightly 1 snapshot (ac55e381d6): 1027 correctness items, 1939 benchmark items (52 suites), zero failures, zero errors, zero skips; corresponding merge point 86c1590783 (#2481)
PTO-ISA organization New repo tilelang-puzzles-ascend (9 commits within the window); related repos such as pto-spec and SuperNPUBench also had pushes (not detailed)
Adopters and community TileKernels and FlashQLA had no pushes within the window; community repos tilelang-debugger (new, 4 commits) and BM1690 pipelines (1 commit) had updates
Google News / Hacker News / arXiv Multiple Chinese and English queries returned zero hits; no HN entries on this topic; no new arXiv preprints (latest remains 07-24 TileSight)

Complete Source List