TileLang Daily Intelligence Report (2026-10-10)
Research window: Past 24 hours (2026-10-09 07:00 ~ 2026-10-10 07:00, Beijing time). This issue is published as usual on a National Day holiday make-up workday (Saturday); the window seamlessly connects with the previous issue (10-09). Sources: GitHub (push verification across 29 repositories in the tile-ai organization; 9 repositories had pushes within the window: tilelang, TileOPs, tilelang-ascend, tilelang-hygon, tilelang-mlir-ascend, TileFoundry, tilelang.github.io, TileOPs.github.io, TileOPs-nightly; 21 merges and 29 new opens within the organization window; 6 merges and 10 new opens in the main repo, 9 merges in TileOPs, 3 merges across the two Ascend repos; open PRs: 177 in the main repo), PTO-ISA organization verification, multiple Chinese and English Google News RSS queries, Hacker News, arXiv, documentation site and release page verification
This Issue’s Index
- Today’s Highlights: Friday merge wave — the organization closed out the day with 21 merges, with the TileOPs refactor batch, the Ascend examples line, and TileFoundry FP8 block scaling all landing on the same day (10-09)
- I. Core Project Progress
- 1.1 Main repo fix batch: T.view scalar bit width, Autotuner input reuse, autodd hang, LDGSTG test gating — four merges (10-09)
- 1.2 Main repo: CodeRabbit intent-routing review merged — review governance enters the toolchain (10-09)
- 1.3 Main repo new batch: ten PRs from #3460 to #3469 — PTO path fixes, A5 double copy, FlashMLA example among them (10-09)
- II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)
- 2.1 Ascend: Expert v2/v3 flash attention examples merged — two generations of implementation benchmarked against v1 at roughly 1.6x to 1.9x speedup (10-09)
- 2.2 Ascend: xAttention example cleanup and shared-prefix decoding kernel merged (10-09)
- 2.3 Ascend: CANN 9.3 self-built image CI verification merged; MLIR path report and argmax merged (10-09)
- 2.4 Hygon: warp size fix merged — gfx92a/gfx946 implement 512 VGPR (10-09)
- 2.5 MetaX / Moore Threads: official repos quiet (10-09 check)
- III. Ecosystem and Adopters
- 3.1 TileOPs: nine merges in a single day — csrc migration completed and linear attention naming unified (10-09)
- 3.2 TileOPs: NSA varlen TMA live-slot ring merged; read-only paged GQA RoPE merged (10-09)
- 3.3 TileOPs: decoding off SM90 and GroupNorm wide-group fix; four performance/fix PRs under review (10-09)
- 3.4 Nightly pipeline: new snapshot with 1030 correctness items and 1945 benchmark items, zero failures (10-10)
- 3.5 TileFoundry: block-scaled FP8 GEMM via SM90 WGMMA bridges HIR to TIR (10-09)
- 3.6 Adopters and docs: TileRT benchmark doc under review; one docs site PR (10-09 check)
- IV. Community, Tutorials, and Events
- 4.1 Community: tilelang-debugger advances with eight dense PRs — nested observation and unified capture (10-09)
- 4.2 PTO-ISA: teaching repo wording revisions; SuperNPUBench continues iterating (10-09)
- 4.3 Media and academia: scattered Google News hits; zero HN entries, no new arXiv preprints (10-09/10-10)
- V. Trend Observations
- 5.1 From “clearing the backlog” to “freezing the increment”: engineering guardrails enter an enforcement phase
- 5.2 Both ends of the queue: 177 open PRs and 1030/1945 items with zero failures
- 5.3 Ascend’s “three-front synchronized cadence” continues and the PTO maintenance surface expands
Today’s Highlights: Friday Merge Wave — 21 Merges in a Single Day to Close Out the Organization, with the TileOPs Refactor Batch, Ascend Example Line, and TileFoundry FP8 Block Scaling All Landing the Same Day (10-09)
Date: 2026-10-09 Sources: Main repo PR list / TileOPs migration wrap-up (#2495) / Ascend Expert v2/v3 (#1863) / TileFoundry block-scaled GEMM (#213) / Nightly snapshot (821084ca)
The organization merged 21 PRs in the window (flat versus the previous window) and opened 29 new ones; 9 repositories saw pushes, with activity concentrated from the afternoon of 10-09 into late night. The 21 merges break down as: main repo 6, TileOPs 9, Ascend adaptation repo 2, MLIR path 1, Hygon 1, TileFoundry 1, supporting docs 1. Four lines landed on the same working day:
- TileOPs 9 merges in one day: csrc migration wrap-up (existing backlog cleared to zero with a guardrail rejecting new additions), linear attention naming unification (GLA/DeltaNet, with no legacy aliases retained), NSA varlen TMA live-slot ring, read-only paged GQA RoPE — refactoring, performance, and features all closed out on the same day (see Chapter 3 for details).
- Ascend dual lines: the Expert v2/v3 flash attention example and the xAttention shared-prefix decode kernel merged in succession (v2 is roughly 1.61 to 1.62x faster than the old version, v3 roughly 1.92x); the same day, the CI verification PR for the self-built CANN 9.3 image (#3458) merged — the example line and the infrastructure line moving in lockstep.
- TileFoundry lands block-scaled FP8 GEMM: #213 takes DeepSeek-V3-style block-scaled FP8 GEMM (e4m3 operands, f32 1×128 and 128×128 scaling) through SM90 WGMMA all the way from HIR to TIR — filling in the MatMul result dtype, the Dim operation cost model, and the Wgmma data type surface.
- Nightly pipeline: a new snapshot in the early hours of 10-10 (corresponding to merge point #2495): 1030 correctness items, 1945 benchmark items (52 suites), zero failures, zero errors, zero skips — test assets refreshed in sync with the day’s merges.
- The 6 main-repo merges are mainly fixes and engineering governance: T.view scalar bit width, Autotuner input reuse, autodd suspension, LDGSTG test gating, CodeRabbit review routing, CANN 9.3 CI.
I. Core Project Progress
1.1 Main repo fix batch: four merges covering T.view scalar bit width, Autotuner input reuse, autodd hang, and LDGSTG test gating (10-09)
Date: 2026-10-09 Source: T.view scalar bit width (#3382) / Autotuner input reuse (#3415) / autodd hang fix (#3385) / LDGSTG test gating (#3463)
- T.view scalar bit width (#3382, merged 17:38):
bits_productreturned 1 directly for empty shapes and skipped dtype bit width and channel count, causing scalarT.viewto accept views with mismatched bit widths while wrongly rejecting scalar/tensor views with equal bit widths (bug #3378); the fix adds CPU-side regressions (bit-width mismatch, vector channels, same-width reinterpretation, scalar/tensor rank). - Autotuner input reuse (#3415, merged 15:33): with
cache_input_tensors=True, comparing TVM dtype against PyTorch dtype directly always mismatched (bug #3412), so compatible inputs were regenerated on every call and missing rank checks were masked; the fix switches to the existingtorch_dtype()conversion, rejecting mismatched ranks first and then comparing dimensions. - autodd hang (#3385, merged 15:34):
python -m tilelang.autodd --jobs 0never returned because no worker threads were spawned; non-positive values are now rejected outright. - LDGSTG test gating (#3463, merged 13:15): the nested-predicate test added in #3390 called the CUDA-only
LowerLDGSTGpass on ROCm CI and errored; arequires_cudagate was added—surfaced by a failing ROCm job and fixed the same day.
1.2 Main repo: CodeRabbit intent-based review routing merged—review governance enters the toolchain (10-09)
Date: 2026-10-09 Source: CodeRabbit review routing (#3456)
- The reviewer distribution list is now maintained directly in
.coderabbit.yaml(using the native “suggested reviewers” directive, with no separate roster file or dispatch workflow): one rule per reviewer, split into four sections—”primary scope / suggested triggers / discouraged triggers / path hints”—matching by change intent rather than filename or keyword, and explicitly excluding the PR author (8 files). - Context: the main repo has 177 open PRs (174 last issue)—review bandwidth governance is shifting from “adding people” to “rule-based routing.”
1.3 New main repo batch: ten PRs from #3460 to #3469—PTO path fix, A5 dual copy, FlashMLA example among them (10-09)
Date: 2026-10-09 Source: PTODSL all-reduce helper re-vendoring (#3469) / A5 dual-copy path conversion (#3468) / FlashMLA manual scheduling example (#3464) / TMA store fence hoisting (#3462)
- PTODSL helper re-vendoring (#3469): upstream PTOAS deleted
ptodsl/_allreduce.pyand expected downstream to vendor it themselves; buttilelang/contrib/ptodsl/simt.pystill referenced the deleted module, causingModuleNotFoundErrorwhen compiling any PTO target operator against the latest CANN Weekly (09-30). This PR moves the implementation intocontrib/ptodsl/allreduce.pyand adds PTO reduction primitive tests—a fresh example of an upstream dependency change directly impacting the main repo’s PTO path. - A5 dual copy (#3467 closed, reopened as #3468): FixPipe on A5 cannot enable dual destination and
quant_presimultaneously, andT.dual_copy’s L0C fp32 to UB f16/bf16 conversion previously required an extravcvt; this PR switches to “convert along the path,” eliminating the extra Vector conversion. - Other new PRs: CANN 9.2.0 documentation requirements (#3460), bisheng wrapper temp file cleanup (#3461), TMA store proxy fence execution order (#3462), Ascend FlashMLA manual scheduling example (#3464), CUDA atomic region rank diagnostics (#3465), Ascend scalar rsqrt sinking fix (#3466)—four of these are Ascend-related.
II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)
2.1 Ascend: Expert v2/v3 Flash Attention Examples Merged — Two Generations of Implementation Show ~1.6–1.9x Speedup Over v1 (10-09)
Date: 2026-10-09 Source: Expert v2/v3 Examples (#1863)
- Added
expert_v2/kernel.pyandexpert_v3/kernel.py: v2 starts from the TileLang approach to organize high-performance computation and pipelining; v3 first writes and validates computation and scheduling in CCE, then considers how to express it in TileLang, using workarounds where direct expression is not possible — both routes coexist, with trade-offs documented separately. - Performance (all 3 fa-public test points): v2 achieves ~1.61–1.62x speedup over the old v1, v3 ~1.92x; all three versions pass correctness checks;
bench_test.shautomatically discovers each version’s entry point. - Title carries an auto-tag (pipeline attribution); the Chinese optimization history is archived alongside the kernel in the same directory.
2.2 Ascend: xAttention Examples Reorganized and Shared-Prefix Decoding Kernel Merged (10-09)
Date: 2026-10-09 Source: xAttention Examples (#1864)
- Added xAttention
expert_v2: the shared-prefix portion uses matrix multiplication, while each beam’s private tokens are computed via vector operations in the output update of the last shared block and merged directly on-chip; v2 uses contiguous BF16 input, with a fixed 32 query heads, 8 K/V heads, and head dimension 128, and does not support paged tables. - Moved the original
xattention.pyandxattention_paged.pyintoexpert_v1/(path move only, contents unchanged), with the CI manifest and related links updated accordingly — the xAttention example line is reorganized into v1/v2 tiers.
2.3 Ascend: CANN 9.3 Self-Built Image CI Verification Merged; MLIR Path Report and argmax Merged (10-09)
Date: 2026-10-09 Source: CANN 9.3 Image Verification (#3458) / MLIR Report and argmax (#193) / NPU Profiling Support (#198)
- CANN 9.3 self-built image (#3458, merged at 18:00): Ascend CI jobs switched to a self-built image (CANN 9.3.0 weekly toolchain with 950 operators, Ubuntu 22.04, Python 3.12, pinned to torch 2.10.0+cpu and torch_npu 2.10.0); the image was pulled via a domestic mirror and validated on the a5 (950) runner — the “regression-ready” infrastructure for 950 went from inception to verification within a single cycle. The companion a5 runner migration PR (#3454) is still under review.
- MLIR path (#193, merged at 11:40): Updated the overall performance test report and added mamba operator tests; removed smoke and torch performance items from the benchmark; integrated per-stage timing statistics and retrospective analysis for agent auto-tuning (folded into the self-evolution mechanism); added the argmax operator and its distillation experience.
- Newly opened: Optional Ascend NPU profiling support (#198) and a full documentation overhaul (#197); the 950 image release PR (#196) remains WIP.
2.4 Hygon: warp size Fix Merged — gfx92a/gfx946 Implement 512 VGPRs (10-09)
Date: 2026-10-09 Source: warp size Fix (#17) / Register Pipeline (#18)
- #17 (merged at 18:54): Supports non-constant expression HIP warpSize; gfx92a and gfx946 implement 512 VGPRs — low-level adaptation on the Hygon side continues to converge (co-author from Hygon, credited as huangteng@hygon.cn).
- #18: Register pipeline and warp-diverge still under review (self-reported example_gemm benchmarked against rocblas).
2.5 MetaX / Moore Threads: Official Repos Silent (10-09 Check)
Date: 2026-10-09 Source: MetaX Repo / Moore Threads Repo
- MetaX’s most recent push was 09-20, Moore Threads’ was 09-30 (v0.1.15+musa.1 release date); no new activity within the window.
III. Ecosystem and Adoption
3.1 TileOPs: Nine Merges in a Single Day — csrc Migration Wraps Up and Linear Attention Naming Unified (10-09)
Date: 2026-10-09 Sources: Inline CUDA migration wrap-up (#2495) / GEMM wave kernel migration (#2493) / GLA rename (#2489) / DeltaNet rename (#2492) / Internal naming alignment (#2494) / Docs follow-up (TileOPs.github.io #68)
- csrc migration wraps up (#2493, #2495): The 1D2D FP8 wave kernel’s
_WAVE_SRCmoved intocsrc/fp8_1d2d_helper.h; the warp-specialized MHA backward’sclaim_tile/retiremoved intocsrc/tile_claim.h, approximate math functions moved intocsrc/approx_math.h(tl::approx_reciprocalandtl::approx_exp2), andelementwise/_prelude.pywas deleted — plus a guardrail was added to “reject new inline CUDA source strings”: the existing backlog was cleared in one pass, and incremental additions are frozen from here on. - Linear attention naming unified (#2489, #2492, #2494):
GLAInferenceFwdOprenamed toGLAFwdOp,DeltaNetInferenceFwdOprenamed toDeltaNetFwdOp, with files renamed accordingly (gla/fwd.py,deltanet/fwd.py); the chunked training side uniformly gained aChunkprefix (including kernel-map keys). No legacy name aliases are retained — external imports must be updated in sync; the docs site followed up afterward (#68). - Refactors and migrations totaled 5 merges, docs 1, making up the bulk of the day’s merges.
3.2 TileOPs: NSA varlen TMA Live-Slot Ring Merged; Read-Only Paged GQA RoPE Merged (10-09)
Date: 2026-10-09 Sources: NSA TMA live-slot ring (#2491) / Read-only paged GQA RoPE (#2484)
- NSA varlen TMA (#2491): Added the SM90 kernel
NSAFwdVarlenTMAKernel— Q and each active K/V block enter 128-byte sniff-aligned tiles via TMA; Q is read into registers once; active slots are traversed by bitmask, with computation of the current block overlapping with loading of the next. - Read-only paged GQA RoPE (#2484): Completed FP16/BF16 paged GQA RoPE — logical cache positions, NeoX and interleaved layouts, partial rotation, and safe tail reads; also fixed NaN with negative
sm_scale(scale logits before masking); public signatures unchanged (closes #2232; the FP8 portion is left to #2231); 48 paged tests on the H200 side included.
3.3 TileOPs: Decode Off SM90 and GroupNorm Wide-Group Fix; Four Performance/Fix PRs Under Review (10-09)
Date: 2026-10-09 Sources: Decode off SM90 (#2487) / GroupNorm wide group (#2486) / Paged GQA Hopper speedup (#2499) / KDA prefill off SM90 (#2498) / DSA seesaw decode (#2497) / GQA cache-append retirement (#2490)
- Decode off SM90 (#2487): The single-token decode kernel lists for DeltaNet, GDN, and KDA were expanded to [80, 89, 90] — they can now run on SM80/89 without SM90 features; the corresponding tests dropped the
sm90marker. - GroupNorm wide group (#2486): Added a split kernel to handle ultra-wide groups that don’t fit in a single block — scenarios such as register rows exceeding 65536 columns and FP16/BF16 rows via shared memory exceeding 16384 columns are handled by sharded reduction followed by merging.
- Four under review: Hopper paged GQA speedup (#2499: self-reported 17/17 comparable benchmark cases faster than FlashInfer on H200, 2.61x for mixed serving, and supports reading active offsets across CUDA Graph replays), KDA chunked prefill off SM90 (#2498), the “seesaw” dual-consumer kernel for DSA sparse MLA decode (#2497), and GQA cache-append prefill retirement (#2490, with the read-only paged API retained).
- TileOPs currently has 6 open PRs.
3.4 Nightly Pipeline: New Snapshot with 1030 Correctness Items, 1945 Benchmarks, Zero Failures (10-10)
Date: 2026-10-10 Sources: Snapshot commit (821084ca) / Snapshot environment metadata
- New snapshot in the early hours of 10-10 (run 37970266975, corresponding to merge point 7c351e304, i.e. #2495): 1030 correctness items with zero failures, zero errors, zero skips; 1945 benchmarks (52 suites) with zero failures — up 3 and 6 items respectively from the previous snapshot, with scale continuing to climb.
- Runtime environment: H200, CUDA 13.2, driver 595.71.05, torch 2.13.0+cu132.
3.5 TileFoundry: Block-Scaled FP8 GEMM via SM90 WGMMA Connects HIR Through to TIR (10-09)
Date: 2026-10-09 Sources: Block-scaled FP8 GEMM (#213) / Fix under review (#220)
- #213 (merged 23:46): Advanced DeepSeek-V3-style block-scaled FP8 GEMM (e4m3 operands, f32 1×128 and 128×128 scaling) from “can’t write it, can’t parse it, can’t schedule it, can’t evaluate it” to a fully runnable end-to-end path — filling in four foundations: MatMul result dtype, Dim operation cost model, Wgmma extended from BF16 to block-scaled FP8, and HIR on-chip matmul result placement changed to write by target MMA (rather than being operand-anchored); supersedes the closed #212.
- The same day, 3 other fix PRs (#216, #217, #219) were opened and then closed (not merged); 1 is under review (#220: mesh-scope calls and scalar window fix).
3.6 Adopters and Docs: TileRT Benchmark Docs Under Review; 1 Docs Site Commit (10-09 Check)
Date: 2026-10-09, 2026-10-10 Sources: TileRT benchmark docs (#58) / Main docs site commit (0158a427) / TileKernels / FlashQLA
- TileRT: #58 (PD deployment benchmark docs for GLM-5.1 and vllm: latency and throughput metrics) is under review, with updates still coming in before deadline; #67 (EFA user-space components) is under review.
- Docs site: 1 build update for tilelang.github.io (autoapi reference pages and search index for the autotuner/tuner modules); TileOPs.github.io followed up with 1 commit for the renames (see 3.1).
- Adopters: TileKernels and FlashQLA had no pushes within the window (both most recently on 09-30); release cadence: the main repo’s v0.1.15 (09-30) remains the latest tag, with no new releases in the window.
IV. Community, Tutorials, and Events
4.1 Community: tilelang-debugger advances with eight dense commits—nested observation and unified capture (10-09)
Date: 2026-10-09 Source: tilelang-debugger / unified capture commit (e838fc88)
- Eight commits within the window (10-09 morning to evening): accepting user-specified kernel source paths, source-selection capture via Python print macros, typed sample recording with nested observation and unified nested control flow, splitting source analysis/instrumentation/emitter/protocol, separating thread selection from the reader, and closing with “unified source capture”; one more commit landed before deadline (10-10 morning) (unified source access capture with TileLang compatibility validation).
- A TileLang-specific debugging workspace targeting H200, maintaining a pace of nearly ten commits per day for a second consecutive window (4 commits on 10-08, the day it was first created).
4.2 PTO-ISA: teaching repo wording revised; SuperNPUBench continues iterating (10-09)
Date: 2026-10-09 Source: tilelang-puzzles-ascend / SuperNPUBench
- Teaching repo (puzzles-ascend): 2 commits within the window (torch track annotates dtype and shape for each tensor; puzzle file wording changed to guide readers to reason on their own rather than pointing directly at answers).
- SuperNPUBench: 2 commits within the window, plus another 2 this morning before deadline—added an FA operator performance workbook, fixed rms_norm 32k test default PE_NUM=4, improved single-block path 32k gamma, and added inf/zero/special triple guards to tail_ocp_fp8.
4.3 Media and Academia: scattered Google News hits; zero HN entries, no new arXiv preprints (10-09/10-10)
Date: 2026-10-09 Source: Google News / Hacker News search / arXiv
- Google News: multiple Chinese and English queries returned 4 hits (1 English tech briefing, 3 Chinese financial articles)—judged by title and source, none are independent coverage of this project; they are background continuation and market interpretation of the “DeepSeek open-sources Ascend components” coverage wave; precise combined-term queries (past 2 days, past 3 days) yielded no other new items.
- Hacker News: zero entries within the window (continuing the thin HN coverage of this topic).
- arXiv: no new preprints within the window (the latest related item retrieved is a 2026-07-05 GPU kernel evaluation study).
V. Trend Observations
5.1 From “clearing the backlog” to “freezing the increment”: engineering guardrails enter an enforcement phase
- TileOPs turns “inline CUDA source strings” from a migration project into a freeze rule (rejecting new additions after the backlog is cleared); linear attention renaming leaves no aliases; on the main repo side, CodeRabbit intent routing regularizes reviewer distribution; Ascend CI is pinned to a self-built CANN 9.3 image. The common thread across these engineering-quality actions: all are upgrading one-off cleanups into long-term constraints—the refactoring phase ends, the maintenance phase begins.
5.2 Both ends of the queue: 177 open PRs and 1030/1945 zero-failure items
- Open PRs: 177 in the main repo (+3 vs. the previous period), 225 in the Ascend adaptation repo—review bandwidth remains the pacing variable; at the other end, the nightly pipeline’s 1030 correctness items plus 1945 benchmark items (52 suites) continue with zero failures, providing a regression foundation for the large pending review queue. Of the 29 newly opened within the window, 10 were merged the same day (mostly TileOPs), with “fast in, fast out” and “long-tail queue” coexisting.
5.3 Ascend’s “three-front simultaneous cadence” continues and the PTO maintenance surface expands
- Features (Expert examples, A5 double copy, rsqrt), infrastructure (CANN 9.3 image validation), and compilation paths (MLIR reports and argmax) all advanced on multiple fronts the same day; meanwhile #3469 flags a new variable: interface changes in upstream PTOAS directly break compilation of the main repo’s PTO targets—the PTO second path enters a phase of “upstream dependency management.” Before deadline (10-10 morning), Ascend-related commits were still updating (#3466, #3468, #3469, and #3321, #1657, etc.).
Appendix: Sources and Verification Notes
| Source | Verification Result |
|---|---|
| tile-ai organization (29 repos) | 9 repos had pushes within the window: tilelang (6), TileOPs (9), tilelang-ascend (2), tilelang-hygon (1), tilelang-mlir-ascend (1), TileFoundry (1), tilelang.github.io (1), TileOPs.github.io (1), TileOPs-nightly (1 snapshot); 21 merges and 29 new opens within the organization window |
| Main repo tilelang | 6 merged (#3382, #3385, #3415, #3456, #3458, #3463); 10 newly opened (#3460 to #3469, of which #3467 was closed and reopened as #3468); 177 open PRs; no additions to releases page or tags within the window (latest v0.1.15, 09-30) |
| Ascend repos (ascend / mlir-ascend) | ascend: 2 merged (#1863, #1864), 2 newly opened (#1866, #1867), 225 open PRs; mlir-ascend: 1 merged (#193), 2 newly opened (#197, #198), #196 still WIP |
| Hygon / MetaX / Moore Threads | Hygon: #17 merged, #18 under review; MetaX (after 09-20) and Moore Threads (after 09-30) had no activity |
| TileOPs | 9 merged (#2484, #2486, #2487, #2489, #2491, #2492, #2493, #2494, #2495); 10 newly opened; 6 open PRs |
| TileOPs-nightly | New snapshot 821084ca (run 37970266975): 1030 correctness items, 1945 benchmark items (52 suites), zero failures, zero errors, zero skips; corresponding merge point 7c351e304 (#2495) |
| TileFoundry | 1 merged (#213); 4 newly opened (#216, #217, #219 closed without merge, #220 under review) |
| TileRT | #58, #67 under review; #58 updated before deadline; no additions to releases page (latest v0.1.6, 09-24) |
| PTO-ISA organization | tilelang-puzzles-ascend: 2 within the window; SuperNPUBench: 2 within the window, plus 2 more this morning before deadline |
| Adopters and community | TileKernels and FlashQLA had no pushes within the window; community repo tilelang-debugger had 8 within the window (plus 1 more before deadline); BM1690 pipeline repo had no updates within the window |
| Google News / Hacker News / arXiv | Google News Chinese and English queries returned 4 hits (1 English briefing, 3 Chinese financial articles); HN zero entries; arXiv no new preprints |
Full Source List
- [1] Main repo PR list — https://github.com/tile-ai/tilelang/pulls
- [2] T.view scalar bit width (#3382) — https://github.com/tile-ai/tilelang/pull/3382
- [3] Autotuner input reuse (#3415) — https://github.com/tile-ai/tilelang/pull/3415
- [4] autodd hang fix (#3385) — https://github.com/tile-ai/tilelang/pull/3385
- [5] LDGSTG test gating (#3463) — https://github.com/tile-ai/tilelang/pull/3463
- [6] CodeRabbit review routing (#3456) — https://github.com/tile-ai/tilelang/pull/3456
- [7] PTODSL all-reduce helper revert (#3469) — https://github.com/tile-ai/tilelang/pull/3469
- [8] PTOAS upstream change commit (f5eff3ee2) — https://github.com/hw-native-sys/PTOAS/commit/f5eff3ee249697f6157088f649c6434fcc9d7c5b
- [9] A5 dual-copy path conversion (#3468) — https://github.com/tile-ai/tilelang/pull/3468
- [10] FlashMLA manual scheduling example (#3464) — https://github.com/tile-ai/tilelang/pull/3464
- [11] TMA store fence hoisting (#3462) — https://github.com/tile-ai/tilelang/pull/3462
- [12] CANN 9.2.0 documentation requirement (#3460) — https://github.com/tile-ai/tilelang/pull/3460
- [13] Atomic region rank diagnostics (#3465) — https://github.com/tile-ai/tilelang/pull/3465
- [14] Ascend scalar rsqrt fix (#3466) — https://github.com/tile-ai/tilelang/pull/3466
- [15] Expert v2/v3 example (#1863) — https://github.com/tile-ai/tilelang-ascend/pull/1863
- [16] xAttention example (#1864) — https://github.com/tile-ai/tilelang-ascend/pull/1864
- [17] CANN 9.3 image verification (#3458) — https://github.com/tile-ai/tilelang/pull/3458
- [18] a5 runner migration (#3454) — https://github.com/tile-ai/tilelang/pull/3454
- [19] MLIR reporting and argmax (#193) — https://github.com/tile-ai/tilelang-mlir-ascend/pull/193
- [20] NPU profiling support (#198) — https://github.com/tile-ai/tilelang-mlir-ascend/pull/198
- [21] Full documentation overhaul (#197) — https://github.com/tile-ai/tilelang-mlir-ascend/pull/197
- [22] 950 image release (#196) — https://github.com/tile-ai/tilelang-mlir-ascend/pull/196
- [23] warp size fix (#17) — https://github.com/tile-ai/tilelang-hygon/pull/17
- [24] Register pipelining (#18) — https://github.com/tile-ai/tilelang-hygon/pull/18
- [25] MetaX repo — https://github.com/tile-ai/tilelang-metax
- [26] Moore Threads repo — https://github.com/tile-ai/tilelang-musa
- [27] Inline CUDA migration wrap-up (#2495) — https://github.com/tile-ai/TileOPs/pull/2495
- [28] GEMM wave kernel migration (#2493) — https://github.com/tile-ai/TileOPs/pull/2493
- [29] GLA rename (#2489) — https://github.com/tile-ai/TileOPs/pull/2489
- [30] DeltaNet rename (#2492) — https://github.com/tile-ai/TileOPs/pull/2492
- [31] Internal naming alignment (#2494) — https://github.com/tile-ai/TileOPs/pull/2494
- [32] Documentation follow-up (TileOPs.github.io #68) — https://github.com/tile-ai/TileOPs.github.io/pull/68
- [33] NSA TMA live slot ring (#2491) — https://github.com/tile-ai/TileOPs/pull/2491
- [34] Read-only paged GQA RoPE (#2484) — https://github.com/tile-ai/TileOPs/pull/2484
- [35] Decode off SM90 (#2487) — https://github.com/tile-ai/TileOPs/pull/2487
- [36] GroupNorm wide group (#2486) — https://github.com/tile-ai/TileOPs/pull/2486
- [37] Paged GQA Hopper acceleration (#2499) — https://github.com/tile-ai/TileOPs/pull/2499
- [38] KDA prefill off SM90 (#2498) — https://github.com/tile-ai/TileOPs/pull/2498
- [39] DSA seesaw decode (#2497) — https://github.com/tile-ai/TileOPs/pull/2497
- [40] GQA cache append retirement (#2490) — https://github.com/tile-ai/TileOPs/pull/2490
- [41] Snapshot commit (821084ca) — https://github.com/tile-ai/TileOPs-nightly/commit/821084caf4f944ff9899a3671498f39c3bfb9363
- [42] Snapshot environment metadata — https://github.com/tile-ai/TileOPs-nightly/blob/snapshots/meta.json
- [43] Block-scaled FP8 GEMM (#213) — https://github.com/tile-ai/TileFoundry/pull/213
- [44] Fix under review (#220) — https://github.com/tile-ai/TileFoundry/pull/220
- [45] TileRT benchmark documentation (#58) — https://github.com/tile-ai/TileRT/pull/58
- [46] EFA user-space component (#67) — https://github.com/tile-ai/TileRT/pull/67
- [47] Main documentation site commit (0158a427) — https://github.com/tile-ai/tilelang.github.io/commit/0158a42782d462781b6afd728f0bcd2530cf79a1
- [48] TileKernels repo — https://github.com/deepseek-ai/TileKernels
- [49] FlashQLA repo — https://github.com/QwenLM/FlashQLA
- [50] tilelang-debugger — https://github.com/superAngGao/tilelang-debugger
- [51] Debugger unified capture commit (e838fc88) — https://github.com/superAngGao/tilelang-debugger/commit/e838fc88d96a562ee6230ee154d6593bac6632f8
- [52] tilelang-puzzles-ascend — https://github.com/PTO-ISA/tilelang-puzzles-ascend
- [53] SuperNPUBench — https://github.com/PTO-ISA/SuperNPUBench
- [54] BM1690 pipeline repo — https://github.com/arcflute/tilelang-tpu-bm1690-pipelines
- [55] Main repo releases page — https://github.com/tile-ai/tilelang/releases
- [56] Google News RSS (Chinese and English queries) — https://news.google.com/
- [57] Hacker News search — https://hn.algolia.com/
- [58] arXiv search — https://arxiv.org/