Research window: Past 24 hours (2026-09-24 07:00 ~ 2026-09-25 07:00, Beijing time). Regular daily update window, no overlap with the previous issue. Sources: GitHub (push verification across 29 repos in the tile-ai organization, 11 repos had pushes within the window; two TileRT release PRs merged the same day; main repo NaN propagation fix merged on two fronts, SIMT im2col and async staging refactor chain taking shape; ten merges in TileOPs with cleanup of legacy LinearAttention paths initiated; fixes merged in TileFoundry, Ascend, Hygon, and MetaX respectively; nightly snapshot correctness 1140 items, benchmarks 1081 items, all zero errors), Google News RSS multi-language queries (via proxy, zero hits), Hacker News, arXiv, community repo verification


In This Issue

  • Today’s Highlight: Two TileRT release PRs merged in one day—v0.1.6 brings GLM-5.2/5.3 to AMD Instinct MI350X/MI355X (ROCm) (09-24)
  • I. Core Project Progress
    • 1.1 Main repo: NaN propagation fix merged on two fronts, two defect tickets closed (09-24)
    • 1.2 Main repo: SIMT im2col and async staging—performance defect ticket and two-stage refactor PR take shape on the same day (09-24)
    • 1.3 Main repo: Parallel loop and async copy fixes merged, RNG queue pending review (09-24)
    • 1.4 TileOPs: LinearAttention consolidation—GLA inference operators and SM90 DeltaNet decode merged (09-24)
    • 1.5 TileOPs: GEMM long-K streaming merged, FFT kernel rewrite proposed (09-24)
    • 1.6 TileOPs: Backend channel and test infrastructure governance wrapped up (09-24)
    • 1.7 TileOPs: Four legacy path cleanup drafts, H200 GLA decode pending review (09-24)
    • 1.8 TileFoundry: isl recursive traversal fix merged, shard permutation refactor draft follows up (09-24)
    • 1.9 Overnight snapshot: 1,140 correctness items and 1,081 benchmark items all zero errors (09-25)
  • II. Multi-Backend Adaptation (Ascend / Hygon / MetaX / Moore Threads, etc.)
    • 2.1 Ascend: Sparse FlashAttention example merged, vector mask and row reduction optimizations landed (09-24)
    • 2.2 Ascend: API documentation verification queue advances in batches, operator migration and race condition fixes pending review (09-24)
    • 2.3 MLIR Ascend: A5 developer branch cleanup wraps up (09-24)
    • 2.4 Hygon: FP32 MMAC K sharing and ds_read_format packing fix merged (09-24)
    • 2.5 MetaX: Upstream main synced and merged into dev, sub-warp guard fix ticket opened (09-24)
    • 2.6 Moore Threads: No pushes to repository (09-25 check)
  • III. Ecosystem and Adopters
    • 3.1 Adopters: TileKernels and FlashQLA no pushes in window (09-25 check)
    • 3.2 Community projects: ETH Zurich TileLang SoftHier extension updated, targeting RISC-V many-core accelerators (09-24)
    • 3.3 Release cadence: TileRT releases page and tags pending update, main repo v0.1.14 now 23 days old (09-25)
  • IV. Community, Tutorials and Events
    • 4.1 Documentation sites: Main repo and TileOPs documentation sites routinely regenerated (09-24)
    • 4.2 Media and academia: Zero Google News hits, no new arXiv preprints (09-25 check)
  • V. Trend Observations
    • 5.1 Release cadence: TileRT running patch line and minor version line in parallel
    • 5.2 Second wave of explicitization: From semantic validation to scheduling and staging
    • 5.3 LinearAttention family enters consolidation phase
    • 5.4 Gaps and risk points
  • Appendix: Sources and verification notes

Today’s Highlight: Two TileRT release PRs merged in one day—v0.1.6 brings GLM-5.2/5.3 to AMD Instinct MI350X/MI355X (ROCm)

Date: 2026-09-24 Source: v0.1.6 release PR (#68)/v0.1.5.post3 release PR (#63)/TileRT releases page

  • On the evening of 09-24 (20:22 and 23:33 Beijing time), TileRT merged two release PRs in succession: v0.1.5.post3 (#63, created 09-07, merged within the window) and v0.1.6 (#68, created and merged the same day, merged by project member Lingxiao Ma). Both entered the main branch.
  • The v0.1.6 release notes state: Added GLM-5.2/5.3 support on AMD Instinct MI350X/MI355X (ROCm path), and includes PD disaggregation (prefill/decode separation) related updates. On the same day, #66 (a performance ticket for decode KV injection copy overlapping across device pairs on the PD side) was closed, with its content folded into this release narrative.
  • It should be noted: as of this report’s verification time (early morning 09-25), TileRT’s GitHub Releases page and tags still show v0.1.5.post2 (2026-08-06)—the release PRs have been merged, but the public releases page has not yet caught up. This daily records it as “release process initiated, page pending update,” and the next issue will continue to verify.

Interpretation: This is the first set of new increments after TileRT’s monthly major version (v0.1.5 series released 08-06): the patch line (post3) and the minor version line (v0.1.6) advance on the same day, and ROCm / AMD Instinct MI350 series are placed in the most prominent position in the release notes. For its positioning as an “ultra-low-latency inference runtime,” model coverage across multiple hardware platforms is beginning to ramp up in parallel with the NVIDIA line.


I. Core Project Progress

1.1 Main Repo: NaN Propagation Fixes Merged on Two Fronts, Two Defect Tickets Closed (09-24)

Date: 2026-09-24 Source: #3273/#3270/#3205/#3024

  • #3273 (merged 17:03 Beijing time): Adds a compile-time error for unsupported reduction NaN-propagation dtypes — the diagnostic reports the operation name, supported output types, and argument dtypes; the check is wired into the native ReduceOp constructor, uniformly enforcing the accumulator dtype contract across local reductions, fragment reductions, and write-back accumulation. The corresponding defect ticket #3270 (reduce_max/min/absmax with nan_propagate=True silently ignored on float32/float64) was closed upon merge — roughly a day and a half from report on 09-23 to fix merge.
  • #3205 (merged 13:44 Beijing time): Fixes T.clamp NaN-propagation semantics — correctly lowers NaN-propagating clamp to the device template. The corresponding defect ticket #3024 was closed; it originated from fuzzing that found a wrong-code class of issue where “clamp silently drops NaN and returns the lower bound.”
  • Both lines landed the same day, meaning the “NaN propagation” semantic gap that surfaced in the previous issue is now closed: one side moves unsupported configurations forward into compile-time failures, the other corrects silent wrong behavior into expected propagation semantics.

1.2 Main Repo: SIMT im2col and Async Staging — Performance Defect Ticket and Two-Stage Refactor PR Take Shape the Same Day (09-24)

Date: 2026-09-24 Source: #3274/#3275/#3279/#3278

  • At 12:48 Beijing time, performance defect ticket #3274 was filed: SIMT im2col is lowered too late, missing scheduling and vectorization opportunities; three minutes later #3275 (CUDA/Transform, hoisting SIMT im2col loads to before scheduling to expose them) was filed — the “report the same day, propose a solution the same day” cadence repeats.
  • At 23:22 that day, #3278 (Transform/CUDA: fix async copy lowering under partitioned layouts) was merged — the fix for async copies under partitioned layouts landed on main.
  • At 00:30 on 09-25, #3279 was filed: “Unified async data staging and transfer-aware lowering” — the author describes it as an aggressive refactor of #3275, with “substantial performance gains” for convolution, mHC prologue, Bounded GEMM, and other scenarios; it remains marked experimental. Key points (per the PR summary): separate logical-level tile operator expansion from physical lowering, with CUDA/CPU/Metal/ROCm/WebGPU pipelines exposing logical accesses before scheduling and layout inference; introduce a shared transfer analysis (describing memory access effects, value conversions, execution guards, and TMA usage) that lets pipeline planning identify prefetch stages and async candidates, lets warp specialization analysis identify SIMT producers, and lets CUDA and ROCm async copy injectors identify convertible transfers; pipeline wait planning begins to support consumers depending on multiple async producer stages; vectorization begins widening access ranges for vectorized cp.async; thread synchronization analysis tracks the physical byte footprint of index and pointer accesses; when a producer prefix writes shared memory, a shared memory synchronization edge is added before TMA publication. It also includes developer documentation and updates the im2col fallback test to make the number of pipeline stages configurable.
  • #3276 (CUDA: preserve FP8 vector copy packing state) was filed the same day and is under review.

Assessment: The three-stage progression from “defect ticket → early access exposure → unified staging and transfer awareness” took shape on the same day, directly targeting the performance skeleton of convolution and GEMM-class workloads. Separating “logical access” from “physical lowering” is consistent with earlier threads such as “random numbers, layout proofs, NaN semantics” — the compiler is systematically making implicit behavior explicit in exchange for composability and diagnosability.

1.3 Main Repo: Parallel Loop and Async Copy Fixes Merged, RNG Queue Pending Review (09-24)

Date: 2026-09-24 Source: #3269/#3267/#3241 et al.

  • #3269 (merged 17:45 Beijing time): Fixes parallel loop lowering when let inlining is disabled — under the “let inlining off” compile configuration, parallel loop lowering produced incorrect results, a defect-class fix for a long-tail configuration combination.
  • #3267 (tl.LowerMagicDiv host-side precomputed magic division) continued to be updated within the window and remains unmerged; #3241–#3244 (CuTeDSL FP4 conversion/packing output, RNG void binding rejection, RNG sequence inference, RNG missing-initialization diagnostics) — four RNG queue tickets — refreshed as a batch between 13:54 and 14:12 Beijing time and remain open.
  • The big items in the main repo review queue (TileIR backend execution PR, magic division, etc.) saw no new merges within the window — the “stamp small fixes, stockpile big changes” cadence continues.

1.4 TileOPs: LinearAttention Consolidation — GLA Inference Operator and SM90 DeltaNet Decode Merged (09-24)

Date: 2026-09-24 Source: #2174/#2194/#2206

  • #2174 (merged 19:33 Beijing time): GLA (Gated Linear Attention) inference operator added, with a tiled prefill path — together with the previous window’s Gated DeltaNet inference operator (#2163), this forms dual-operator coverage of the linear attention family.
  • #2194 (merged 19:34 Beijing time): SM90-platform Gated DeltaNet decode kernel — completing the DeltaNet inference chain through the decode stage on the Hopper architecture.
  • #2206 (merged 22:26 Beijing time): Fixes GLA roofline state naming consistency, aligning the performance model (roofline) state definitions with the operator implementation; this commit then became the build target of that night’s snapshot (see 1.9).

Assessment: The linear attention direction completed multi-point placement across “training side → inference side → platform decode → performance model definitions” within two days. Combined with the cleanup drafts in 1.7, one can say this family has moved from an “expansion phase” into a “consolidation phase.”

1.5 TileOPs: GEMM Long-K Streaming Merged, FFT Kernel Rewrite Opened (09-24)

Date: 2026-09-24 Source: #2185/#2187/#2172/#2178

  • #2185 (merged 19:36 Beijing time): W4A16 GEMM’s decode scaling latency + long-K loop streamed across all SMs — opened in the previous window, merged in this one, landing performance engineering for the long-K scenario on the 4-bit weight quantization path.
  • #2187 (filed 10:34 Beijing time, under review): FFT kernel rewrite — replaces the LUT (lookup table) kernel with a new approach and extends power-of-two coverage to 2^28 — replacing the lookup table means the FFT path is converging toward a computational implementation.
  • The queue also holds #2172 (dense GEMM ping-pong main loop to hide the epilogue + 176-tile option) and #2178 (varlen GQA scheduling pre-arrangement), both open and continuously updated.

1.6 TileOPs: Backend Channel and Test System Governance Wraps Up (09-24)

Date: 2026-09-24 Source: #2186/#2181/#2183 et al., six tickets (full list in appendix)

  • The governance mainline started early in the morning (dense merges from 11:17 to 13:41 Beijing time): #2181 (in-tree kernels become the default target of the test suite), #2186 (fixes DeltaNet prefill hang and GEMM roofline anomaly, corresponding to the previous issue’s benchmark timeout problem), #2175 (report failure reason when roofline formula synthesis fails), #2183 (operator behavior tests run on all backends, in-tree assertions take effect only in-tree).
  • Afternoon wrap-up (17:18–17:19): #2179 (external backends serve external targets at operator boundaries — following the earlier removal of device-kind checks, now formally merged), #2193 (fixes memory retention caused by _require_one_device keeping per-call inputs alive).
  • With this, the two items launched in the previous window — “external backend channel governance + test default target switch” — were fully merged within 24 hours, the most complete stretch of engineering debt repayment this week.

1.7 TileOPs: Four Legacy Path Cleanup Drafts, H200 GLA Decode Pending Review (09-24)

Date: 2026-09-24 Source: #2200/#2201/#2196 et al., eight tickets (full list in appendix)

  • From morning to evening (15:05–16:57 Beijing time), four refactor/maintenance drafts were filed in a batch (all in draft state): #2196 (remove unowned Gated DeltaNet kernels), #2197 (remove legacy varlen GQA path), #2200 (remove legacy GLA inference path), #2201 (rename GLA forward and chunked operators) — all pointing uniformly toward “delete old, rename, converge interfaces.”
  • Also in the same period: #2198 (D64 varlen GQA scheduling to WS kernel), #2199 (H200-platform GLA dense decode), #2202 (migrate paged GQA decode to unified operator), #2203 (build workload inputs on caller-specified devices), among other drafts.
  • Within the window, #2195, #2205, #2180 were also closed without merging after review; their content (lifecycle ownership, backend call consistency, varlen scheduling persistence) is being advanced in split form.

Assessment: Four “remove legacy path” drafts opened in a single day, plus the unification of paged GQA, indicate the operator library is proactively shrinking its interface surface after functional expansion. Such cleanups typically precede an interface freeze or version release — worth watching whether a new release cadence follows.

1.8 TileFoundry: isl Recursive Traversal Fix Merged, Shard Arrangement Refactor Draft Follows (09-24)

Date: 2026-09-24 Source: #184/#185

  • #184 (merged 12:50 Beijing time): fix(ir) — avoids recursive traversal of isl ranges (a fix triggered by recursion depth issues), a stability-class fix for the compile-time IR analysis path.
  • #185 (continued updating into the early hours of 09-25): refactor(shard) — a refactor draft declaring only one arrangement per grid level, continuing to sort out shard semantics.

1.9 Nightly Snapshot: 1140 Correctness Items, 1081 Benchmark Items, All Zero Errors (09-25)

Date: 2026-09-25 (snapshot committed 02:39 Beijing time) Source: snapshot commit/environment metadata

  • The new snapshot corresponds to commit 85455e46c2 (i.e., the #2206 merge point), with an environment of NVIDIA H200, CUDA 13.2, driver 595.71.05, torch 2.13.
  • Correctness: 1140 items, zero failures, zero errors, 2 skipped; Benchmark: 50 suites, 1081 items, zero failures, zero errors, 3 skipped (up 31 items from the previous snapshot’s 1050).
  • The file-level timeout on the long-prefill configuration of bench_deltanet.py in the previous snapshot no longer occurs — the corresponding fix #2186 was merged that day, and the nightly pipeline is fully green again.

II. Multi-Backend Adaptation (Ascend / Hygon / MetaX / Moore Threads, etc.)

2.1 Ascend: Sparse FlashAttention Example Merged, Vector Mask and Row Reduction Optimizations Landed (09-24)

Date: 2026-09-24 Source: #1741/#1829/#1826

  • #1741 (merged 15:16 Beijing time): examples/cann-bench/sparse_flash_attention operator — the example library extends into sparse attention benchmark scenarios (opened 09-04, merged after two weeks of review).
  • #1829 (merged 10:28 Beijing time): AscendC vector mask management + FP32 row reduction optimization — low-level vector instruction-level performance engineering.
  • #1826 (merged 15:12 Beijing time): Restore Select/Transpose tests and in-place validation — a test-coverage restoration fix.

2.2 Ascend: API Documentation Verification Queue Advances in Batch, Operator Migration and Race Condition Fix Pending Review (09-24)

Date: 2026-09-24 Source: #1657/#1681/#1828/#1834

  • Between 18:50 and 20:29 Beijing time, a dozen-plus API documentation verification PRs were updated in batch: covering interfaces such as T.tile.bitwise_xor/bitwise_not/bitwise_and/bitwise_lshift, relu/leaky_relu/sigmoid/silu, sin/cos, clamp, fill, clear, axpy — adding documentation fact-checks, test cases, and API descriptions for each operator. These reviews had been pending since mid-August; this window saw concentrated progress.
  • Pending review queue: #1828 (migrating the wy_fast_bwd_split backward split operator to Ascend NPU) and #1834 (race condition fix for scalar GM stores and MTE3 DMA write-back cache incoherence, corresponding to defect ticket #1304) both continue to be updated.

2.3 MLIR Ascend: A5 Developer Branch Cleanup Wraps Up (09-24)

Date: 2026-09-24 Source: #93/#130

  • Between 14:41 and 14:43 Beijing time, the repository deleted four branches in a batch (a5-dev, a5-branch, a5-main-sync, revert-86-A5-dev), with corresponding sync PRs #93, #136, #121, #66 closed; no new commits on the main branch during the window — all pushes to this repository in this window came from branch cleanup.
  • #130 (A5 test sync) remains open; test merging on the A5 line has yet to conclude.

2.4 Hygon: FP32 MMAC K Sharing and ds_read_format Packing Fix Merged (09-24)

Date: 2026-09-24 Source: #12

  • #12 (merged 11:10 Beijing time): fix(hcu) — sharing the K dimension of FP32 MMAC and packing ds_read_format fragments. A data layout fix for the Hygon HCU matrix core, following the example support merge (#11) from the previous window.

2.5 MetaX: Upstream main Synced into dev, Sub-warp Guard Fix Opened (09-24)

Date: 2026-09-24 Source: #158/#159

  • #158 (merged 16:06 Beijing time): syncing upstream tilelang main into the dev branch (including merge conflict resolution), keeping the adaptation branch aligned with the mainline.
  • #159 (opened 17:37 Beijing time, under review): [fix bug][MetaXGPU] adds a sub-warp block size guard for T.gemm_sp — out-of-bounds protection for sparse GEMM at small block sizes.

2.6 Moore Threads: No Pushes to Repository (09-25 Check)

Date: 2026-09-25 Source: tilelang-musa repository/TileOPs repository

  • tilelang-musa had no pushes during the window (last push 09-17). The MUSA governance tickets on the TileOPs side (builder registration, test paths, etc.) also saw no new activity this window. Among the four domestic adaptation efforts, MetaX and Hygon maintain a daily update cadence, while Moore Threads is silent for the second consecutive week.

III. Ecosystem and Adopters

3.1 Adopters: No pushes from TileKernels or FlashQLA within the window (09-25 check)

Date: 2026-09-25 Source: TileKernels repo / FlashQLA repo

  • DeepSeek’s TileKernels (last push 2026-04) and Qwen’s FlashQLA (last push 09-18) saw no pushes within the window; no new activity on the adopter side this period.

3.2 Community project: ETH Zurich’s TileLang SoftHier extension updated, targeting RISC-V many-core accelerators (09-24)

Date: 2026-09-24 Source: tilelang-softhier repo

  • The individual community repo tilelang-softhier updated its README and adjusted repository files between 17:15 and 17:19 Beijing time. Per the repo’s own description: this work extends TileLang to tiled many-PE accelerators — beyond computation, kernels must also decide tile ownership, tile movement over the NoC (network-on-chip), and synchronization placement; the target platform SoftHier is a configurable multi-cluster RISC-V accelerator modeled with GVSoC (a semester project at ETH Zurich’s Integrated Systems Laboratory under Professor Luca Benini; the repo is a fork of TileLang).
  • Core design principle: kernel authors declare only “what to compute and where the data is,” while the compiler derives cluster placement, collective-communication routing, and synchronization. Results table: a 2D SUMMA GEMM with a 4×4 cluster configuration and BM=BN=BK=128, N=7168, K=2048 (matching the DeepSeek-V3 routed-expert dimensionality-reduction projection size); at M=512, compiler-generated code runs in 1.01 ms vs. 0.96 ms for hand-written C (a 5.14% gap).
  • Corroborating evidence: the community also has Chinese-language learning materials such as the TVM/TileLang Cookbook (30 stars), whose pushes within the window were star-chart automation (chart branch), with no new content on the main branch; there are also several scattered personal experiment repos with no substantive content.

3.3 Release cadence: TileRT release page and tags pending update, main repo v0.1.14 now 23 days old (09-25)

Date: 2026-09-25 Source: TileRT releases page / tilelang releases page

  • TileRT: two release PRs have been merged, but GitHub Releases and tags remain at v0.1.5.post2 — the publicly available version has not been updated (see today’s highlights).
  • Main repo tilelang: v0.1.14 was released on 09-02 and has had no new version for 23 days as of this period; TileOPs has no standalone release; TileFoundry’s latest remains v0.0.2 (09-10).

IV. Community, Tutorials, and Events

4.1 Documentation sites: routine regeneration of main repo and TileOPs doc sites (09-24)

Date: 2026-09-24 Source: tilelang doc site / TileOPs doc site

  • The main repo doc site received two Update docs bot commits (13:48 and 17:08 Beijing time); the TileOPs doc site had its routine daily gh-pages push (08:19 Beijing time). All are automated regenerations with no content-level changes.

4.2 Media and academia: zero Google News hits, no new arXiv preprints (09-25 check)

Date: 2026-09-25 Source: Google News / Hacker News / arXiv

  • Multiple Google News RSS queries in Chinese and English (via proxy) returned zero hits within the window; Hacker News items in the window are unrelated to this topic; the latest full-repository arXiv search still returns the 07-24 TileSight paper (whose code was open-sourced on 09-23), with no new preprints in the window.

V. Trend Observations

5.1 Release cadence: TileRT running patch and minor release lines in parallel

TileRT merged two release PRs in a single day on 09-24: v0.1.5.post3 and v0.1.6 — the former the third patch on the v0.1.5 line, the latter folding in ROCm (AMD Instinct MI350X/MI355X) support for GLM-5.2/5.3 along with PD-disaggregation updates. Relative to August’s single release, this signals a denser release pipeline; however, the release page and tags lag behind the release PRs, and the public-channel information gap warrants follow-up next period.

5.2 The second wave of explicitization: from semantic validation to scheduling and staging

The previous window’s “NaN propagation” work concluded with dual-track merges of compile-time validation plus semantic fixes and two closed defect tickets; the same principle then shifted to the compiler scheduling layer: early exposure of SIMT im2col loads (#3275), unified async staging with logical access separated from physical lowering (#3279), and fixes for async copies under partitioned layouts (#3278). “Declare first, derive later, implicit to explicit” is extending from language semantics into performance-critical paths, directly serving convolution, GEMM, and mHC scenarios.

5.3 The LinearAttention family enters a consolidation phase

From the DeltaNet inference operator (09-23) to GLA inference + SM90 DeltaNet decoding (09-24), and then to four “remove legacy paths / rename” drafts the same day, the linear attention family’s iteration pattern is fully on display: expansion (training side → inference side) → platformization (SM90/H200) → consolidation (deleting old paths to converge). The migration of paged GQA to a unified operator (#2202) is part of the same mainline on the attention side.

5.4 Gaps and risk points

First, review bandwidth for large items: the main repo’s TileIR backend execution PR and major changes such as magic division saw no new merges within the window, and combined with the 23-day release stall, the pattern of “small PRs merge fast, large PRs review slowly” persists. Second, although the Ascend documentation verification queue is advancing in batches, the main body has remained open since being suspended in mid-August, with no explanation of the acceptance and merge path. Third, the TileRT release page lags behind release PRs, creating a short-term blind spot for version tracking on this topic. Fourth, Moore Threads has been silent for a second consecutive week, with uneven activity levels among the four domestic adaptation vendors.


Appendix: Sources and Verification Notes

Source Verification Result
tile-ai organization (29 repos) 11 repos had pushes in the window: tilelang, TileRT, TileOPs, TileOPs-nightly, TileFoundry, tilelang-ascend, tilelang-metax, tilelang-hygon, tilelang-mlir-ascend, tilelang.github.io, TileOPs.github.io
Main repo tilelang 4 merges (#3205, #3273, #3269, #3278); new issues #3274, #3275, #3276, #3279; bug reports #3270, #3024 closed
TileRT v0.1.5.post3, v0.1.6 release PRs merged; releases page and tags still at v0.1.5.post2; #66 closed
TileOPs 10 merges (#2174, #2179, #2181, #2183, #2185, #2186, #2175, #2193, #2194, #2206); 8 new drafts; #2195, #2205, #2180 closed without merge
TileOPs-nightly 1 snapshot (c796eb479b, corresponding to 85455e46c2, i.e., the #2206 merge point): correctness 1140 items, zero failures, zero errors; benchmarks 1081 items, zero failures, zero errors
TileFoundry #184 merged; #185 draft updated
tilelang-ascend 3 merges (#1741, #1826, #1829); documentation verification queue updated in a batch of about 10
tilelang-mlir-ascend Branch cleanup (4 branches deleted, corresponding PRs closed); no commits on main branch
tilelang-metax #158 upstream sync merged; #159 newly opened
tilelang-hygon #12 merged
tilelang-musa No pushes (last push 09-17)
deepseek-ai/TileKernels, QwenLM/FlashQLA No pushes
Community repos SoftHier extension updated; Cookbook main branch has no new content (pushes from chart automation); personal experiment repos have no substantive content
Google News / Hacker News / arXiv Zero hits for Chinese and English queries; HN irrelevant entries; arXiv no new preprints (latest remains the 07-24 TileSight paper)

Complete Source List