Reporting period: Monday, September 21, 2026 to Friday, September 25, 2026 — 5 working days, 5 issues of daily briefing material (corresponding to ISO week 39) Source material: 5 issues of this publication’s daily TileLang updates (09-21, 09-22, 09-23, 09-24, 09-25); supplemented by a 168-hour window verification across 29 repositories in the tile-ai organization and a review of release metadata for each repository (see appendix)


1. Weekly Highlights

  • Two release PRs merged into TileRT in one day; v0.1.6 brings GLM-5.2/5.3 to AMD Instinct MI350X/MI355X (09-24): At 20:22 and 23:33 Beijing time on 09-24, TileRT merged two release PRs back-to-back: v0.1.5.post3 (#63) and v0.1.6 (#68). The v0.1.6 release notes cite added GLM-5.2/5.3 support on AMD Instinct MI350X/MI355X (ROCm path) along with prefill-decode (PD) disaggregation updates; on the same day, the PD-side KV injection copy cross-device pair overlap performance ticket #66 was closed, with its content folded into this release narrative. Caveat: as of the end of this cycle, TileRT’s GitHub Releases page and tags remain at v0.1.5.post2 (2026-08-06) — the releases page lags behind the release PRs, and the publicly available version has not yet been updated.

  • Exact FP4-to-FP8 conversion fused into the main repo; DeepSeek V4.1 expert GEMM speeds up by up to 5.38x (09-22): Main-repo tilelang #3204 merged on 09-22 at 19:36 (opened 09-11, 11 days of refinement, 161 lines added). The change fuses the default projection chain — “E2M1 (FP4) weights converted to E4M3 (FP8) via an FP32 intermediate” — into two __byte_perm instructions per four elements, covering scalar and five vector widths (2, 4, 8, 16, 32), entirely in registers with no new memory allocation. On H100, speedups across twelve configurations range from 2.97x to 5.38x (e.g., 1×2304×5120 drops from 547.94 µs to 108.65 µs; 2048×5120×2304 drops from 1760.22 µs to 591.92 µs), with register usage rising from 128 to 156 and no spills. Validation covers source-generated tests, per-word encoding comparison for each packed word, and current-head CI for CUDA/ROCm/Metal/CuTeDSL; caveat: Blackwell execution and full-model serving remain unverified.

  • TileSight officially open-sourced: tile-centric analytical performance model lands from paper to official repo (09-23): At 22:54 on 09-23, the newly added TileSight repo under the tile-ai org completed its first commit (177 files, ~56.7k lines, MIT license), bringing the org’s total repo count from 28 to 29; the “open-source upon publication” promise at the end of the July 24 arXiv paper (2607.22432) has now been fulfilled. The repo covers 20-plus hardware profiles (NVIDIA, AMD Instinct, and Intel Arc discrete GPUs), a multi-level cache reuse-distance model, pipeline overlap analysis, and CUDA/HIP microbenchmark probes; per the paper: single-GPU kernel latency combined MAPE of 12.35% across four generations of platforms from A100 to B6000, rising to weighted errors of 16.18% and 13.52% for fused distributed kernels and vLLM end-to-end serving, respectively, when scaled to a 32-GPU cluster.

  • TileOPs sees 21 merges in a week: governance consolidation and peak performance season in parallel (09-21 ~ 09-24): TileOPs totaled roughly 21 merges this cycle. On the governance track, the large roofline formula PR #2158 made the manifest formula the single source of execution and corrected 15 classes of formula errors (covering 175 implemented operators), followed by #2167 removing the last bypass and #2175 turning synthesis failures into named reports — “trust engineering for metrics” consolidated; on the performance track, seven merges landed in a single day on 09-23 (across seven subsystems: GEMM, Engram, attention, MoE, Mamba, mHC, and linear attention), among which the DeltaNet inference operator (#2163) achieved up to 3.7x speedup over FLA on H200; channel governance and interface contraction advanced in tandem (see Section 4).

  • GLM-5.3 k-pool chain content merged into mainline: full example set and 4 AMD test assets included (evening of 09-20): At 17:06 on 09-20, tilelang #3254 merged in compressed form (11 files, +2520/-4), bringing the previous cycle’s k-pool series into the mainline: four runnable examples under examples/kpool/ (in-pool softmax pooling, BF16 round-trip with Hadamard normalization, FP8 quantization, paged cache readback with 32-head MQA weighted scores) plus four ROCm tests; at the language layer, FP8 dtype selection in tilelang/language/fp8.py now supports per-device specification. Validation: exact-head CI all green, with ROCm 7.2 reporting 2433 passing tests on MI300X. Caveat: four stacked PRs — #3250, #3251, #3253, #3255 — remain open and unclosed (their content is already in the mainline), with no closing action seen by the end of this cycle.

  • Main-repo fix depth: 6 merges in a single day, Fuzzer cleanup, AllReduce silent-error interception, and dual-track NaN semantics closure (09-21 ~ 09-24): The main repo saw 14 merges over the week, forming three layers of depth — compiler frontend and boundaries (two Fuzzer ICE fixes 3206/3207, with the original tickets opened by the same reporter on 08-17 the same day being precisely cleared), runtime wrapper layer (NVRTC fix #3260), and code generation and scheduling (AllReduce stride validation #3266 intercepting “runs but silently computes wrong” errors). On 09-24, NaN propagation semantics closed on two tracks (#3273 compile-time error + #3205 clamp lowering fix, with both defect tickets closed), and the same day saw the SIMT im2col and async staging refactor chain take shape (#3274 → #3275 → #3279).

  • Ascend’s one-week “merge—revert—restore” loop; A5 real-hardware defect made public (09-20 ~ 09-24): After 11 merges in a single day on 09-20 (the largest, #1753, being a 72-file, +10117/-3618 FP32 row reduction and synchronization fix), 09-22 saw three wholesale reverts (#1818 precisely reverting #1753), and 09-24 closed out with daily tests passing 2448/2448 — the full chain of “merge, validation failure, revert, restore to all-green” completed within two days. Meanwhile, three defect tickets were opened on 09-23, among which #1830 points out that on real Ascend 950 (dav-3510), the device path is compiled with a 910B-generation label, throwing a device exception at startup, and silently computing wrong results after changing the label — exposing a workaround blind spot in the validation matrix where “real hardware covers the 910B generation, simulation covers the 950 generation.”

  • Multi-backend and ecosystem: Hygon example support merged, MetaX daily updates, Sunrise release, Moore Threads silent (09-21 ~ 09-24): Hygon #11 example support merged on 09-23 (37 files, +1104/-130, with GEMM switched to HCU matrix-core built-in instructions), and #12 merged on 09-24 (FP32 MMAC K sharing and ds_read fragment packing); MetaX completed upstream main synchronization (#158) and opened a sub-warp guard fix (#159); Sunrise released 0.1.14+sunrise.1.1.0 on 09-21 (CI system expansion and sparse mask fix); the Moore Threads repo was silent all week (last push 09-17), with the mechanistic gap in its external backend integration being absorbed and closed by TileOPs’ grouped queue governance.


II. Core Project Progress (Main Repo tilelang)

Versions and Releases

  • The main repo’s latest tag remains v0.1.14 (released 09-02); by the end of this period it has gone 23 days without a new release, approaching the previous ~30-day release interval; across the entire organization, only Sunrise had a new release this period (09-21), and TileRT’s two release PRs have been merged but the releases page has not been updated (see Section IV).
  • A downstream hard constraint has emerged: since #2172, TileOPs has imposed a >= 0.1.14 version floor on tilelang (adding the floor in pyproject.toml and enforcing it in the install script), and the release cadence and downstream adoption windows are beginning to constrain each other’s timing.
  • Adaptation-line tags remained unchanged all week: Ascend v0.1.2.000 (09-09), MLIR Ascend v0.1.2.020.sglang.poc (09-09), Moore Threads v0.1.14+musa.1 (09-11), MetaX v0.1.14 (09-17).

Language Layer and Frontend

  • Pythonic frontend (#3230, merged 09-21): compile-time loops now accept Python iterables and comprehensions, 5 files +1084/-24 (including 766 lines of new tests), with explicit rejections for misuse of PrimExpr, Buffer, etc.; two long-standing feature requests were closed at the same time (ForFrame not iterable, T.unroll expression-list indexing).
  • Loop target binding fix (#3232, merged 09-23, +196/-14): establishes explicit bindings for loop targets (including tuple targets), avoiding writing induction variables into existing bindings; bug report #3231 was closed upon merge.
  • Copy width clamping (#3246, merged 09-23): coalesced_width for T.copy and T.async_copy changed from “fatal log if not evenly divisible” to clamping to the achievable vector width — warning on downgrade, rejection of non-positive values; bug report #3007 closed.
  • Dual-track NaN propagation semantics (merged 09-24): #3273 adds a compile-time error for unsupported reduction NaN-propagation dtypes (the diagnostic gives the op name, supported output types, and argument dtype, with validation hooked into the native ReduceOp constructor); #3205 correctly lowers T.clamp NaN propagation to the device template. The former stems from report #3270 and the latter from fuzzing bug #3024; both were closed upon merge.
  • Cleanup items: #3262 removes the warning for rebinding immutable variables (merged 09-21, net deletion of 14 lines).

Compiler, Codegen, and Runtime

  • Fuzzer backlog cleared with two items (merged 09-22): #3206 lowers plain assert in kernel bodies to legal device-side checks (previously host-side error-reporting calls were written into __global__ kernels and nvcc rejected compilation outright, original issue #3019); #3207 fixes a call-time crash when “dynamic-shape outputs precede shape-providing inputs” (original issue #3017).
  • AllReduce silent-error interception (#3266, merged 09-21): the stride check for XOR butterfly reduction is folded into CheckAllReduceWidth, a static assertion is added to the template, and the planner falls back to a wide scheme for unsupported forms — intercepting a class of errors that “seem to run, but overflow shared-memory exchange exactly at 96/192 threads.”
  • NVRTC line: #3260 fixes a warp reduction compilation issue (merged 09-22, about 14.5 hours from filing to fix); subsequently #3265 (sinking launch-invariant integer arithmetic to the host side) and #3263 (wrapper layer failing to pass host scalars) are under review.
  • Dynamic-shape division: feature request #3261 and implementation #3267 (tl.LowerMagicDiv magic division, precomputed from the divisor’s runtime value) continue to be updated and await merge.
  • SIMT im2col and async staging, a three-step progression (09-24): performance bug #3274 (lowering happens too late, missing scheduling and vectorization opportunities) → three minutes later patch #3275 (hoisting loads ahead of scheduling to expose them) → early 09-25 refactor draft #3279 (“unified async data staging and transfer-aware lowering,” experimental: separating logical access from physical lowering, shared transfer analysis, pipeline wait planning, and vectorization widening). On the same day #3278 (fix for async copy lowering under partitioned layouts) was merged; #3276 (preserving FP8 vector copy packing) is under review.
  • Engineering support: #3268 (CMake preserving SDK paths given by backend environment variables) is newly opened and awaiting review.
  • CUDA Tile IR execution backend (#3247): within a week it grew from 8 commits to 31 commits, +36506/-112, 139 files, covering JIT, caching, autotuning, and a DeepSeek V4 sparse attention tuning kernel; still under review — the “seasoned but not stamped” state continues.

Performance

  • FP4→FP8 conversion fusion (#3204, see Top News): 2.97x to 5.38x speedup across twelve H100 configurations; validation in four layers (source generation, literal comparison, bitwise-identical output, cross-platform CI).
  • SM120 block-scaling line (09-21): only 21 minutes from feature request #3256 to patch #3257 (fragment A, odd warp grids, and compact scaling), with the roadmap driven by SageAttention3’s actual needs; #3257 awaits review.
  • Cadence observation: most small fixes during the week were “filed the same day, cleared the same week” (two Fuzzer items, one NaN item patched the same day), while the center of gravity for substantive increments remains in the review queue.

Review Queue Observations

  • Large items piling up: TileIR backend (#3247, 31 commits), magic division (#3267), launch invariants (#3265), four CuTeDSL and RNG fixes (#3241 through #3244 refreshed as a batch), GEMM dtype combination interception (#3245), SM120 patch (#3257) — the pattern of “small fixes stamped, large changes accumulating material” persisted throughout the week.
  • Four leftover items from the k-pool stack (#3250, #3251, #3253, #3255): their content has already entered the mainline via #3254, but the issues remain open without resolution and have seen no activity for several consecutive days; they need to be closed out via rebase or closure.
  • Queue-release linkage: after TileOPs imposed the >= 0.1.14 floor, the timing relationship between the main repo’s release window (already 23 days full) and downstream adoption has become a cross-repository variable.

III. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads, etc.)

Ascend: A Three-Act Week of Merge Day, Revert Day, and Recovery Day

  • 09-20 (Dense merge day): 11 merges in a single day (09:36 to 16:29) — the largest being #1753 (72 files, +10117/-3618: compile-time shape-known FP32 row-wise max/min/sum implementation and fixes for two classes of auto-sync defects); the correctness fix group #1804 (row slice stride, fixing uint16 truncation data corruption on 910B2), #1809 (block_sparse_mqa_attn async read/write); the operator and docs group #1621 (mhc_pre example) and a docs series of 6.
  • 09-21 (The tail emerges): #1633 (mhc_bwd, Sinkhorn implicit conjugate gradient example), under review across two weeks, merged at 17:16, with 3.1% speedup on 2048/4096 sequences and reference CG validation folded into the PASSED criteria; two new fixes #1821 and #1822 opened; daily test failure ticket #1817 and two follow-up PRs (#1815, #1816) appeared in succession.
  • 09-22 (Revert day): Three reverts merged in succession — #1818 precisely withdrew #1753 (72 files, +3618/-10117, exact symmetry), #1819 and #1820 withdrew two docs enhancements; daily tests #1823 recovered with 2448/2448 all passing (against a pre-merge-wave baseline of 1936, the increment of 512 items did not fall back).
  • 09-23 (Defect trio): Two mHC example reverts (#1832, #1833, both in-repo for under three days); three new defect tickets — #1830 (the A5 device path compiles for a real Ascend 950 (dav-3510) using the fixed label dav-2201, causing device exception 507015 on kernel launch and silent miscalculation after changing the label; the README validation matrix happens to bypass this label in both directions), #1824 / #1825 (two boundaries of auto-sync: incorrect merging of sync state in conditional branches, and mixed configurations that should be rejected at compile time); #1834 opened a fix plan for a scalar GM store race (reproduction script mismatch count dropped from 12,967,510 to 3,496,136, reaching zero once the opt-in rewrite path is enabled, with 113 existing suites passing).
  • 09-24 (Recovery advances): #1741 sparse FlashAttention example merged (created 09-04, reviewed across two weeks), #1829 vector mask and FP32 row reduction optimization, #1826 restored Select/Transpose tests and in-place validation; daily tests #1835 with 2448/2448 all passing; the API docs verification queue refreshed in a batch of about 10 (the bulk suspended since mid-August); #1828 (operator migration) and #1834 pending review.
  • Assessment: The validation cost of large merges was paid head-on — from merge, to validation failure, to revert, to recovery to all-green in under two days; the A5 real-hardware check signals that cross-generation hardware adaptation is moving from “it compiles” into the stage of “daring to calibrate on real hardware.”

MLIR Ascend (tilelang-mlir-ascend): Large Mamba Merge and A5 Branch Cleanup

  • 09-20: #187 (TileOPs benchmark and reporting infrastructure, ~+5.4k lines) and #188 (optimized Mamba operators and per-agent model selection) both landed.
  • 09-22: #191 Mamba-2 SSD chunked scan NPU expert kernel optimization merged (42 files, +4204/-1568, about 3.5 hours from open to merge), with tuning conclusions distilled into the pattern library and case records.
  • 09-23: #178 NPU launcher memory leak and stream sync fix merged (rtMalloc changed to stream-aware allocation, unified launch path).
  • 09-24: Four A5 development branches deleted in a batch (corresponding sync-class PRs closed); #130 (A5 test sync) still open; #189, #190 pending review.

Hygon: Example support merged, matrix core data layout fix landed

  • Example support #11 merged 09-23 (37 files, +1104/-130: adapted autotune and LDS configuration, resolved kernel layout conflicts, switched GEMM to HCU matrix core built-in instructions, added regression coverage) — the Hygon-side phase of “getting the TileLang example suite running on HCU” is now merged.
  • Two advances on the feature branch: FP32 MMAC K sharing and ds_read packing (09-21) → expanded to a five-file version and cleaned up the tvm_ffi FP8 export workaround (09-22).
  • #12 merged on 09-24: shared FP32 MMAC K dimension and packed ds_read_format fragment.

MetaX: Upstream sync and sub-warp guard

  • 09-21: main fast-forwarded to the latest upstream eab74a4a (syncing nearly three weeks of upstream increments).
  • 09-24: #158 (upstream main synced into dev, including conflict resolution) merged; #159 (T.gemm_sp sub-warp block size guard) opened pending review.

Moore Threads (MUSA): Silent all week, governance shifted to the operator library queue

  • The tilelang-musa repository’s most recent push was 09-17, with no new activity this cycle; the mechanistic gaps in MUSA external backend integration (three issues #2164 to #2166: CUDA-only input validation, zero-input rejection, CUDA device query before dispatch) are being absorbed and closed by TileOPs as a grouped governance effort (see Section IV), with all three tickets continuously updated within the window.

Other: Sunrise releases 0.1.14+sunrise.1.1.0

  • A new version tag was released on 09-21 (candidate branch pushed at 11:06 → release PR #6 merged at 15:33 → tagged): the content is mainly engineering-focused — fixed sparse mask execution, stopped CI after device recovery failure; the CI system was greatly expanded (run/install/junit to JSONL scripts); added compilation speed benchmarks and compile-only, lower-trace tool documentation; z3 prover switch and integer analysis updates.

IV. Ecosystem and Adopters

TileOPs: A Week of 21 Merges — Governance Consolidation, Performance Season Peak, and Channel Governance

  • Governance consolidation: #2158 (merged 09-21) makes the roofline manifest formula the single source of execution (one source, one evaluation surface), fixes 15 classes of formula errors, and has the structural oracle recompute 175 operators per PR; #2167 (merged 09-22) removes the last bypass that “silently defers to entries that define an evaluation method” and adds explicit errors for two classes of misuse; #2175 (merged 09-24) turns formula synthesis failures into named reports (full determination possible without a torch environment); #2161 (issue) records that the manifest’s shape_rules helpers are never evaluated and 234 rules are silently skipped, with the fix pending a decision.
  • Performance season peak: Seven merges in a single day on 09-23 — W4A16 prepacked weight order (#2168), Engram decode split (#2173, 8.2x at batch 4), varlen GQA unified operator migration (#2160), MoE excess routing grouped by 16 rows (#2169, rising to the 0.99 tier relative to vLLM), Mamba DaCumsum (#2176), mHC prefill projection split along K (#2177), DeltaNet inference operator (#2163); 09-24 added the GLA inference operator (#2174) and SM90 DeltaNet decode (#2194), completing multiple points across the linear attention family — “training side → inference side → platform decode → performance model calibration”; W4A16 long-K streaming (#2185) landed; in the queue, the GEMM ping-pong main loop with 176 tiles (#2172), the varlen GQA pre-planning trio (#2178 / #2180 / #2184), and the FFT kernel rewrite (#2187) continue to advance.
  • Channel governance (external backends): #2179 (merged 09-24, +797/-256) removes roughly 38 device-kind checks at the operator layer, shifting device declaration to the kernel at the tensor level, and adds a lint hook to prevent regression; #2181, #2183 (test default target switch, operator behavior tests running on all backends), #2186 (DeltaNet prefill hang and GEMM roofline anomaly fix), and #2193 (input memory retention fix) merged in the same window; #2182 (issue) records the BuildKernel protocol gap, pending extension.
  • Interface contraction: Four legacy-path cleanup drafts — #2196 (remove unowned GDN kernels), #2197 (remove legacy varlen GQA), #2200 (remove legacy GLA inference path), #2201 (GLA operator rename), plus #2198 / #2199 / #2202 / #2203; #2195, #2205, #2180 closed unmerged (content split and advanced).
  • Read: The ledger left by governance is turning into credit assets for performance work — every performance PR carries its own comparison baseline (torch.compile, cuBLASLt, vLLM, FLA) and bitwise-consistency declaration, and can be independently verified by the manifest and benchmarks.

TileOPs Nightly Pipeline: Benchmarks from 1040 to 1081, Correctness Steady Above 1140

  • Six snapshots visible this cycle (corresponding to merge points from #1996 to #2206): benchmark cases started at 1040, rising through 1046 and 1050 to 1081; correctness cases ranged between 1118 and 1141, with zero failures except for one snapshot that omitted a results file (later republished, a release oversight).
  • The 09-24 snapshot saw a file-level timeout — the DeltaNet long-prefill configuration had no test start within 900 seconds (the cost between a new operator’s “functional merge and benchmark affordability” remains to be calibrated); after the corresponding fix #2186 merged the same day, the 09-25 snapshot returned to zero failures and zero errors.
  • Environment consistent all week: H200, CUDA 13.2, driver 595.71.05, SM locked frequency, torch 2.13 — the reproducibility elements of the snapshots remain intact.

TileRT: Two Release PRs Merged in One Day (see News Brief)

  • Patch line and minor-version line advance in parallel: v0.1.5.post3 (created 09-07) is the third patch on the v0.1.5 line; v0.1.6 (created and merged the same day, merged by a project member) folds in ROCm (AMD Instinct MI350X/MI355X) support for GLM-5.2/5.3 and PD-disaggregation updates; compared with a cadence of just one release in August, the release pipeline has clearly intensified.
  • The release page and tags lag behind the release PRs, leaving an information gap in public channels; continue to verify next issue.

TileFoundry: A Dense Week of Compile-Time Analysis

  • Two merges on 09-21: #170 (loop-aware buffer lifetime solving, structured SSA intervals with reuse constraints, replacing logical lifetime summation), #174 (CuTe-style Swizzle functor with layout algebra specialization, CUDA codegen output as-is).
  • Two merges on 09-22: #175 (analysis supports loop start points of dependency units), #176 (memory metadata and traffic normalization, 263 lines changed in a single analysis-spec file).
  • Large merge on 09-24: #179 (memory footprint and reuse windows, +3207/-539, 44 files) — derives temporal and spatial reuse axes from access relations, producing recomputable capacity conclusions of the form “b holds=176.00MB fits=no (vs. 47.68MB L2)”; #184 (isl recursive traversal fix) merged the same day.
  • #185 (shard layout restructuring draft) follows up; #168, #171, #172, #173, #178 closed.

TileSight: Tile-Centric Performance Model Officially Open-Sourced (see News Brief)

  • Division of labor: TileSight predicts cross-layer costs during the modeling phase, complementing TileFoundry’s compile-time IR analysis; its hardware profiles cover AMD and Intel discrete GPUs, naturally serving the multi-backend narrative. The three-stage rollout (07-24 paper → 09-18 community docs repo → 09-23 official code repo) elevates performance explainability from a personal thread to an organizational project.

Adopters: TileKernels and FlashQLA Saw No Pushes All Week

  • TileKernels’ most recent push remains 04-23; FlashQLA’s is 09-18 (three SM100/SM120 merges belonging to the previous cycle). Adopter code has not moved, but their inference workloads continue to drive upstream optimization — this week’s FP4→FP8 fusion merge specifically targets the DeepSeek V4.1 expert GEMM.

V. Community, Tutorials, and Events

Docs Sites and Routine Updates

  • The main repo’s docs site saw multiple automated “Update docs” commits (09-20, 09-21, 09-23, 09-24); the TileOPs docs site deployed routinely (09-20 to 09-22, 09-24); all are automated regenerations, with no substantive site content changes observed.

Community Projects: TPU Backend, Bootcamp, Teaching Repos, and SoftHier

  • Community TPU extension repo (Sophgo BM1690): Three commits on 09-22 — multi-core tiled scan S3/P6 validation, P10 performance matrix, host-side source verification fix; the repo retains TileLang’s Python frontend and adds TPU lowering, code generation, and a JIT runtime, making it the first publicly available TPU backend path in the TileLang community.
  • MetaX C500 bootcamp materials repo: Two commits on 09-22 — TileLang and Nine-Teeth comparison exercises for Lecture 2 “Vector Addition,” and supplementary operational proxy configuration for the remote instance workflow; the teaching sequence advances as “environment → addition comparison → Softmax and GEMM → AI agent assistance → Llama operators.”
  • Teaching and localization repos: tilelang-tutorials added init_core group examples (PTO frontend in pure Cube, pure Vector, and hybrid forms, 09-21, author from Huawei); tilelang-docs-l10n completed a full-language README translation sync.
  • SoftHier (ETH Zurich): Extends TileLang to a tiled multi-PE RISC-V many-core accelerator (GVSoC modeling, a semester project supervised by Luca Benini’s lab) — kernel authors declare only “what to compute and where the data is,” while the compiler derives cluster placement, collective communication routing, and synchronization; for a 4×4 cluster SUMMA GEMM with BM=BN=BK=128 (N=7168, K=2048, matching the DeepSeek-V3 routed expert down-projection dimensions), compiler-generated code runs in 1.01 ms at M=512 versus 0.96 ms for hand-written C (a 5.14% gap).

Media and Academia: Zero Hits All Week

  • Multiple Google News queries in Chinese and English returned zero hits across five consecutive windows; no topic-relevant Hacker News entries; no new arXiv preprints (the most recent remains the 07-24 TileSight performance model paper — its code was open-sourced this cycle). The media side has been quiet for several consecutive weeks, with primary content driven by GitHub organization activity.

VI. Trend Observations

  • TileRT’s release cadence intensifies, and AMD platform weight rises: The patch line and minor-version line advance on the same day, placing the ROCm (AMD Instinct MI350X/MI355X) plus GLM-5.2/5.3 combination in the most prominent position of the release notes — for a “ultra-low-latency inference runtime” positioning, model coverage is beginning to ramp up in parallel with the NVIDIA line; the release page and tag lag is a short-term blind spot that needs to be closed out next issue.
  • Optimization focus moves into “the ledger of every layer’s instructions”: A 161-line conversion fusion buys up to 5.38x end-to-end speedup, but only because it sits on a repeatedly read weight path; TileOPs’ second performance wave (scheduling planning, register splitting, grid utilization, stream-K) likewise interrogates “how many times the same block of data is read and at which cache level it flows” — gains continue to migrate toward data movement and format conversion.
  • “Silent errors moved forward” extends to the scheduling and staging layers: The dual-track NaN propagation, AllReduce static assertions, and Fuzzer clearing continue the principle of semantic validation; SIMT im2col early exposure, separation of logical access from physical lowering, and unified asynchronous staging push “declare first, derive later, implicit to explicit” into the performance-critical path, directly serving convolution, GEMM, and mHC scenarios.
  • The linear attention family enters a consolidation phase: Within two days, multiple points landed — DeltaNet inference → GLA inference and SM90 decode — immediately followed by four drafts to “delete old, rename, converge interfaces”; such cleanups typically precede interface freezes or version actions — watch whether TileOPs shows a release cadence afterward.
  • Multi-backend engineering debt becomes explicit, governance advances in groups: The final stretch of external backends from “can register” to “can schedule” begins to be paved (three blocking items down to 38 checks removed plus lint to prevent regression, advanced through bilateral coordination); the Ascend “merge—revert—restore” loop completes within two days, with A5 real-hardware validation pushing cross-generation adaptation into deep water; Hygon and MetaX maintain a daily-update rhythm, Moore Threads remains silent, and the four show diverging momentum.
  • Performance explainability elevated to an organizational asset: TileSight goes from paper to official repo, and TileFoundry pushes compile-time analysis to recomputable capacity conclusions — beyond operator libraries and runtimes, the maturity of analysis tooling is becoming a differentiating asset for TileLang and provides a modeling foundation for the multi-backend narrative.
  • Watch items for next issue: The main repo’s release window (now 23 days overdue) and whether the TileIR large item lands; the rework form of the reverted Ascend #1753 and the scheduling of the A5 output correctness issue; TileRT release page and tag follow-up; whether TileOPs cleanup drafts merging triggers a version action; the four-item wrap-up of the k-pool stack.

Appendix: Sources and Verification Notes

  • Sources: Five issues of this publication’s daily TileLang intelligence brief (09-21 to 09-25), with a collection window from 09-20 07:00 to 09-25 07:00 (Beijing time); the 09-21 issue covers 24 hours starting the morning of 09-20, including items such as the k-pool chain merge on the evening of 09-20. Activity over the weekend following this period (09-26 to 09-27) will be folded into next Monday’s brief.
  • Supplementary verification (168-hour window, as of the preparation of this report): 13 of the 29 repositories in the tile-ai organization had pushes within the window (tilelang, TileRT, TileOPs, TileOPs-nightly, TileFoundry, TileSight, tilelang-ascend, tilelang-metax, tilelang-hygon, tilelang-mlir-ascend, tilelang-sunrise, tilelang.github.io, TileOPs.github.io); release and tag metadata were re-checked repository by repository; adopters and related repositories were cross-checked for their most recent push times (TileKernels 04-23, FlashQLA 09-18, tilelang-musa 09-17).

Source Verification Table

Source Verification Result
tile-ai organization (29 repositories) 13 repositories had pushes within the window; new repository TileSight completed its first commit (177 files, ~56.7k lines)
Main repository tilelang 14 merges over the week (09-20 evening to 09-24); large review items (TileIR backend 31 commits, magic division, RNG queue, etc.) continue to pile up
TileRT v0.1.5.post3 and v0.1.6 release PRs merged; releases page and tags still at v0.1.5.post2
TileOPs 21 merges over the week, plus multiple opens and closes; nightly snapshot correctness 1118 to 1141, benchmarks 1040 to 1081 (one file-level timeout fixed)
TileFoundry 6 merges (#170, #174, #175, #176, #179, #184); #185 draft in progress
Ascend Three-act closed loop: 11 merges (09-20) → 5 reverts (09-22, 09-23) → restored 2448/2448 (09-24); tickets opened for A5 defect and sync boundary defect
MLIR Ascend Multiple landings from #187 to #191 and A5 branch cleanup; #130 still open, #189/#190 pending review
Hygon / MetaX Hygon #11, #12 merged; MetaX #158 merged, #159 opened
Moore Threads / Sunrise No pushes for Moore Threads (most recent 09-17); Sunrise released 0.1.14+sunrise.1.1.0
Adopters (TileKernels, FlashQLA) No pushes all week (most recent 04-23 / 09-18)
Community repositories TPU extension, bootcamp materials, teaching repository, and SoftHier each had updates
Media and academic channels Zero hits across Google News / Hacker News / arXiv all week
  • Scope and boundaries: All performance data are self-reported in merge notes or PR descriptions (measured on H100, H200); this publication has no corresponding hardware and has not performed independent reproduction; nightly benchmarks are single pipeline records, not head-to-head evaluations; release status is based on re-verification of each repository’s release metadata (as of the preparation of this report), and TileRT’s releases page lagging behind its release PRs reflects the repository-side release process cadence and does not mean the release content is invalid.
  • Suggested next period: 2026-09-28 ~ 2026-10-02.

Key References

Data sources: Daily TileLang intelligence brief (five issues, 2026-09-21 to 2026-09-25) and hands-on verification during this period; reporting period is September 21 to September 25, 2026.