Research window: Past 24 hours (2026-09-29 07:00 ~ 2026-09-30 07:00, Beijing time; the previous report was on 09-29, so the window is seamlessly connected with no overlap). Sources: GitHub (push verification across 29 repositories in the tile-ai organization; 7 repositories had pushes within the window: tilelang, TileOPs, TileOPs-nightly, tilelang-ascend, tilelang-hygon, TileFoundry, TileOPs.github.io; the main repo had 2 merges and 7 new opens, TileOPs had 17 merges, the Ascend repo had 2 fixes, Hygon and TileFoundry each had 1 merge, and the nightly snapshot had 1 build with zero failures), Google News RSS multi-language queries (via proxy, zero hits), Hacker News, arXiv, community repository verification


In This Issue

  • Today’s Highlight: Ascend 950 backend publicly released — 134-commit native backend PR opened, version number bumped to 0.1.15 in tandem (09-30)
  • I. Core Project Progress
    • 1.1 Main repo: Simplify unused binding eliminator merged — read-after-write, side effects, and metadata references all covered by regression safeguards (09-29)
    • 1.2 Main repo: Windows wheel build restores ccache hits — clang-cl flags split into separate writes (09-29)
    • 1.3 Main repo queue: Seven new opens and an atomic operation fix line, open PRs total 116 (09-29/09-30)
    • 1.4 TileOPs: Four INT8 quantization and dequantization kernels merged — bit-exact alignment with torch reference (09-29/09-30)
    • 1.5 TileOPs: INT4 per-group quantization kernel merged, two block quantization PRs queued (09-30)
    • 1.6 TileOPs: Sampling operator layer merged — six kernels queued last issue now formally landed (09-29)
    • 1.7 TileOPs: Four attention PRs — FP8 expression sinking, varlen outbound to shared memory, SM89 head dim 512 fix (09-29/09-30)
    • 1.8 TileOPs: Two PRs on kernel interfaces and shared memory governance (09-29)
    • 1.9 TileOPs: Three PRs on CI and engineering standards — eight-concurrency smoke tests, absolute imports, single-pass manifest validation (09-29/09-30)
    • 1.10 TileOPs: Two element-wise speedups — division fallback once per thread, f32 dual 16-byte loads (09-29/09-30)
    • 1.11 Nightly snapshot: 1030 correctness items, 1355 benchmark items, zero failures and zero errors (early 09-30)
  • II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)
    • 2.1 Ascend: Further fixes within the 950 release branch window — broadcast vectorization and scalar reduction restored (09-30)
    • 2.2 Ascend: Two fixes in the ascend repo — one-line TVM patch reveals the true cause of regression (09-29)
    • 2.3 Hygon: Async copy and cache swizzle lowering improvements merged (09-29)
    • 2.4 MetaX / Moore Threads / MLIR Ascend: Official repos silent (09-30 check)
  • III. Ecosystem and Adopters
    • 3.1 TileFoundry: AtomSched phase four merged — HIR lowering to TIR and scheduling CLI (09-29)
    • 3.2 Docs site: Quantization, sampling, and shared-expert MLP operators enter API reference (09-30)
    • 3.3 Adopter check: TileKernels and FlashQLA have no pushes within the window (09-30)
    • 3.4 Release cadence: 0.1.15 bump PR opened; TileRT and TileFoundry release pages stalled (09-30)
  • IV. Community, Tutorials, and Events
    • 4.1 Community repo check: No new commits within the window (09-30)
    • 4.2 Media and academia: Zero Google News hits, no new arXiv preprints (09-30)
  • V. Trend Observations
    • 5.1 From contract to kernel: Quantization family delivered within 48 hours
    • 5.2 Ascend 950: Multi-backend narrative moves from “adaptation repo” to “main repo native”
    • 5.3 Engineering capacity: CI refactoring under a 21% increase in test volume
    • 5.4 Gaps and risk points
  • Appendix: Sources and verification notes

Today’s Highlight: Ascend 950 Backend Publicly Released — 134-commit Native Backend PR Opened, Version Number Bumped to 0.1.15 in Sync (09-30)

Date: 2026-09-30 Source: Ascend 950 Backend PR (#3308) / Version Bump PR (#3309) / Release branch ascend-950-0930

In the early hours of 09-30, the TileLang main repo opened two landmark PRs in quick succession: the Ascend 950 native backend, labeled “Public Release 9/30”, and a version bump to 0.1.15. The former, weighing in at 134 commits, 379 files, and +79956/-733, promotes the Ascend line from a standalone adaptation repo to a first-class backend in the main repo:

  • End-to-end path: Adds an Ascend language dialect covering kernel launch, memory allocation, data movement, and compute operations, including hardware-specialized lowering, auto-scheduling and synchronization, device code generation, Bisheng compilation, and NPU runtime loading; users can develop Ascend 950 kernels directly with target="ascend".
  • Programming model: Combines Cube (AIC) T.gemm with Vector (AIV) computation within a single kernel; T.SimdVF / T.SimtVF regions mix SIMD and SIMT programming; explicit UB / L1 / L0 allocation; T.copy / T.dual_copy for tiled movement and cross-core transfer; includes MXFP8 / MXFP4 block-scaled GEMM low-precision paths.
  • Compiler and scheduling: Automatic inference and storage normalization for Ascend fractal layouts; AutoSchedule performs dependency- and latency-aware Cube/Vector scheduling, pipelining, and multi-buffering; schedule-aware intra-core and cross-core synchronization insertion, with redundant synchronization elimination and automatic flag allocation and reuse.
  • Build integration: The USE_ASCEND build switch, Bisheng toolchain, runtime loading, and kernel launch all land in the main repo; integrated into the multi-backend architecture via existing backend interfaces, reusing shared compiler infrastructure.
  • Evaluation methodology (within the PR): Benchmarked against Torch NPU using BF16 GEMM, FP8 conversion, and GQA backward, each across 4 shapes; GEMM and GQA report TFLOP/s, FP8 conversion reports effective GB/s.
  • Scale and attribution: The branch is in sync with main and 134 commits ahead, with fixes still being appended within the window (the latest commit restores broadcast vectorization and scalar reduction, early 09-30); 10 co-authors, including 2 developers with deepseek.com attribution.
  • Related: The 0.1.15 version bump PR opened the same day changes only one line in VERSION (0.1.14 to 0.1.15), landing adjacent to the 950 backend; whether they ship as one batch remains to be seen.

Assessment: If merged, Ascend 950 will land for the first time as a native backend in the main repo (previously Ascend support was mainly carried by a standalone adaptation repo), upgrading “multi-backend” from mirror-repo adaptation to a trunk capability; for users, Ascend 950 operator development and DeepGEMM-style kernels can be written directly in TileLang. The risk lies in the sheer size (379 files, roughly 80k lines added) and the review and post-merge stability costs it entails, plus the fact that NPU-side performance currently has only the public figures from charts within the PR.


I. Core Project Progress

1.1 Main Repo: Simplify Unused-Binding Eliminator Merged — Read-After-Write, Side Effects, and Metadata References All Covered by Regression Tests (09-29)

Date: 2026-09-29 Source: Simplify Unused Bindings (#3293)

  • The item newly opened in the previous issue was merged in this window: UnusedBindRemover was added to tl.Simplify, repeatedly removing unused BindNodes when their values can be discarded without dropping associated side effects or volatile buffer reads.
  • Boundary rules: reflection fields and containers are traversed to preserve bindings referenced by metadata; bindings used in buffer definitions and load predicates are always preserved; regression cases cover unused macro bindings, read-after-write, side effects, volatile reads, metadata references, and shared expression nodes.
  • Significance: the expansion of tl.Simplify’s cleanup capability moves in the same direction as last week’s modulo-remainder bound fix (#3294) — the transformation layer’s “what may be deleted, what must be kept” is gradually being written into a testable, explicit checklist.

1.2 Main Repo: Windows Wheel Build Restores ccache Hits — clang-cl Argument Split into Separate Form (09-29)

Date: 2026-09-29 Source: Windows Build Fix (#3305)

  • Problem: after ccache was restored for the Windows wheel build, TVM was still repeatedly recompiled — about 32 minutes per run, 554 misses, a hit rate of only 20.63%, with the same pattern across three consecutive nightly builds. The root cause was CMake emitting the concatenated -imsvc<path> form for clang-cl, which ccache 4.9.1 folds into the preprocessor hash; whenever the PEP 517 dependency directory changed, the cache key changed. The existing wrapper could normalize preprocessor text but not eliminate command-line differences.
  • Fix: -imsvc <path> was changed to the separated form, stabilizing the cache key; created and merged the same day, with review annotations automatically summarized for re-check.

1.3 Main Repo Queue: Seven New PRs and an Atomic-Operation Fix Line, 116 Open PRs (09-29/09-30)

Date: 2026-09-29, 2026-09-30 Source: Main Repo PR List / Atomic Operation Fix (#3307) / DeepSelect Example (#3304) / Packed Vector Comparison (#3303)

  • Seven new PRs opened in the window: #3303 covers FP16/BF16 packed vector comparison (fixes bug #3302; mask builtins normalized to 0/1, correct not-equal for NaN; scalar fallback retained for CUDA below 12 and low compute capability); #3304 submits a TileLang example of DeepSeek DeepSelect Top-K (no inline PTX; depends on #3303 and #3296 under review; provides 129 measured cases against torch.topk on an RTX 5090); #3306 and #3307 form the atomic-operation fix line (int64 atomic max/min and AtomicStore crashes for int64/bf16/fp16, traceable to the type normalization in #1716; #3306 was closed unmerged, #3307 continues under review); #3308 and #3309 are covered in today’s highlights.
  • Existing updates: SM120 register A GEMM (#3286), z3 thread inference (#3291), C source import line breaks (#3289), popcount extension (#3300), and others all received updates within the window.
  • Queue temperature: 116 open PRs (112 in the previous issue); 2 merged in this window (#3293, #3305), continuing the “many opened, few merged” pattern.

1.4 TileOPs: Four INT8 Quantization and Dequantization Kernels Merged — Bit-Exact Against torch Reference (09-29/09-30)

Date: 2026-09-29, 2026-09-30 Source: Per-Tensor Dequantization (#2304) / Per-Channel Dequantization (#2306) / Per-Tensor Quantization (#2307) / Per-Channel Quantization (#2308)

  • The previous issue’s “define the contract first, then implement” has entered the kernel-delivery phase: four operator layers in the INT8 family — quantization and dequantization — landed on the kernel side, their manifest status moved to implemented, and outputs are bit-exact against the torch reference implementation (including the boundary case where scale underflows to zero).
  • Per-tensor dequantization (#2304): q is read as a flat sequence; each thread converts one 16-byte output vector per step, with four steps per 64-thread block and all loads issued first; code points are converted via exponent-biased integer addition plus FADD (bypassing I2F); block outputs go to shared memory first and are then written out in bulk (using cp.async.bulk on SM90); a new L1 evict-last load helper was added (streaming_load.h).
  • Per-channel dequantization (#2306): vectors spanning rows select between the two rows’ scales, running as a single program for any K; per-channel quantization (#2308): one CTA owns one row with the row resident in registers, scales use the reciprocal plus two FMAs for correct rounding, and the row head/tail share vectors with neighboring rows under masking.
  • Per-tensor quantization (#2307): completed in a single cooperative launch — the first stage reads in and exchanges amax shards (grid-level barrier), the second stage quantizes and writes out; the kernel is built on each call, with the compilation boundary explicitly declared.
  • Progress: of the 9 spec-only entries in the quantization family, 7 have entered kernel implementation (six INT8 plus one INT4); FP8 per-block and SmoothQuant have yet to see kernel PRs.

1.5 TileOPs: INT4 Per-Group Quantization Kernel Merged, Two Block-Quantization PRs Queued (09-30)

Date: 2026-09-30 Source: INT4 Per-Group Quantization (#2317) / Per-Block Quantization (#2316, under review) / Per-Block Dequantization (#2318, under review)

  • INT4 per-group quantization landed two kernels: packed_weight, weight_scale, and weight_zero are bit-exact against the torch reference, and it adopts the packing format of GemmW4A16FwdOp.repack — format alignment with the W4A16 weight path is the most practical compatibility point here.
  • Kernel highlights: each lane holds one 32-element block, so no cross-lane data movement is needed within a 128-element K step; a four-by-four word transpose produces 16-byte contiguous storage; group sizes support powers of two from 32 to 1024; a row-kernel variant is also provided.
  • Two under review: INT8 per-block quantization (#2316) and per-block dequantization (#2318) — both follow the scheduling template of #2304, completing the last piece of the “tensor/channel/block” three-tier set.

1.6 TileOPs: Sampling Operator Layers Merged — Six Items Queued Last Issue Formally Landed (09-29)

Date: 2026-09-29 Source: Sampling Operator Layers (#2299)

  • The six sampling operator layers recorded as “queued for review” in the previous report (TopKMask, MinPMask, TopPMask, TopKTopPMask, SamplingFromProbs, ChainSpeculativeSampling) were merged on the afternoon of 09-29: operator layers, workloads, and tests were all prepared in one go, with reference values aligned to FlashInfer 0.6.16 (bf16/fp16/fp32), and random numbers implemented via Philox4x32-10 integer-tensor arithmetic, consistent with Random123 known-answer vectors. The operator layers remain declarative (kernel_types = {}), with kernel implementations pending future PRs.

1.7 TileOPs: Four Attention PRs — FP8 Expression Lowering, varlen Outbound via Shared Memory, SM89 Head-Dim 512 Fix (09-29/09-30)

Date: 2026-09-29, 2026-09-30 Source: FP8 GQA Leaf Operators (#2294) / varlen Prefill Outbound (#2313) / SM89 Head-Dim 512 (#2312) / Sliding-Window Test Tiering (#2311)

  • FP8 GQA lowering (#2294): output accumulator zeroing now uses T.clear, the KV tail mask is expressed as a layout-aware T.Parallel loop, and two now-useless CUDA helpers were removed; this is an increment on #2113. The evaluated T.tanh replacement for softcap was deliberately excluded due to regressions — the approximate tanh path is retained.
  • varlen outbound via SMEM (#2313): varlen GQA prefill output is no longer written out scattered from the acc_o fragment, but written out in full 16-byte rows through the now-dead query-block shared buffer, at zero extra shared-memory cost; partial blocks retain guarded direct writes.
  • SM89 head-dim 512 fix (#2312): the single-warpgroup 32-row block no longer keeps the score block in registers (a layout-inference conflict previously meant the block_m=32 candidate never compiled successfully); a 32×16 candidate was added to the default configuration (requires 68 KiB, below the SM89 limit).
  • Test tiering (#2311): the 8 sliding-window varlen cases are marked SM90-only — previously they appeared as failures rather than skips on SM89.

1.8 TileOPs: Two PRs on Kernel Interfaces and Shared-Memory Governance (09-29)

Date: 2026-09-29 Source: Kernel Interfaces and Dispatch (#2300) / Shared-Memory Limit Table (#2303)

  • Kernel interfaces (#2300): kernel interfaces are declared for operators and dispatched through a single lookup — seven concepts including “kernel interface, implementation, invocation spec, availability” are written into the ops-design document; replacement implementations are validated against the interface contract rather than by key name alone, and backends can add their own implementations alongside the built-in ones.
  • Shared-memory governance (#2303): shared-memory limits are consolidated into a unified table; FP8 GEMM declares supported_archs = [90] (SM89 invocations are explicitly rejected rather than mistakenly selecting a TMA/WGMMA kernel); sparse MLA evaluates a lower bound for invocations that do not fit in shared memory and rejects them (e.g., the case on SM86/SM89 with 64 heads per block and d=512 requiring no less than 100 KB, exceeding the 99 KB limit, which previously failed at launch).

1.9 TileOPs: Three PRs on CI and Engineering Standards — Eight-Way Concurrent Smoke, Absolute Imports, Single-Pass Manifest Validation (09-29/09-30)

Date: 2026-09-29, 2026-09-30 Source: GPU Smoke Eight-Way Concurrency (#2314) / Absolute Import Standard (#2310) / Single-Pass Manifest Validation (#2319)

  • Smoke speedup (#2314): PR smoke tests changed from single-process single-GPU to eight worker processes; signature checks are generated lazily; and a barrier conflict in bs1 decode was fixed along the way. The background is worth recording: on main-branch pushes, test count rose from 4281 to 5174 (+21%) and pytest time from 324 seconds to 605 seconds (+87%), with the increment coming mainly from CPU work in the manifest system rather than kernel tests.
  • Standards (#2310): all imports under src/tileops/ were changed to absolute imports and enforced with ruff TID252 (ban-relative-imports = "all", 234 files rewritten); design docs updated accordingly.
  • Manifest validation (#2319): validate_manifest.py runs once per PR instead of three times (this test occupies 22 to 32 seconds of the critical path in GPU smoke), and the FMA saturation metric now follows the sampled GPU.

1.10 TileOPs: Two Elementwise Speedups — Division Fallback Once per Thread, f32 Dual 16-Byte Loads (09-29/09-30)

Date: 2026-09-29, 2026-09-30 Source: Division Fallback (#2305) / f32 Loads and bool Stores (#2309)

  • Division family (#2305): after FloorDivide, Remainder, and floor-mode Div were made bit-exact in #2291, they regressed by 1.01x to 1.28x; now fast_func performs a uniform fast check per thread, reading back operands and executing the exact fallback only on failing threads, keeping the kernel within the 32-register limit; the 16-bit fast path uses a * (1/b) with upward correction.
  • f32 loads (#2309): same-shape binary kernels now use two 16-byte loads per thread (matching inductor’s XBLOCK 1024 four-warp shape); bool results for comparison and logical ops are now stored once per contiguous per-thread segment — the rows previously running 1.002x to 1.006x of torch-compile have thus converged.

1.11 Nightly Snapshot: 1030 Correctness Items, 1355 Benchmark Items, Zero Failures and Zero Errors (Early 09-30)

Date: 2026-09-30 Source: Snapshot Commit b8f4329a35 / Snapshot Environment Metadata

  • One snapshot in the window (09-30 02:53, run 36608962204; corresponding to TileOPs main branch 2ce972f98f, i.e., the #2313 merge point): 1030 correctness items (2 skipped), zero failures and zero errors; 1355 benchmark items (39 suites), zero failures and zero errors. Environment: H200, CUDA 13.2, torch 2.13, tilelang 0.1.11+cu132.
  • Compared with the previous snapshot (b7f5a1b4bd, 1027 correctness, 1335 benchmark): correctness +3, benchmark +20; the standard remains “zero failures.” Note: this snapshot’s merge point predates the quantization kernel batch (#2304 to #2317); that batch’s regression status will appear in the next snapshot.

II. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads)

2.1 Ascend: Another Fix Within the 950 Release Branch Window — Broadcast Vectorization and Scalar Reduction Restored (09-30)

Date: 2026-09-30 Source: Release branch ascend-950-0930

  • The release branch received its latest fix commit in this window (early morning 09-30, commit 39691ebaac): broadcast vectorization and scalar reduction restored — one of the latest fixes within the window for the “Public Release 9/30” PR; the branch is synced with main and 134 commits ahead, with the main PR still open. See today’s highlights for technical details.

2.2 Ascend: Two Fixes in the ascend Repo — a One-Line TVM Patch Reveals the Real Cause of the Revert (09-29)

Date: 2026-09-29 Source: TVM backtrace out-of-bounds patch (#1846) / CI submodule cleanup (#1847)

  • The real cause of the revert comes to light (#1846): a one-line patch added to the pinned TVM submodule fixes an out-of-bounds read when TVM generates an erroneous call stack (std::string tmp(symname, symsize) changed to the single-argument constructor). The backdrop is the “#1829 reverted in full 4 days after merge” recorded in the previous report: 9 compiler negative cases expected to throw tvm.error.InternalError consistently SIGSEGV on the x86 runner while consistently passing on the ARM runner; a single-variable A/B experiment on the same machine showed the crash stems from this TVM defect and is unrelated to the logic of #1829. The patch follows the existing apply_tvm_patches.sh mechanism (automatically applied by the three build entry points, skipped if already applied, and failing the build outright rather than silently when no longer applicable in the future).
  • CI cleanup (#1847): on self-hosted runners, actions/checkout does not clean submodules, and git submodule update only moves the pointer when the commit changes, causing the TVM patch applied by the previous PR to leak into all subsequent builds (the #1846 build failed to apply the patch due to this residue); the fix resets the submodule before building.
  • Assessment: the revert of #1829 was not due to a defect of its own — whether the reverted feature is resubmitted against a clean baseline is the next thing to watch.

2.3 Hygon: Async Copy and Cache Swizzle Lowering Improvements Merged (09-29)

Date: 2026-09-29 Source: Hygon async copy (#13) / Merge commit 145fc5a8e7

  • The open PR noted in the previous report was merged in this window (9 files, +497/-41): improved lowering for async copy and cache swizzle. The changes are concentrated in codegen_hcu (+121/-29), op/copy.cc (+269/-8), tl_templates/hcu/copy.h (+47), and amd_buffer_addressing.hpp (+20), plus a minor LDS policy tweak for GEMM; Co-authored by Teng Huang. Hygon is the only domestic backend besides Ascend with code progress this period.

2.4 MetaX / Moore Threads / MLIR Ascend: Official Repos Silent (09-30 Check)

Date: 2026-09-30 Source: MetaX repo / Moore Threads repo / MLIR Ascend repo

  • MetaX (last push 09-24), Moore Threads (09-17), and MLIR Ascend (09-24) official repos saw no pushes within the window; no new commits on the community side either. Domestic backend activity this week is concentrated in Ascend (950 backend + two fixes) and Hygon (#13 merged).

III. Ecosystem and Adopters

3.1 TileFoundry: AtomSched Phase 4 merged — HIR lowered to TIR with scheduling CLI (09-29)

Date: 2026-09-29 Source: AtomSched Phase 4 (#192) / TileFoundry repo

  • A phased merge of 99 files, +4885/-2885: adds a new module-level ConvertHIRToTIR pass that lowers authored scheduling HIR into validated TIR; adds three CLI report groups tilefoundry schedule finalize / facts / candidates, and removes the rejected matched workflow.
  • Design direction: IR declarations and access relations become the single source of truth; duplicate problem plans, pre-scans, inferred distributions, and hand-written accessors are removed; report listings, ordering, pattern rendering, and HIR/TIR candidate pairing move to shared reflection and a neutral registry. This is the consolidation phase of the “declaration-driven scheduling” line.

3.2 Docs site: quantization, sampling, and shared-expert MLP operators enter the API reference (09-30)

Date: 2026-09-30 Source: Docs site update (#57)

  • The TileOPs docs site updated its API reference in sync after the quantization, sampling, and shared-expert MLP operators were merged — following the same-day cadence of operator launches in #2299, #2304 through #2317.

3.3 Adopter check: no in-window pushes from TileKernels or FlashQLA (09-30)

Date: 2026-09-30 Source: TileKernels repo / FlashQLA repo

  • DeepSeek TileKernels remains stalled at 2026-04; Qwen FlashQLA’s last push was 09-28 (previous window, already reported for SM100 tcgen05), with no new commits this window. Both adopters’ public activity is in an intermittent phase.

3.4 Release cadence: 0.1.15 bump PR opened; TileRT and TileFoundry release pages stalled (09-30)

Date: 2026-09-30 Source: Main repo releases / TileRT releases / TileFoundry releases

  • Main repo v0.1.14 (released 09-02) has now been out for 28 days; the 0.1.15 version bump PR opened this window (see today’s highlights), with no official release yet; TileOPs has no standalone release.
  • The TileRT release page remains at v0.1.5.post2 (the v0.1.6 release PR reported last period was merged but no official publication has appeared); TileFoundry’s latest is still v0.0.2 (09-10).

IV. Community, Tutorials, and Events

4.1 Community repo check: no new commits in-window (09-30)

Date: 2026-09-30 Source: Community tutorials repo / Sophgo BM1690 pipeline repo / S5000 profiling tool / tvm_tilelang_cookbook

  • All four community repos (tilelang-tutorials, bm1690-pipelines, Tilelang_musa, tvm_tilelang_cookbook) had no new commits in-window; cookbook only had an automated star-history push on the chart branch (not substantive activity, not counted). The tutorial site’s CPU examples and BM1690 pipeline progress reported last period predate this issue and are not repeated here.

4.2 Media and academia: zero Google News hits, no new arXiv preprints (09-30)

Date: 2026-09-30 Source: Google News / Hacker News / arXiv

  • Google News RSS returned zero hits across multiple Chinese and English queries (via proxy) in the 24-hour window; “Ascend 950” hits in both Chinese and English were pre-window stock-market and macro pieces, unrelated to the TileLang backend release (dated 9/30, media has yet to follow up). Hacker News in-window items are unrelated to this topic; the latest arXiv search for “tilelang” is still TileSight from 07-24.

V. Trend Observations

5.1 From contract to kernel: the quantization family delivered within 48 hours

Last period’s observation was that “this batch of operators has no kernels yet; watch the kernel landing cadence.” In under 48 hours, 7 of the 9 quantization-family operators entered kernel implementation: four INT8 kernels merged with bit-exact alignment to the torch reference, one per-group INT4 kernel merged and aligned to the W4A16 packing format, and two per-block INT8 kernels under review. A PR description paradigm has also taken shape — each entry notes bit-exact consistency with the reference implementation, compilation boundaries, manifest status, and rejection semantics, following “prove correctness first, then discuss performance.” At this pace, the landing of the remaining FP8 per-block and SmoothQuant kernels is the next window’s focal point.

5.2 Ascend 950: the multi-backend narrative moves from “adaptation repo” to “main-repo native”

The Ascend 950 backend was submitted directly to the main repo as a PR with 379 files and roughly 80,000 added lines, opened the same day as the 0.1.15 version bump — TileLang’s “multi-backend” is upgrading from mirror-repo adaptation to trunk capability, with the benchmark baseline directly choosing Huawei’s own Torch NPU. The authorship shows a deepseek.com email alongside multiple core community developers, 10 co-authors in total, indicating the Ascend line is a full-scale, multi-party investment rather than a single-point adaptation. A sober view is warranted: the PR is not yet merged, its size is enormous, and performance data currently comes only from charts within the PR; post-merge mainline stability and follow-up maintenance costs will be longer-cycle observation points.

5.3 Engineering capacity: CI rework under +21% test volume

Main-branch test count rose from 4281 to 5174 (+21%), with pytest runtime up +87%, the increment coming mainly from CPU-side work on the manifest system — the direct cost of the “200 operators, contract-first” strategy. The same-day response was three engineering measures: eight-way concurrent GPU smoke tests, lazy signature checks, and single-pass manifest validation (removing two redundant validation rounds per PR). Reforging CI from “good enough” to “scale-bearing” during a capability-expansion phase is the underrated infrastructure narrative of this week.

5.4 Gaps and risks

First, multi-backend divergence continues: MetaX (since 09-24), Moore Threads (since 09-17), and MLIR Ascend (since 09-24) are silent, with activity concentrated in Ascend and Hygon. Second, the main repo’s “many opened, few merged” pattern has not reversed — 2 merged and 7 newly opened this window, with 116 open PRs; review bandwidth remains the bottleneck, and the merge cadence of the large Ascend 950 PR will directly test this. Third, the re-raising of the reverted feature in #1829, the official publication of 0.1.15, and the belated release-page entry for TileRT v0.1.6 all remain unresolved. Fourth, validation of quantization and sampling-style kernels is currently limited to bit-exact alignment and single-machine benchmarks, with no public validation of end-to-end model accuracy impact.


Appendix: Sources and Verification Notes

Source Verification Result
tile-ai organization (29 repos) 7 repos had pushes in the window: tilelang (2 merged, 7 newly opened), TileOPs (17 merged), TileOPs-nightly (1 snapshot), tilelang-ascend (2 merged), tilelang-hygon (1 merged), TileFoundry (1 merged), TileOPs.github.io (1 commit)
Main repo tilelang Merged #3293, #3305; newly opened #3303 through #3309, 7 total (#3306 closed without merge); 116 open PRs; release branch and version-bump branch active in the window
TileOPs 17 merged (#2299 through #2319 range); 17 newly opened in the same window; 4 under review (#2172, #2315, #2316, #2318); quantization family dominates
tvm (submodule repo in the same organization) No commits in the window
TileOPs-nightly 1 snapshot b8f4329a35: 1030 correctness items (2 skipped), zero failures, zero errors; 1355 benchmark items (39 suites), zero failures, zero errors
TileOPs.github.io 1 commit (#57 quantization, sampling, and MLP operators added to API reference)
Five domestic backend repos Ascend: 950 release PR #3308 and two fixes in the ascend repo; Hygon: #13 merged; MetaX (after 09-24), Moore Threads (after 09-17), MLIR Ascend (after 09-24) no activity
Adopters and community TileKernels, FlashQLA no pushes in the window; four community repos no new commits (cookbook only automated branch pushes)
Google News / Hacker News / arXiv Zero hits for Chinese and English queries (Ascend 950 hits were all pre-window stock market and macro pieces); HN unrelated to this topic; arXiv latest remains 07-24 TileSight

Complete Source List