Reporting period: September 30, 2026 (Wednesday) to October 7, 2026 (Wednesday), 8 days total — National Day holiday special consolidated edition (daily reports suspended from October 1 to 7; this issue is the holiday merged edition) Sources: The final pre-holiday issue of this publication’s daily TileLang updates (2026-09-30, window through 07:00 that day); holiday-period activity is based on empirical verification of a 216-hour window across the tile-ai organization (29 repositories) (commits, PRs, tags and releases, nightly pipelines; verification extended through the afternoon of 10-08); the news side consists of public searches across multiple keyword sets over the holiday window (approximately 9 days) (see appendix)


I. Weekly Highlights

  • TileLang v0.1.15 released before the holiday: Ascend 950 native backend lands with the release (09-30): At 07:31 Beijing time, the Ascend 950 backend mega-PR labeled “Public Release 9/30” (#3308, 134 commits, 379 files, +79956/-733) and the version bump PR (#3309) were merged in succession; at 08:29, v0.1.15 was officially released (release page). Four highlights of the release: Ascend 950 native support (end-to-end NPU backend: native code generation, automatic Cube/Vector scheduling and synchronization, SIMD/SIMT hybrid programming), CUDA automatic warp specialization (dispatching TMA loads, MMA compute, and stores to dedicated warp groups by role), unified block-scaled GEMM (unified T.gemm_blockscaled semantics, SM100 instruction selection and SM120 fragment expansion), and a more expressive Python frontend (compile-time iteration and comprehensions). At the time the pre-holiday daily brief went to press, both PRs were still open; this publication verifies and completes the record: merged on the morning of that day, with the release page updated in sync; Ascend A2/A3 support remains in the community-maintained adaptation repository.

  • DeepSeek open-sources infrastructure components for Ascend, with the Ascend version of TileLang included (09-30): On September 30, DeepSeek announced the open-sourcing of infrastructure components for Huawei’s Ascend compute platform, covering the TileLang high-level language compilation tool, compute libraries, and distributed communication libraries (as reported: a one-to-one correspondence with the NVIDIA platform components). Organizational corroboration: deepseek-ai had push records that day across multiple repositories including DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA, and DeepSelect. The report quoted that “every TileLang operator used in DeepSeek training has a corresponding high-performance implementation on Ascend.” Chinese and international media followed up intensively starting that day (approximately 210 hits within the window across ten keyword groups; see Chapter 5 for details).

  • PTO backend merged into the main repository: Ascend 950 gains a second code path on the same day (09-30): #3310 (+11406/-31, 36 files) was merged at 12:08 — registering the pto target and execution backend, generating PTO device code, running via kernel wrapping plus a Cython runtime interface, and compiling PTODSL kernels with ptoas and BiSheng; accompanying helper functions cover GEMM, SIMT/SIMD, synchronization, hybrid kernels, cache-bypass access, and Philox random numbers. More than ten follow-up items have been opened (PTO execution routing, import isolation, DeepGEMM code generation, etc.), still under review.

  • TileOPs merges 137 PRs over the 8-day holiday: quantization, sampling, and linear attention advance on three fronts in force (09-30 ~ 10-07): 137 merges, 143 newly opened (plus 2 more on the 10-08 return-to-work day); 09-30 saw a single-day peak of 32. The sampling line landed six kernels covering the full chain from masking to rejection sampling; the linear attention line saw intensive progress on GLA, Gated DeltaNet, DeltaNet, and Kimi Delta Attention (including tensor-core WY inverse #2365); the quantization and GEMM lines covered multiple forms of INT4/INT8/FP8; kernel interface refactoring, the PyPI release pipeline, and a README overhaul landed in sync (see Chapter 4 for details).

  • Moore Threads MUSA prepares public 0.1.15 and releases v0.1.15+musa.1 (09-30): 53 commits in a single day (51 MUSA-side features plus 2 upstream syncs) — the mp31 architecture batch (tme series, sqmma, wmma, warp specialization, cache hints, peer memory, etc.) and the runtime batch (base target and backend entry points, device code generation, JIT wrapping, ipc/vmm, dlpack compatibility, explicit streams, etc.); after “Prepare public TileLang MUSA 0.1.15” landed in the evening, release v0.1.15+musa.1 was tagged at 17:59 (the previous version, v0.1.14+musa.1, was on 09-11).

  • Hygon lands three consecutive merges (09-30): matrix storage (gfx946, #14), multi-XCD Cython launch support (#15), and LDS policy conflict fix (#16) — continuous progress following the pre-holiday async copy improvement (#13); Hygon is the only domestic backend besides Ascend with code progress this period.

  • Nightly pipeline produces 12 uninterrupted snapshots: benchmark suite expands to 1,939 items (09-30 ~ 10-08): 12 snapshots produced day by day, correctness cases from 1,009 to 1,044 (fluctuating during the test reorganization period), with the latest snapshot at 1,027 items and zero failures; benchmark measurements expanded from 1,355 items (39 suites) to 1,939 items (52 suites); the only failure across the entire period — an FFT case in a 10-07 snapshot due to CUDA out of memory (environmental) — was restored to all-green in the next snapshot.

  • Adopters follow up on the same day: TileKernels v2.0.0 released with Ascend support, FlashQLA updated to v0.1.3 (09-30): DeepSeek TileKernels released v2.0.0 (09:58), with the README adding “Huawei Ascend support”: the dual backend is automatically selected by the runtime, and the same Python API runs on both NVIDIA GPUs and Huawei NPUs (requires TileLang 0.1.15+ and CANN 9.2.0+); Qwen FlashQLA updated to v0.1.3 (21:02), accompanied by sm90/sm100 autocp coefficients and benchmark updates. Adopters and upstream releases landed on the same day, validating the linkage.


II. Core Project Progress (Main Repo tilelang)

Versions & Releases: v0.1.15 Officially Released (09-30)

  • Timeline: At 09-30 07:31, #3308, labeled “Public Release 9/30”, and the version bump #3309 were merged; at 08:29, v0.1.15 was released—28 days after the previous version v0.1.14 (09-02), consistent with the existing monthly release cadence.
  • Ascend 950 Section (excerpt from release notes): Added tilelang.ascend.language and target="ascend" (corresponding to dav-3510); combining Cube GEMM and Vector computation within the same kernel (T.SimdVF / T.SimtVF regions); explicit UB/L1/L0 storage, tile copies, cross-core transfers, and MXFP8/MXFP4 block-scaled GEMM; automatic scheduling, pipelining, multi-buffering, layout inference, and synchronization insertion; integration with BiSheng compilation, tvm_ffi and Cython execution, PyTorch NPU tensors and streams, and NPU profiling; examples include GEMM, DeepGEMM-style kernels, FlashAttention forward and backward, RMSNorm, and FP8 quantization. A2/A3 support remains in the community-maintained TileLang-Ascend project.
  • CUDA Section: Role-based automatic warp specialization (can be enabled via environment variable); SM120 block-scaled GEMM extensions (fragment-resident A operand, row-major scale fragments, odd per-warp atomic grids); T.fma/T.fmul round-to-nearest; FP8 sparse MLA forward example (DeepSeek V3.2, Hopper, #3224); WGMMA/UMMA K-panel stride fix and sparse MMA async fence.
  • Language, Compiler & Runtime Section: Kernel launch encoding target neutralization; Python compile-time iteration and comprehensions; coalesced_width as a hint; AllReduce thread stride and reduction NaN propagation dtype validation; cache parameters switched to JSON, atomic publishing of CUDA binaries, wall-clock benchmarking, wheel verification on GPU runners, performance regressions now fail CI.
  • ROCm, CPU & Metal Section: ROCm support for DeepSeek V3.2 Top-K selector, wave64 Hadamard, and GLM-5.3 k-pool examples; CPU scalar GEMM computes per accumulation dtype; Metal 32-bit atomic add.
  • Compatibility Notes: T.Pipelined removes unused sync/group parameters, T.gemm_sp removes k_pack; T.symbolic becomes a deprecated alias (new code should use T.dynamic); TILELANG_CACHE_VERIFY_HASH removed, cache hash verification is now mandatory.

Holiday Fix Line & Small Engineering Steps (10-02 ~ 10-08)

  • Correctness Fixes: #3361 supports bfloat16 scalar parameters (10-02); #3334 decodes the reserved UE4M3 NaN scale codeword in NVFP4 (10-03); #3429 preserves explicit casts during index promotion (10-06); #3441 LDSM copies preserve shared region offsets (10-07); #3439 fp8 T.gemm on SM89 now raises a gated error instead of crashing on launch (10-07).
  • JIT & Tests: #3435 lazy kernel builds preserve variadic arguments (10-07); #3443 removes unsupported BF16 scalar regression test (10-07).
  • Engineering Misc: #3418 merges two identical branches (10-03); #3431 pre-commit auto-update (10-07).
  • First Post-Holiday Merge (10-08): #3299 fixes round-to-nearest-even for T.round on ROCm—merged midday on release day, with post-holiday cadence resuming as expected.
  • Summary: 12 merges across the window (three on 09-30 as the release batch, nine from 10-02 onward as small holiday fixes), slightly down from 14 in the single week before the holiday; the substantive increment is concentrated in the review queue (next section).

Review Queue Observations

  • Scissors Gap: 96 new PRs opened and 12 merged during the window; about 45 opened on 09-30 alone, comprising two categories—an Ascend-themed batch (about 20 PRs: #3316 through #3331, covering PTO execution routing, AutoSimtVF, T.PersistentKernel, torch_npu task queue, operator-generation agent, DeepGEMM code generation, RNG reset, etc.) and a “reject unsupported patterns” batch of systematic small-fix hardening (#3332 through #3372, over forty PRs).
  • Tests & Infrastructure: #3313 improves test infrastructure, #3315 speeds up Windows wheel builds.
  • Pressure Assessment: The pre-holiday daily report cited 116 open PRs; the holiday added further net pending review tasks. Post-holiday review bandwidth will directly determine the merge speed of the 950 follow-up queue and the cadence of the next release.

III. Multi-Backend Adaptation (Ascend / MetaX / Hygon / Moore Threads, etc.)

Ascend: Dual backends (950 and PTO) land; adaptation repo wraps up after the holiday

  • Main repo (releases and backends): See Chapter II — #3308 (950 native backend) and #3310 (PTO backend) merged on the same day.
  • Main repo (follow-up queue, still under review): PTO execution routing #3316, import isolation #3327, DeepGEMM code generation #3330, AutoSimtVF pass #3320, causal sparse attention example #3322, gemm_l1 syntactic sugar #3323, torch_npu task queue #3324, T.PersistentKernel #3325, aibin single-stage compilation #3326, L0C dual-copy conversion #3328, operator generation and tuning agent #3329, RNG spacing reset #3331.
  • Adaptation repo tilelang-ascend (default branch ascendc_pto): Three merges — #1746 documentation completion (09-30); #1848 “Explicit and automated C/V ownership” (10-08) — independently restores “Part 2” of the previously reverted #1829 (default partitioning of C/V with execution ownership checks); #1786 rejects non-flat contiguous BufferRegions in T.tile binary operators (fixes #1680, 10-08). Five new bot-authored PRs opened (09-30): Vector mask state reuse #1850, sync dependency lifetime #1851, FP32 Reduce2D shape planning and workspace contract #1852, atomic publication of cached kernel libraries #1853, automatic cross-core sync configuration validation #1854.
  • Assessment: The main repo’s “native backend” and the adaptation repo’s “community-maintained line” proceed in parallel; two merges on the first day after the holiday (#1848, #1786) show momentum resuming immediately upon return to work, and the re-submission path for reverted features (#1829 Part 2) has been validated.

MLIR Ascend (tilelang-mlir-ascend): Dialect branch takes off

  • On 10-06, the mlir-dialect branch was created along with PR #195 “mlir dialect setup with add example” (dialect hookup and addition example); #193 (report update and argmax kernel) and #194 (Tlir add pipeline expert mode), opened on 09-28, remain pending review. The repo’s direction is advancing from operator examples toward the compiler dialect layer.

Hygon: Three consecutive merges

  • #14 matrix storage (gfx946, 11:26), #15 multi-XCD Cython launch support (15:47), #16 LDS policy conflict handling (17:49) — three PRs on the same day, continuing the matrix core and multi-XCD direction after #13 (async copy and cache swizzle).

MetaX: Silent throughout the period

  • No pushes in the official repo during the window (most recent push 09-24), consistent with the pre-holiday daily report; no PR or tag activity.

Moore Threads (MUSA): Public 0.1.15 preparation and v0.1.15+musa.1 release

  • Commit batch: 53 commits (51 on the MUSA side plus 2 upstream syncs), concentrated in two waves from midday to evening on 09-30 — the foundational batch (target and backend entry points, release configuration, musa language frontend defaulting, lazy pass pipeline, device code generation, runtime device registration, JIT wrapping, lazy-loading stubs, skipping driver unload on shutdown) and the feature batch (mp31: tme load/store/im2col/descriptor prefetch/out-of-bounds zero-fill/swizzle/fp8, sqmma, wmma, warp specialization, lsu cache hints, peer memory, copy select; general: dp4a, vectorized cast, fast divmod, ldg/stg, atomic family, async copy, math and transpose/fill/scan/reduce/debug, threadblock swizzle; runtime: symmetric ipc, ipc and optional vmm, dlpack device compatibility, system-scope fence, explicit streams).
  • Release: At 17:59, v0.1.15+musa.1 was tagged (corresponding to main repo 0.1.15); the previous version v0.1.14+musa.1 was on 09-11. The synchronization cadence between the public version and the main repo’s version line continues.

Other: ROCm backport in the tvm submodule repo

  • tile-ai/tvm (tilelang_main branch) merged PR #79 “Backport HIP device existence checks” (10-07) — a maintenance backport, with no other activity during the window.

IV. Ecosystem and Adopters (TileOPs, TileRT, TileFoundry, TileKernels, FlashQLA, etc.)

TileOPs: A Panorama of 137 Merges in 8 Days

  • Total volume: 137 merges (plus 2 on 10-08 as work resumed: #2482, #2483), 143 newly opened. Day by day: 32 on 09-30 (peak), 14 on 10-01, 23 on 10-02, 15 on 10-03, 11 on 10-04, 15 on 10-05, 13 on 10-06, 14 on 10-07 — averaging about 17 per day, several times the previous period (21 for the whole week, about 4 per day). The old impression of “running at a low level during the holiday” does not hold.
  • Sampling line (full chain of six kernels): #2326 min-p mask, #2327 top-k mask, #2332 top-k fused top-p single kernel, #2334 top-p mask, #2341 token extraction from probability rows, #2346 speculative draft chain verification; wide-row serving and single-CTA limit governance (#2445, #2448).
  • Linear attention line: GLA forward prefill and decode (#2355, #2368, #2379), Gated DeltaNet and DeltaNet prefill/decode (#2351, #2366, #2372, #2375, #2369), Kimi Delta Attention with tensor-core WY inverse progressive tree (#2365); accompanied by GDN/KDA naming unification (#2440, #2443).
  • Quantization and GEMM line: int8 W8A8, FP8 grouped expert GEMM, causal convolution, Hadamard transform, and others landed on the list (#2400, #2404); FP8 per-request scaling for packed varlen GQA (#2396), 16-bit paged GQA contract (#2386), fused RoPE (#2377); SM89 FP8 GEMM and quantization kernel off-SM90 compatibility (#2399, #2425).
  • Engineering line: kernel interface-based selection (in batches since #2321), naming conventions (#2410, #2434), PyPI self-tag release (#2359), README “Built for agents” (#2358), unified correctness checks and nightly coverage (#2429), bench.Runner benchmark refactor (#2470). The foundry-prefixed operator supply batch ran throughout the period.
  • Assessment: the four family lines (quantization, sampling, attention, linear attention) reached a “runnable plus verifiable” state within 8 days, with engineering closure actions (release, naming, interfaces) equally dense in parallel — capacity is converging toward pipelineization.

TileOPs Nightly Pipeline: 12 Snapshots, Benchmarks Expanded to 1,939 Items

  • Snapshots: 12 within the window (1 to 2 per day from 09-30 to 10-08); correctness cases ranged from 1,009 to 1,044 (fluctuating during the test organization refactor), skips converged from 2 to 0, and the latest (10-08) shows 1,027 items with zero failures.
  • Only failure: one snapshot on 10-07 showed a single failure, localized to CUDA out-of-memory in an FFT case (attempted to allocate 16 GiB, environmental); the next snapshot returned to all green.
  • Benchmarks: measurement cases expanded from 1,355 items (39 suites, 09-30) to 1,939 items (52 suites), with zero failures — new coverage and case cleanup advanced in tandem during the holiday. Data comes from parsing the test and benchmark results of the nightly pipeline branch snapshot by snapshot (as of the afternoon of 10-08).

TileFoundry: Continuous Small Steps in Scheduling and Analysis (6 Merges)

  • #196 CI checks PR titles and completes the contribution template (09-30); #193 scheduling gap closure and tutorial (10-01); #198 analysis projection reuse (10-01); #199 per-scope work and traffic statistics (10-02); #208 broadcast support on symbolic grids (10-04); #210 explicit access relations (10-05).

DeepStack: Collective Communication and NoC Fixes (3 Merges)

  • all_gather collective communication wrapper implementation (10-05), switch instance differentiation fix in the full-exchange hierarchy (10-05), and one defect fix (10-04).

Adopters: TileKernels v2.0.0 and FlashQLA v0.1.3

  • TileKernels (DeepSeek): released v2.0.0 (09-30 09:58, merged PR #34); the README added a “Huawei Ascend support” entry the same day — automatic selection at runtime between dual backends, with the same Python API running on both NVIDIA GPUs and Huawei NPUs (requires TileLang 0.1.15+, CANN 9.2.0+, Ascend 950). The adopter’s release and the main repo’s release landed on the same day, tightly coupled.
  • FlashQLA (Qwen): updated to v0.1.3 (09-30 21:02), with four commits including sm90/sm100 autocp coefficients and benchmark result updates.

TileRT: Silent Throughout the Period

  • No pushes within the window; the release page stagnation continues (already recorded in the previous period’s report), with no new entries on either the release line or the tag line, pending review next period.

V. Community, Tutorials, and Events

Media: DeepSeek Ascend Open Source Sparks a Wave of Coverage (09-30 ~ 10-07)

  • Scale: Ten groups of Chinese and English keywords totaled approximately 210 hits within the window (including cross-group and reprint duplicates), constituting the highest peak of coverage on this topic to date.
  • English highlights: SCMP “DeepSeek opens tools to help Huawei chips supplant Nvidia in AI” (09-30), Tom’s Hardware “DeepSeek and Huawei release open-source Ascend AI programming tools…” (10-01), GIGAZINE, The Next Web, Pandaily, etc.; a discussion thread appeared on Hacker News on 10-02.
  • Chinese highlights: Cailian Press “Aiming at Nvidia CUDA! DeepSeek open-sources full Huawei Ascend component suite,” InfoQ “DeepSeek open-sources Ascend platform infrastructure components, covering TileLang, compute libraries, and distributed communication libraries,” Guancha, Leiphone, OSCHINA, MyDrivers, Machine Heart, etc.
  • Industry-side extensions: Xu Zhijun publicly stated that “Ascend’s China market share has surpassed Nvidia” and that “950 SuperNodes are insufficient even domestically, with no plans for full overseas expansion for now” (10-03 to 10-07); the first Ascend 950 testing base was completed in Tongzhou, Beijing (10-04); reports on an Ascend 950 SuperNode landing in Shanghai (10-01).
  • Assessment: The coverage wave and Huawei-side disclosures reinforced each other, placing “TileLang Ascend Edition” at the center of the “de-CUDA-ization” narrative; a sober boundary—engineering details and performance claims should be judged by repositories and release notes, and end-to-end validation has yet to be made public.

Documentation Sites and Routine Updates

  • TileOPs documentation site (11 commits): API reference additions for quantization/sampling/shared-expert MLP (#57), sampling benchmark page and manifest decision page (#58), kernel interface documentation (#59), manifest and kernel distribution user guide (#60), development guide and Blog section (#61), stylesheet versioned by content (#62), naming follow-up batch (#63 to #66), and benchmark authoring guide (#67).
  • Main repo documentation site (5 commits): routine “Update docs” refreshes on 09-30 (two commits), 10-02, 10-07, etc., synchronized with the release cycle.

Community Project: BM1690 TPU Pipeline Continues to Advance

  • The community TPU extension repo (Sophgo BM1690, arcflute/tilelang-tpu-bm1690-pipelines) had 20 commits from 09-30 to 10-01: item-by-item records of compilation and correctness for three variants each of Add/Sub/Mul, latency records for coarse-grained blocks and 1024 sizes, and the PCIe handoff download pipeline (including GitHub API transfer and timeout recovery); the recording strategy is “trace every item, claim no unverified speedups.”
  • Other community repos (tutorial repo, MetaX teaching repo, cookbook) had no commits within the window.

Academic Side: No New arXiv Papers

  • No new TileLang-related preprints appeared within the search window; the most recent remains the TileSight performance model paper from 07-24.

VI. Trend Observations

  • Ascend 950’s “Day-1 versioning” marks a turning point in the multi-backend narrative: The 950 backend was delivered with v0.1.15 as a full package—”large PR plus version release plus second code path (PTO) plus usage guide”—while A2/A3 remain in the community repo, with clear division of labor between generations and maintenance—”multi-backend” has officially become a trunk capability narrative. Watch points: real-hardware performance data for the 950 and longer-cycle maintenance costs.
  • DeepSeek and Huawei push TileLang to the center of the domestic compute agenda: Coverage volume hit a historic peak (approximately 210 hits); adopter TileKernels released v2.0.0 the same day with Ascend support attached, marking the first time “upstream release, downstream version bump” has been so tightly coordinated. Risk side: narrative heat exceeds verifiable data, and end-to-end precision and large-scale training validation remain to be made public.
  • Holiday development intensity trending upward: TileOPs at 137 commits over 8 days, MUSA at 53 commits in a single day, 12 nightly pipeline runs, and two merges into the adaptation repo on the first post-holiday day—the old assumption of “low-intensity operation during holidays” no longer holds.
  • “Many opens, few merges” remains a structural tension: The main repo saw 96 new opens versus 12 merges, with the open queue growing net; the large Ascend batch and forty-plus hardening fixes are queued for review—post-holiday review bandwidth will directly determine the pace of 950 follow-up and the next release.
  • Multi-backend heat diverges while the layering stabilizes: Ascend (dual backend) and MUSA (public release) form the dual core of activity; Hygon advances steadily in small steps; MetaX and MLIR Ascend are silent; the tvm submodule repo is mainly maintenance-oriented backports—the layered pattern of “native main repo plus adaptation repo follow-up plus public release sync” has stabilized.
  • Engineering closure actions are密集: PyPI release pipeline, bench.Runner, atomic cache publishing, test unification, and naming conventions—the bottleneck is shifting from “writing operators” to “release and validation pipelines,” a typical sequence as a library matures into its engineering phase.

Appendix: Sources and Verification Notes

  • Sources: This is the last pre-holiday issue of the daily TileLang intelligence briefing (2026-09-30, collection window 09-29 07:00 to 09-30 07:00 Beijing time); daily briefings are suspended from October 1 to 7, with holiday developments subject to hands-on verification; developments after this period (from 10-08) will be covered in the next issue.
  • Supplementary verification (216-hour window, through the afternoon of 10-08): 12 of the 29 repositories in the tile-ai organization had pushes—tilelang, TileOPs, TileOPs-nightly, tilelang-ascend, tilelang-musa, tilelang-hygon, tilelang-mlir-ascend, TileFoundry, DeepStack, tvm, TileOPs.github.io, tilelang.github.io; commits, PRs, tags, and release metadata were checked repository by repository; 12 nightly snapshots were parsed one by one for test and benchmark results; adopters (TileKernels, FlashQLA), deepseek-ai organization pushes, and community repositories (BM1690 pipelines, etc.) were verified.

Source Verification Table

Source Verification Result
tile-ai organization (29 repos) 12 repositories had pushes within the 216-hour window; holiday development did not stall
Main repo tilelang 12 merges, 96 new opens; v0.1.15 released (09-30); 1 more added on the first post-holiday day (#3299)
Ascend (main repo plus adaptation repo) Main repo: Ascend 950 backend and PTO backend merged; adaptation repo: 3 merges (including #1829 Part 2 restoration), 5 new bot PRs
MLIR Ascend mlir-dialect branch and dialect PR #195 created on 10-06; #193, #194 pending review
Hygon Three consecutive merges (#14, #15, #16) (09-30)
MetaX No pushes throughout the period (most recent 09-24)
Moore Threads 53 commits, released v0.1.15+musa.1 (09-30)
TileOPs 137 merges, 143 new opens (plus 2 more upon return to work on 10-08)
Nightly pipeline 12 snapshots; correctness 1009 to 1044; only 1 environmental failure; benchmarks expanded to 1939 items
TileFoundry and DeepStack 6 and 3 commits
TileRT No pushes throughout the period; release page stagnation continues
Documentation sites TileOPs.github.io 11 commits, tilelang.github.io 5 commits
Adopters TileKernels v2.0.0 (with Ascend support), FlashQLA v0.1.3
Community repositories BM1690 pipelines 20 commits; no commits in other community repos
Media and academia About 210 hits across ten keyword groups (DeepSeek Ascend open-source wave); no new arXiv entries
  • Scope and boundaries: All times are Beijing time (UTC+8); merges and commits are counted by committer time; performance and scale data are as self-reported in PRs or release notes, and this publication has no corresponding hardware and has not independently reproduced them; nightly results are single pipeline records, not cross-comparative benchmarks; “210 hits” is the total of raw hits across multiple keyword groups (including cross-group and repost duplicates, not deduplicated item by item).
  • Suggested next period: 2026-10-08 to 2026-10-09 (the two post-holiday working days, to be covered by the next Saturday weekly report).

Main References

Data sources: Daily TileLang intelligence briefing (last pre-holiday issue, 2026-09-30) and hands-on verification for this period; reporting period is September 30 to October 7, 2026 (National Day holiday special summary period).