TileLang Daily Intelligence Report (2026-09-29)
Research window: Past 24 hours (2026-09-28 07:00 ~ 2026-09-29 07:00, Beijing time; the previous report was on 09-28, with a seamless window and no overlap). Sources: GitHub (push verification across 29 repositories in the tile-ai organization; 7 repositories had pushes within the window: tilelang, TileOPs, TileOPs-nightly, tvm, tilelang-hygon, TileOPs.github.io, tilelang-ascend; TileOPs had 14 merges, the main repo had 2 merges plus TVM submodule follow-up, and 1 nightly snapshot with zero failures; adopter FlashQLA and community repositories had updates), Google News RSS multi-group queries in Chinese and English (via proxy, zero hits), Hacker News, arXiv, community repository verification
Issue Index
- Today’s Highlights: TileOPs quantization family lands at the implementation layer first—nine operator-layer kernels merged overnight, six sampling-family kernels queued the same day (09-28/09-29)
- I. Core Project Progress
- 1.1 Main repo: TVM tail guardrail fix—inverted modulo boundary could cause out-of-bounds reads, merged after real-target reproduction (09-28)
- 1.2 Main repo: Default single-config build adds -O2—lowering geometric mean 1.19x faster, library size halved (09-28)
- 1.3 Main repo queue: Five new PRs opened and three updated, 112 open PRs total (09-28)
- 1.4 TileOPs: Silent-error kernels fixed in batch, a dozen-plus “returns result but no error” cases addressed together (09-29 early morning)
- 1.5 TileOPs: Attention consolidation continues—MHA and GQA paged decode unified, backward WS kernel merged (09-28)
- 1.6 TileOPs: Six batch speedups for reduction and norm (09-28)
- 1.7 TileOPs: Two engineering items—source migrated into csrc, guard exemptions cleaned up (09-28)
- 1.8 Overnight snapshot: 1027 correctness items, 1335 benchmark items, zero failures zero errors (09-29 early morning)
- II. Multi-Backend Adaptation (Ascend / MLIR Ascend / MetaX / Hygon / Moore Threads)
- 2.1 Hygon: Async copy and cache swizzle lowering improvements enter review (09-28)
- 2.2 Ascend: Last week’s merged features reverted wholesale, main repo adds ascend-0930 branch (09-28/09-29)
- 2.3 MetaX / Moore Threads: Official repos silent, community-side S5000 performance tooling updated (09-28)
- 2.4 Community side: Sophgo BM1690 TPU pipeline adaptation opens new repository (09-28)
- III. Ecosystem and Adopters
- 3.1 Adopter: FlashQLA migrates to SM100 tcgen05, prepare_h measured 1.43x (09-28)
- 3.2 Adopter check: TileKernels no pushes, TileFoundry no in-window commits (09-29 check)
- 3.3 Release cadence: v0.1.14 at 27 days, TileRT release page 5 days behind (09-29)
- IV. Community, Tutorials, and Events
- 4.1 Community tutorials: tilelang-tutorials updates CPU examples (09-29)
- 4.2 Docs site: Two TileOPs docs site items with deletion-merge sync (09-28)
- 4.3 Media and academia: Google News zero hits, arXiv no new preprints (09-29 check)
- V. Trend Observations
- 5.1 The execution phase of contract-first: From spec-only to operator layer within 24 hours
- 5.2 Silent errors and the “kernel zone”: A concentrated reckoning of accumulated debt
- 5.3 Compiler soundness begins to interlock with upstream
- 5.4 Gaps and risk points
- Appendix: Sources and verification notes
Today’s Highlight: TileOPs Quantization Family Lands Implementation Layer First — Nine Operator Layers Merged Overnight, Six Sampling Operators Queued the Same Day (09-28/09-29)
Date: 2026-09-28, 2026-09-29 Source: Six Quantization Operator Layers (#2298) / INT8 Dequantization Operator Layers (#2297) / Sampling Operator Layers (#2299) / 200-Operator Tracking Issue (#2271)
Less than 24 hours after 24 spec-only entries and the 200-operator plan were opened late on 09-27, TileOPs began converting “contracts” into “operator layers” one by one, at the following pace:
- Six quantization operators merged first (#2298, merged 09-29 06:54, +798/-78, 16 files): INT8 per-tensor, per-channel, and per-block quantization, FP8 per-block quantization, INT4 per-group quantization, and SmoothQuant — six operator layers in total (named from
INT8QuantPerTensorFwdOptoSmoothQuantFwdOp) landed intileops.quantization; each haskernel_types = {}— calls first run generative checks, then raiseOpNotAvailableError(declared, pending implementation). AddedQuantizeCall(rows, cols, dtype, group_size); reference implementations migrated from test files intoworkloads/quantization.py, with tests changed to call the same math. The checklist also spells out engineering constraints: INT4 grouping and GEMM W4A16 requireK % 128 == 0(packing order defined by 128-element stride), INT4 group semantics given item by item (zero-point range [0, 15], scale floored at the smallest normal number, quantized following torch.ao and AWQ conventions); every divergence from vLLMscaled_int8_quant,per_token_group_quant_int8, and DeepGEMMper_block_cast_to_fp8is explicitly recorded. - Three INT8 dequantization operators followed (#2297, merged 09-29 06:20, +431/-29, 11 files): INT8 per-tensor, per-channel, and per-block dequantization operator layers (three
FwdOps includingINT8DequantPerTensorFwdOp) landed the same day, with entries marked as remaining spec-only; the reference implementation(q.float() * scale).to(out_dtype)matches torch’s three dequantize variants bit-for-bit; per-block handles the last block with scale shapeceil_div(K, 128), and divergences from the SGLang implementation are recorded. - Six sampling operators queued the same day (#2299, open, +1065/-151, 28 files): operator layers, workloads, and tests for
TopKMask,MinPMask,TopPMask,TopKTopPMask,SamplingFromProbs, andChainSpeculativeSamplingare all prepared in one go; reference implementations align with FlashInfer 0.6.16 (bf16/fp16/fp32 tiers), random numbers use Philox4x32-10 for deterministic integer-tensor implementation (aligned with Random123 known-answer vectors); and the boundary tie-handling of the old top-p reference was fixed (76 bf16 token results differ within 8 rows). - Context: Among the 24 spec-only entries opened in #2279 (9 quantization, 6 sampling, 7 MLA/KV, 1 each norm and linear attention), all 9 quantization operators have entered the operator layer, and the 6 sampling operators have entered PR review — “define contracts first, then land implementations” has entered a steady execution rhythm.
Interpretation: The benefit of operator layers coming first is that “interfaces, workloads, reference numerics, and rejection behavior” are frozen upfront, so kernel implementations can follow one by one without further changes to the external contract; for users, OpNotAvailableError provides a clear capability boundary. Note that none of this batch of operator layers includes kernel implementations, so no usable performance is produced before the first kernel is merged; follow-up focus points are the pace of kernel landing and the impact of shape constraints such as K % 128 on model integration.
I. Core Project Progress
1.1 Main Repo: TVM Tail Guardrail Fix — Inverted Modulo Remainder Bounds Could Cause Out-of-Bounds Reads, Merged After Real-Target Reproduction (09-28)
Date: 2026-09-28 Source: #3294 / TVM-side fix (#77)
- Problem: With dynamic reduction over
B * 192elements and 128 threads,(192 * B + 127) % 128was given an inverted interval[127, 63]by bounds analysis (correct should be[63, 127]). Based on this,tl.Simplifycould mistakenly provei * 128 + threadIdx.x < B * 192always true and delete the tail guardrail; when B = 1, the last 64 threads should skip the second round, but with the guardrail removed they would read beyond the next layer or even the last layer — actually observed on the aux-loss display kernel. - Fix: #3294 advances the
3rdparty/tvmsubmodule from907a88c879to4211874e9e(containing only this bounds fix and its TVM regression), and adds a portable regression case (directly constructing TIR, without assuming B > 0). Verification: 48 transform-related tests pass, 76 reducer v2 tests pass, 110 TVM-side tests pass (8 xfail); all 48 parameter combinations reproducing aux-loss on B300 pass. - Significance: A complete closed-loop example of “compiler soundness defect → real-target reproduction → upstream fix → submodule advance + regression backstop”; fix commit #77 in the same submodule repo (tile-ai/tvm) writes “modulo remainder bounds are ordered and reliable” into the TVM core.
1.2 Main Repo: Default Single-Config Build Adds -O2 — Lowering Geometric Mean 1.19x Faster, Library Size Halved (09-28)
Date: 2026-09-28 Source: #3191
- Change: When no explicit build type is given, single-config CMake builds add
-O2for TileLang object compilation; explicit Debug/Release and multi-config generators remain unchanged (1 file, +10/-0). - Measured (178 examples): total lowering time 1548.8 → 1376.0 seconds, geometric mean 1.186x, median 1.161x; 168/178 examples improved lowering by over 5%, none regressed by more than 5%; full kernel compilation geometric mean 1.053x. Side effect:
libtilelang.soshrinks from 43.5 MB to 18.9 MB. - Background: This change also speeds up the compiler performance regression cases from #3180 by about 5x (geometric mean), converging the “default parameters worse than manual parameters” issue in the compilation pipeline.
1.3 Main Repo Queue: Five New PRs and Three Issue Updates, 112 Open PRs (09-28)
Date: 2026-09-28 Source: Main repo PR list / #3296 / #3297 / #3299 / #3300 / #3293
- Five new PRs opened in the window (all small-step fixes): #3293 adds an unused binding eliminator to
tl.Simplify(with regressions for read-after-write, side effects, volatile, metadata references, etc.); #3296 fixes alignment and cache hints for bulk TMA copies (corresponding to defect issue #3295); #3297 supports uint8 sparse GEMM operands; #3299 fixes ties-to-even rounding forT.roundon ROCm; #3300 adds CUDA int32 popcount support. - Defects and requirements: Three issue updates in the window — #3295 (bulk TMA cache hint compilation failure, shared memory over-alignment reducing occupancy), #3292 (
T.copydefault coalesced_width striding issue on METAL), #3298 (apache-tvm-ffi 0.1.13 support request). - Queue temperature: 112 open PRs; the carryover queue from the previous period (four RNG PRs #3241~#3244, async staging #3279, TileIR backend #3247, magic division #3267, SIMT im2col #3275) saw no updates in the window; one older PR (#3187, vec integer bitwise_not crash fix) was closed without merging in the window.
1.4 TileOPs: Concentrated Fix of Silent-Error Kernels, Over a Dozen “Return Results Without Erroring” Cases Handled Together (09-29 Early Morning)
Date: 2026-09-28, 2026-09-29 Source: #2291
- Scale: A single PR with +2456/-1161 across 116 files, fixing over a dozen kernels at once that compute wrong results without erroring: MLA decode tail heads and empty cache, CountNonzero over 2^24, FloorDivide/Remainder/Div (floor and trunc semantics), MoePermuteAlign over 1024 experts, TopkSelector overflow, SSD d_state sharding, Conv “same” odd padding, FP8 dense prefill causal, GLA sharding, integer mean accumulator; MLA decode now returns zeros matching torch behavior under empty cache.
- Governance trio: ① SM90-specific cases switch to
@pytest.mark.sm90(26 functions, 5 files), manual capability checks retired — SKIP can no longer mask failures on SM90; ② kernels, benchmark timing, and warmup plugins no longer take configuration from environment variables (a new lint rule prohibits reading environment in source/workloads/benchmarks), instead using explicit parameters that enter the build identifier; ③ kernel limits converge from scattered locations into a single declaration in the “kernel region”, with out-of-bounds calls rejected as “no implementation serves this call” — configuration, rejection, and caching behavior are now consistently reproducible. - Supporting:
bench_kernelgains an event fallback switch; empty-output calls return empty; cases requiring CUDA get proper markers.
1.5 TileOPs: Attention Consolidation Continues — MHA and GQA Paged Decode Unified, Backward WS Kernel Merged (09-28)
Date: 2026-09-28 Source: #2296 / #2295 / #2281 / Issue #2293
- Decode unification (#2296, +588/-1156):
GQADecodePagedKernelnow serves both MHA and GQA paged decode operators, andMHADecodePagedKernelis deleted entirely; the MHA-side interface, manifest entries, and operator count remain unchanged; also fixes two issues: NaN when GQA reads non-finite values outside the cache, and masked keys still receiving weights whensm_scale=0. Test nodes 5047 → 5048; external benchmarks (H200, device_busy): four serving rows at 1.34~1.76x flashinfer. - WS kernel fix and expansion (#2295, +454/-425): The per-row guard introduced in #2273 had slowed four single-token rows by up to 20%, now restored by guarding edge blocks only at whole-block granularity; the warp-specialized kernel begins serving multi-query rows (constrained by
B*H*S_q*N_kv*D <= 2^27workload), reducing speculative decoding scenarios from 25.46 microseconds to 6.27 microseconds (2.37x versus FA3’s 14.88 microseconds), with four serving rows at 1.40~1.90x the control. - Backward kernel (#2281, +806/-486): MHA backward enters a new persistent warp-specialized kernel on SM90, speeding up calls with head dim 128 and 128-row key-block partitioning by 2.08~4.21x; launches per call 8 → 3; fixes two issues: WGMMA serialization in the pipeline and premature release of Q/dO staging;
MultiHeadAttentionBwdOpis deleted (merged into GQA backward directly serving MHA); benchmark scripts now count only backward (previously mixed forward and backward).
1.6 TileOPs: Six Batched Speedups for Reduction and Norm (09-28)
Date: 2026-09-28 Source: #2290 / #2292 / #2289 / #2287 / #2284 / #2286
- Reduction engine rebuild (#2290, +1436/-1294): Leading-axis prod switched to in-place reduction, no longer transposing inputs —
[4, 128, 4096](dim=0) drops from 469 microseconds to 3.3 microseconds (versus torch-compile 3.2 microseconds); adds a 16-byte streaming load header (L1/L2 evict-first cache policy); a single-register folded-row kernel serves sum/mean/amax/amin/prod and L1/L2/Inf norms together, with mlp-intermediate rows dropping from 28.6~30.1 microseconds to 24.8~25.1 microseconds (versus torch-compile 24.4~25.4). - The other five: #2292 logsumexp splits statistics across lane folding while preserving torch’s inf semantics; #2289 logical reduction folds once into registers at the input’s own byte width; #2287 InstanceNorm training mode converges from 20 launches to 1 (image-track-stats 27.28 → 2.62 microseconds), with rounding aligned to torch (previously off by 1 ulp); #2284 BatchNorm backward directly indexes the caller’s layout (resnet-stage3 15.5 → 4.9 microseconds, 1.86x), and fixes NaN caused by split statistics accumulating across blocks (shapes like (3, 5, 300, 301)); #2286 RMSNorm shares blocks for short rows, zero-padding only in registers.
1.7 TileOPs: Two Engineering PRs — Source Moved into csrc, Guard Exemption Cleanup (09-28)
Date: 2026-09-28 Source: #2285 / #2288
- #2285 moves all C++/CUDA source into
src/tileops/csrc(build layout normalized); #2288 removes the guard exemption list and fixes the issues it was masking — the failure surface is no longer diluted by a whitelist.
1.8 Nightly Snapshot: 1027 Correctness Items, 1335 Benchmark Items, Zero Failures Zero Errors (09-29 Early Morning)
Date: 2026-09-29 Source: Snapshot commit b7f5a1b4bd47 / Snapshot environment metadata
- One snapshot in the window (09-29 04:11, run 36462093045; corresponding to TileOPs
fa8acf71b9, i.e., the #2291 merge point): 1027 correctness items (2 skipped) zero failures zero errors; 1335 benchmark items (38 suites) zero failures zero errors. Environment: H200, CUDA 13.2, torch 2.13, tilelang 0.1.11+cu132. - Compared with the previous snapshot (09-28 early morning, 1093/1318): correctness count -66, benchmark +17 — the count changes accompany this week’s multiple test refactors and exemption cleanups, with both showing zero failures; the tracking criterion is “zero failures”, and counts do not directly represent coverage changes.
II. Multi-Backend Adaptation (Ascend / MLIR Ascend / MetaX / Hygon / Moore Threads)
2.1 Hygon: Async Copy and Cache Swizzle Lowering Improvements Enter Review (09-28)
Date: 2026-09-28 Source: hygon PR #13 / branch commit 4bf3ac58c4
- After merging #12 on 09-24, the Hygon line moved to branch development: the
feat/hcu-async-copy-cache-swizzlebranch gained a new commit (author Teng Huang, 16:29) improving async copy and cache swizzle lowering; the corresponding PR #13 was updated the same day and remains open. Hygon is the only domestic backend with code progress this period.
2.2 Ascend: Last Week’s Merged Feature Fully Reverted, Main Repo Adds ascend-0930 Branch (09-28/09-29)
Date: 2026-09-28, 2026-09-29 Source: revert PR #1841 / reverted feature #1829 / tilelang main repo
- At 09:44 a full revert was merged (+3621/-10333, 84 files), withdrawing #1829 (the “compiler-managed AscendC Vector mask and FP32 row reduction optimization” merged on 09-24, originally +10333/-3621) — the feature was entirely reverted 4 days after merging, with no revert rationale attached in the public record.
- Also: the main repo created the
ascend-0930branch in the early hours of 09-29 (author LeiWang1999) — an Ascend-related development signal, logged as an observation item.
2.3 MetaX / Moore Threads: Official Repos Silent, Community-Side S5000 Performance Tooling Updated (09-28)
Date: 2026-09-28 Source: tilelang-metax / tilelang-musa / Tilelang_musa community repo
- The official repos for MetaX (last push 09-24) and Moore Threads (09-17) saw no pushes during the window. On the community side there was activity in the Moore Threads direction:
Rankf/Tilelang_musaorganized TileSight and TileLang-MUSA source code and removed build artifacts (09-28 15:33); it describes itself as an “MTT S5000 operator-level performance analysis framework based on the DeepStack engine,” serving two scenarios — manual operator optimization and Agent operator evaluation — and includes 30 P0~P7 calibration documents, 64 scripts, and 24 benchmarks.
2.4 Community Side: Sophgo BM1690 TPU Pipeline Adaptation Opens a New Repo (09-28)
Date: 2026-09-28 Source: tilelang-tpu-bm1690-pipelines / migration and verification notes
- New repo (created 09-28, 8 commits within the window): building on existing tilelang-tpu development, it creates pipeline versions of six operator categories (Elementwise, RMSNorm, RoPE, SwiGLU, Matmul, FlashAttention) for SOPHGO BM1690 / SG2260E; it supports two programming models — BM1690 eight-core TPU-Kernel and SG2260E’s RV Tensor — with compilation, CModel, and PCIe execution running through a unified PPL 1.7 toolchain.
- The methodology is noteworthy: CModel only verifies numerical behavior, and “CModel wall-clock time is not BM1690 performance” is explicitly stated; each step produces evidence via a controlled runner (constrained by process count, a 120-second timeout, and a 4096 MiB memory cap); commits within the window cover FP16 matmul 1024³ CModel verification and pipelining of each operator (SSA-safe nested loops, double-buffered descriptors), etc. On-board measurements await a remote hardware environment becoming available.
III. Ecosystem and Adopters
3.1 Adopter: FlashQLA migrates to SM100 tcgen05, prepare_h measured at 1.43x (09-28)
Date: 2026-09-28 Source: PR #42 / FlashQLA repo
- Qwen’s FlashQLA (a TileLang-based GDN linear attention kernel library) merged optimizations targeting the Blackwell data-center tier:
prepare_h’s CP state-transition recurrence (M, Z) moved from the consumer-groupT.gemmto the MMA warp’s tcgen05 pipeline, with M resident in TMEM; register limits reconfigured from 168/160/160/24 to 152/104/104/152; whenstore_h=False, sync and the left-state half are published one beat earlier. - Measured (B200, 24 active CP rows):
prepare_hlatency 153.8 → 107.5 microseconds (1.43x, -30.1%); correctness gate 48/48 specializations bit-identical, forward integration 15/15 and 42/42 rows bit-identical. The adopter has advanced from SM90/103/120 adaptation to deep utilization of SM100 tcgen05.
3.2 Adopter check: TileKernels no pushes, TileFoundry no in-window commits (09-29 check)
Date: 2026-09-29 Source: TileKernels / TileFoundry
- DeepSeek TileKernels remains stalled at 2026-04; TileFoundry’s most recent push was early morning 09-28 (previous window, already reported as the AtomSched series), with no new commits in this window.
3.3 Release cadence: v0.1.14 at 27 days, TileRT release page lagging 5 days (09-29)
Date: 2026-09-29 Source: tilelang releases / TileRT releases / TileFoundry releases
- Main repo v0.1.14 (09-02) has gone 27 days without a new version; TileOPs has no standalone release; TileRT’s v0.1.6 release PR merged on 09-24 but the release page and tag remain at v0.1.5.post2 (5-day lag); TileFoundry’s latest remains v0.0.2 (09-10).
IV. Community, Tutorials, and Events
4.1 Community tutorial: tilelang-tutorials updates CPU examples (09-29)
Date: 2026-09-29 Source: tilelang-tutorials
- The community tutorial repo
easy-tilelang/tilelang-tutorialsupdated its CPU examples (09-29 00:17); the repo contains six notebook groups — 00_basic / 01_tvm / 02_tilelang / 03_tilelang_cpu / 04_tilelang_ascend / 05_tilelang_work — among which the CPU JIT and Ascend lowering/compile/JIT examples are useful references for getting started on the platform.
4.2 Docs site: two TileOPs docs-site commits synced with deletions/merges (09-28)
Date: 2026-09-28 Source: docs site PR #55 / PR #56
- Two commits: generate docs from the annotated
__all__and remove the deleted GQA operator (#55); remove the deletedMultiHeadAttentionBwdOpfrom the attention API page (#56) — synced with the deletion/merge actions in #2296/#2281.
4.3 Media and academia: zero Google News hits, no new arXiv preprints (09-29 check)
Date: 2026-09-29 Source: Google News / Hacker News / arXiv
- Multiple Google News RSS queries in Chinese and English (via proxy) yielded zero worthwhile items in the 24-hour window (the 9 candidates scraped were all GitHub repo pages or irrelevant content); Hacker News window items were unrelated to this topic; an arXiv search for “tilelang” still returns the latest as TileSight (a tile-centric GPU performance model) from 07-24, with no new preprints.
V. Trend Observations
5.1 The execution phase of contract-first: from spec-only to operator layer within 24 hours
Yesterday produced 24 spec-only entries and a 200-operator plan; today, quantifying 9 (merged) and sampling 6 (queued) already completes the full encapsulation of “operator layer + workload + reference numerics + rejection behavior,” with boundary semantics checked item by item against vLLM / FlashInfer / DeepGEMM. This shows the Manifest system has engineered contract checkpoints to the point of pipeline output; the next observation point is whether kernel implementations can deliver at the same cadence (this batch’s operator layer still has no kernels).
5.2 Silent errors and the “kernel region”: a concentrated governance debt
#2291 fixes a dozen-plus silent-wrong cases in one commit (from MLA tail heads to integer mean accumulators) — these are the most insidious: wrong results but no error thrown, tests pass, benchmarks normal. Viewed alongside “kernel limits converge to region declarations,” “environment variables exit the kernel path,” and “guard exemptions retired,” TileOPs is turning “under what conditions a kernel is correct” from scattered code into machine-checkable declarations; this is also the technical foundation enabling the 200-operator plan to make external coverage commitments.
5.3 Compiler soundness begins to link with upstream
The fix chain in #3294 (real-target reproduction → submodule repo fix #77 → submodule advance + portable regression) shows the main repo’s response path for “compiler silently changes semantics” defects has taken shape; for downstreams depending on TileLang (including various backend forks), the follow-up cadence of submodule-level fixes will directly determine the soundness level they obtain.
5.4 Gaps and risk points
First, multi-backend divergence: Hygon continues evolving on a branch, Ascend was reverted wholesale with no explanation, and MetaX and Moore Threads remain silent on the official side (Moore Threads now 12 days), widening the progress gap among domestic backends. Second, the attention backward line and decode line completed “deletion/merge” on the same day, but speculative decoding and similar scenarios still lag FA3 by a single-row 0.75x (the #2296 benchmark table). Third, the main repo has 112 open PRs, with the four RNG PRs, async staging, and TileIR and other large items untouched for days — digestion capacity and review bandwidth remain bottlenecks. Fourth, the main repo merged 2 and opened 5 in this window, and the “more opened than merged” pattern has not reversed.
Appendix: Sources and Verification Notes
| Source | Verification Result |
|---|---|
| tile-ai organization (29 repos) | 7 repos had pushes in the window: tilelang, TileOPs, TileOPs-nightly, tvm, tilelang-hygon, TileOPs.github.io, tilelang-ascend |
| Main repo tilelang | 2 merges (#3294, #3191); 5 new PRs opened (#3293, #3296, #3297, #3299, #3300); 112 open PRs total; bug reports #3292, #3295 and feature request #3298 updated within the window |
| TileOPs | 14 merges (in the #2281–#2298 range); additionally #2279 and #2283 landed 7–29 minutes after this window began (content covered in the previous report, not repeated here); 5 open PRs (#2172, #2294, #2299, #2300, #2301); issue #2293 closed along with #2296 |
| tvm (submodule repo in the same organization) | 1 merge (#77 modulo remainder bound fix, linked to main repo #3294) |
| TileOPs-nightly | 1 snapshot b7f5a1b4bd47: 1027 correctness items (2 skipped) with zero failures and zero errors; 1335 benchmark items with zero failures and zero errors |
| TileOPs.github.io | 2 PRs (#55, #56 for docs and deletion/merge sync) |
| Five domestic backend repos | Ascend: revert #1841 (withdrawing #1829); Hygon: branch commit + PR #13 update; MetaX (after 09-24), Moore Threads (after 09-17), MLIR Ascend (after 09-24) had no activity |
| Adopters and community | FlashQLA merged #42 (SM100 tcgen05); TileKernels had no pushes; community repos tilelang-tutorials, Tilelang_musa, bm1690-pipelines had updates |
| Google News / Hacker News / arXiv | Zero hits for Chinese and English queries; HN unrelated to this topic; arXiv had no new preprints (latest remains 07-24 TileSight) |
Complete Source List
- [1] Main repo TVM tail guard fix (#3294) — https://github.com/tile-ai/tilelang/pull/3294
- [2] TVM modulo remainder bound fix (#77) — https://github.com/tile-ai/tvm/pull/77
- [3] Main repo build -O2 (#3191) — https://github.com/tile-ai/tilelang/pull/3191
- [4] Main repo Simplify unused bindings (#3293) — https://github.com/tile-ai/tilelang/pull/3293
- [5] Main repo bulk TMA alignment fix (#3296) — https://github.com/tile-ai/tilelang/pull/3296
- [6] Main repo TMA bug report (#3295) — https://github.com/tile-ai/tilelang/issues/3295
- [7] Main repo uint8 sparse GEMM (#3297) — https://github.com/tile-ai/tilelang/pull/3297
- [8] Main repo ROCm rounding fix (#3299) — https://github.com/tile-ai/tilelang/pull/3299
- [9] Main repo int32 popcount (#3300) — https://github.com/tile-ai/tilelang/pull/3300
- [10] Main repo METAL bug report (#3292) — https://github.com/tile-ai/tilelang/issues/3292
- [11] Main repo tvm-ffi feature request (#3298) — https://github.com/tile-ai/tilelang/issues/3298
- [12] Main repo z3 thread inference (#3291) — https://github.com/tile-ai/tilelang/pull/3291
- [13] Main repo bitwise_not fix (#3187, closed without merge) — https://github.com/tile-ai/tilelang/pull/3187
- [14] Main repo PR list — https://github.com/tile-ai/tilelang/pulls
- [15] Main repo releases page — https://github.com/tile-ai/tilelang/releases
- [16] TileOPs six quantization operator layers (#2298) — https://github.com/tile-ai/TileOPs/pull/2298
- [17] TileOPs INT8 dequantization operator layer (#2297) — https://github.com/tile-ai/TileOPs/pull/2297
- [18] TileOPs sampling operator layer (#2299) — https://github.com/tile-ai/TileOPs/pull/2299
- [19] TileOPs kernel slot interface (#2300) — https://github.com/tile-ai/TileOPs/pull/2300
- [20] TileOPs varlen GQA hd512 (#2301) — https://github.com/tile-ai/TileOPs/pull/2301
- [21] TileOPs FP8 GQA leaf operator (#2294) — https://github.com/tile-ai/TileOPs/pull/2294
- [22] TileOPs GEMM epilogue (#2172) — https://github.com/tile-ai/TileOPs/pull/2172
- [23] TileOPs batch fix for silent errors (#2291) — https://github.com/tile-ai/TileOPs/pull/2291
- [24] TileOPs single-kernel paged decode (#2296) — https://github.com/tile-ai/TileOPs/pull/2296
- [25] TileOPs paged decode correction (#2295) — https://github.com/tile-ai/TileOPs/pull/2295
- [26] TileOPs MHA backward WS kernel (#2281) — https://github.com/tile-ai/TileOPs/pull/2281
- [27] TileOPs decode single-kernel issue (#2293) — https://github.com/tile-ai/TileOPs/issues/2293
- [28] TileOPs reduction in-place prod (#2290) — https://github.com/tile-ai/TileOPs/pull/2290
- [29] TileOPs logsumexp folding (#2292) — https://github.com/tile-ai/TileOPs/pull/2292
- [30] TileOPs logical reduction folding (#2289) — https://github.com/tile-ai/TileOPs/pull/2289
- [31] TileOPs InstanceNorm single launch (#2287) — https://github.com/tile-ai/TileOPs/pull/2287
- [32] TileOPs BatchNorm backward (#2284) — https://github.com/tile-ai/TileOPs/pull/2284
- [33] TileOPs RMSNorm short rows (#2286) — https://github.com/tile-ai/TileOPs/pull/2286
- [34] TileOPs csrc migration (#2285) — https://github.com/tile-ai/TileOPs/pull/2285
- [35] TileOPs guard exemption cleanup (#2288) — https://github.com/tile-ai/TileOPs/pull/2288
- [36] TileOPs calibration matching (#2283, previous content, landed this period) — https://github.com/tile-ai/TileOPs/pull/2283
- [37] TileOPs spec-only entry (#2279, previous content, landed this period) — https://github.com/tile-ai/TileOPs/pull/2279
- [38] TileOPs 200-operator tracking issue (#2271) — https://github.com/tile-ai/TileOPs/issues/2271
- [39] TileOPs PR list — https://github.com/tile-ai/TileOPs/pulls
- [40] Nightly snapshot commit (b7f5a1b4bd47) — https://github.com/tile-ai/TileOPs-nightly/commit/b7f5a1b4bd47
- [41] Snapshot environment metadata — https://github.com/tile-ai/TileOPs-nightly/blob/snapshots/meta.json
- [42] Hygon async copy PR (#13) — https://github.com/tile-ai/tilelang-hygon/pull/13
- [43] Hygon branch commit (4bf3ac58c4) — https://github.com/tile-ai/tilelang-hygon/commit/4bf3ac58c4
- [44] Ascend revert PR (#1841) — https://github.com/tile-ai/tilelang-ascend/pull/1841
- [45] Ascend reverted feature (#1829) — https://github.com/tile-ai/tilelang-ascend/pull/1829
- [46] MetaX repo — https://github.com/tile-ai/tilelang-metax
- [47] Moore Threads repo — https://github.com/tile-ai/tilelang-musa
- [48] MLIR Ascend repo — https://github.com/tile-ai/tilelang-mlir-ascend
- [49] Community S5000 profiling tool — https://github.com/Rankf/Tilelang_musa
- [50] Sophgo BM1690 pipeline repo — https://github.com/arcflute/tilelang-tpu-bm1690-pipelines
- [51] BM1690 migration and verification notes — https://github.com/arcflute/tilelang-tpu-bm1690-pipelines/blob/main/docs/bm1690-pipelines.md
- [52] tilelang-tpu upstream repo — https://github.com/xwhzz/tilelang-tpu
- [53] Community tutorials repo — https://github.com/easy-tilelang/tilelang-tutorials
- [54] FlashQLA SM100 optimization (#42) — https://github.com/QwenLM/FlashQLA/pull/42
- [55] FlashQLA repo — https://github.com/QwenLM/FlashQLA
- [56] TileKernels repo — https://github.com/deepseek-ai/TileKernels
- [57] TileRT releases page — https://github.com/tile-ai/TileRT/releases
- [58] TileFoundry releases page — https://github.com/tile-ai/TileFoundry/releases
- [59] TileOPs docs site PR one (#55) — https://github.com/tile-ai/TileOPs.github.io/pull/55
- [60] TileOPs docs site PR two (#56) — https://github.com/tile-ai/TileOPs.github.io/pull/56
- [61] TileOPs docs site — https://github.com/tile-ai/TileOPs.github.io
- [62] Google News RSS (Chinese and English queries, via proxy) — https://news.google.com/
- [63] Hacker News search — https://hn.algolia.com/
- [64] arXiv search — https://arxiv.org/