TileLang Daily Intelligence Report (2026-09-28)
Research window: Past 3 days (2026-09-25 07:00 ~ 2026-09-28 07:00, Beijing time; includes weekend period, last report was 09-25). This issue is a Monday three-period merged window, with no overlap with the previous issue. Sources: GitHub (push verification across 29 repos in the tile-ai organization, 6 repos had pushes within the window; TileOPs 35 merges and the full “Manifest parameterized signature” migration, 200-operator release plan opened; main repo SM120 block-scaled GEMM two merges and FP8 vector copy speedup benchmarked; TileFoundry AtomSched seven merges; all 5 nightly snapshots zero failures; adopter FlashQLA fixes and docs site updates), Google News RSS multi-query in Chinese and English (via proxy, zero hits), Hacker News, arXiv, community repo verification
Issue Index
- Today’s Highlights: TileOPs “Manifest parameterized signature” full migration complete—17 merges, 200-operator release plan opened the same day (09-26/09-27)
- I. Core Project Progress
- 1.1 Main repo: SM120 block-scaled GEMM two merges—fragment residency support and layout conflict rejection (09-26/09-27)
- 1.2 Main repo: FP8 vector copy retains packing, conditional copy kernel benchmarked at 1.42x (09-25)
- 1.3 Main repo: ten new PRs opened within the window, RNG queue and async staging refactor still pending review (09-25~09-27)
- 1.4 TileOPs: attention consolidation—SM90 prefill triple and decode unification, legacy paths removed (09-25~09-28)
- 1.5 TileOPs: linear attention, FFT rewrite, and operator misc merges (09-25~09-28)
- 1.6 TileOPs: test and benchmark system refactor, manifest consistency cases landed (09-26/09-27)
- 1.7 TileFoundry: AtomSched—instructionized tensor scheduling and staged memory analysis, seven merges (09-25~09-28)
- 1.8 Nightly snapshots: all 5 within the window zero failures, latest 1093 correctness + 1318 benchmark items (09-28 early morning)
- II. Multi-Backend Adaptation (Ascend / MLIR Ascend / MetaX / Hygon / Moore Threads)
- 2.1 Five repos zero pushes, zero PR updates within the window (09-28 verification)
- III. Ecosystem and Adopters
- 3.1 Adopter: FlashQLA fixes Blackwell path kernel synchronization (09-26)
- 3.2 Adopters and community: TileKernels no pushes, SoftHier extensions no updates within the window (09-28 verification)
- 3.3 Release cadence: main repo v0.1.14 at 26 days, TileRT release page lagging 4 days (09-28)
- IV. Community, Tutorials, and Events
- 4.1 Docs site: TileOPs docs site three consecutive updates following Manifest migration; main repo docs site routine updates (09-26/09-27)
- 4.2 Media and academia: Google News zero hits, arXiv no new preprints (09-28 verification)
- V. Trend Observations
- 5.1 Institutionalizing “contract first”: from roofline self-checks to release plans
- 5.2 Small-step fast iteration on the SM120 line and lowering-only test strategy
- 5.3 Attention kernels: from multi-path parallelism to a single lineage
- 5.4 Gaps and risk points
- Appendix: Materials and verification notes
Today’s Highlight: TileOPs “Manifest Parametric Signature” Full Migration Complete — 17 PRs Merged, 200-Operator Release Plan Opened the Same Day (09-26/09-27)
Date: 2026-09-26, 2026-09-27 Source: Design doc PR (#2209) / Signature checking and generation (#2212) / 200-operator tracking issue (#2271) / 24 spec-only entries (#2279)
From 09-26 08:02 to 09-27 11:57 (Beijing time), 17 PRs in the TileOPs Manifest series were merged in succession, rewriting the operator inventory from YAML tables into “parametric signatures,” and opening a 200-operator release plan in the same window:
- Design (#2209, merged 09-26 08:02): Each operator is written in the manifest as a single parametric signature —
forallindices and kinds, type families and algebraic data types (ADT), a closed expression language, read/write effects, workload rows with generators, and a validation checklist; roofline pricing is changed to evaluating inline expressions overix; operator design docs, the trust model, testing guidelines, and agent domain rules are rewritten in sync (13 files, +633/-1193, more deletions than additions). - Implementation (#2212 to #2227, throughout 09-26 to midday 09-27): Signature checking and operator method generation are changed to be derived from the manifest (#2212); operator families are converted in batches — elementwise (#2216), moe (#2218), reduction/normalization (#2221), linear_attention (#2222), gemm/quantization/convolution (#2223), attention (#2224), mamba (#2214), pool/scan/rope/sequence-modeling/FFT (#2217); the closing PR (#2225, 09-27 10:32) finishes converting the last batch of gemm entries and removes the legacy manifest path; manifest YAML is moved into the
spec/directory with unified file naming (#2227). Accompanying behavior change: calling an operator on a device with no in-tree kernel declaration now throwsOpNotAvailableError(#2219). - Release plan (evening of 09-27): The tracking issue “TileOPs 200-operator release plan” (#2271) is opened — targeting 200 manifest entries (counted as implemented + spec-only; new operators only, extensions and scheduling excluded), with a worst-case projection table; the same day, PR #2279 adds 24 spec-only entries and a new
samplingfamily to the plan, and per that PR, the repo count badge goes from 172/177 to 172/201. The 24 new entries cover four directions: 9 quantization entries (INT8/INT4 per-tensor, per-channel, per-block, and SmoothQuant, FP8 per-block), 6 sampling and speculative decoding entries (TopK/TopP/MinP/TopKTopP masks, SamplingFromProbs, ChainSpeculativeSampling), 7 MLA and KV cache entries (MLA Paged/Varlen, PagedKVCache write and gather, MergeAttentionStates, DeepSeekSparseAttentionPaged), 1 norm entry (FusedQKNormRope), and 1 linear attention entry (KimiDeltaAttention). - Accompanying constraints (#2279): A new
test_spec_reference.pyruns reference implementations per workload for spec-only entries and validates input writes and rejected calls; the invariant “implemented ⊆ public ⊆ manifest” is established; the issue also lists 5 spec-only completion trackers (two GQA, GatedDeltaNet, DeltaNetInference, GLAInference) and 11 kernel coverage subtasks (FP8 Q/K/V, RoPE, FP8 KV cache, GDN variants, etc.).
Interpretation: This is an institutional upgrade of TileOPs’ “trust engineering” over the past two weeks (following roofline self-checks and formula corrections) — contracts first, then implementation, with checking taken over by generators; the 200-operator plan makes the roadmap public, bringing operator directions for domestic large models such as quantization, speculative decoding, MLA, and KimiDeltaAttention to the surface. For users, the “source of contract” for operator behavior shifts from code comments to a verifiable manifest, allowing coverage gaps to be assessed before integration.
I. Core Project Progress
1.1 Main Repo: Two SM120 Block-Scaled GEMM Merges — Fragment Residency Support and Layout Conflict Rejection (09-26/09-27)
Date: 2026-09-26, 2026-09-27 Source: #3257 / #3284 / Issue #3283 / Preceding feature request #3256
- #3257 (merged 09-26 23:39, author sepcnt): Adds fragment-resident A operand and row-major scale fragment support for SM120 block-scaled GEMM, and fixes compact scale row selection under odd per-warp MMA atom grids (including 8x1 partitioning); 5 files, +230/-50; includes regression coverage for fragment operands/scales, transposed A, compact layouts, and multi-warp strategies.
- #3284 (merged 09-27 19:18, author LeiWang1999): Fixes the issue where SFA and SFB sharing the same scale fragment silently overwrites layout requirements, causing incorrect scale reads — layout inference now compares full fragment layouts first, and on conflict raises an error suggesting separate scale fragments or shared-memory scales; adds 18 CUDA lowering regression cases (three warp strategies, shared/fragment A, aliased and separate scale fragments, aliased shared-memory scales), with 138 tests passing and 21 skipped across the two related test files.
- Testing scope note: Both PRs explicitly lower to
sm_120awithout device compilation (SM120 hardware numerical execution is not covered by this batch of tests; explicitly noted in #3284). - Background: SM120 is the consumer/workstation tier of Blackwell (GeForce RTX 50 series, RTX PRO series); this line has continued from the 09-16 mxf4nvf4 block-scaled MMA PR (#3081, still open), with two more follow-ups added within the window (#3286, #3290, see 1.3).
1.2 Main Repo: FP8 Vector Copy Stays Packed, Conditional Copy Kernel Measured at 1.42x (09-25)
Date: 2026-09-25 Source: #3276
- #3276 (merged 09-25 14:17, author LeiWang1999; recorded as “under review” in the previous issue): Defines packed copy constructors and assignments for CUDA FP8 vector types (E4M3/E5M2/E8M0, 2/4/8/16/32 lanes) — whole-object moves via matched integer carriers (uint16/uint32/uint2/uint4/ulonglong4). Motivation: NVCC lowers default member-wise copies into redundant unpack/repack instructions even when surrounding memory accesses are already vectorized; the fix introduces no codegen special cases, only header changes.
- Measured (B300 SXM6,
sm_103a, NVCC 13.1): E4M3 conditional copy kernel 44.07 → 31.07 microseconds (1.42x); registers 77 → 71; removes 80 PRMT, 16 SHF, and 32 packing-related LOP3 instructions in conversion workloads, with E8M0 PRMT dropping from 241 to 155. Timing uses do_bench (warmup 100 / rep 500, median of five alternating runs); static SASS comparison and timing results are presented separately.
1.3 Main Repo: Ten New PRs Opened in the Window, RNG Queue and Async Staging Refactor Still Pending Review (09-25~09-27)
Date: 2026-09-25 ~ 2026-09-27 Source: #3286 / #3290 / #3285 / #3288 / #3280 / Main repo PR queue
- Ten new PRs opened in the window: two SM120 PRs (#3286 K-atom unrolling and per-atom scale loads for register A; #3290 W4A16 NVFP4 decode example, draft), four CUDA/compiler PRs (#3285 NVRTC device compilation support, TVM-FFI; #3288 preserve warp-specialized stages until consumed by moved non-TMA copies — fixes #3287; #3282 TIRX source range visualization tool; #3291 use z3 to infer all threads satisfying constraints), two platform PRs (#3281 CPU/Windows JIT shared library compilation and export fix; #3289 add missing newline to imported C source files), and one fix (#3280 deepseek_v32 topk selector buffer overflow).
- Ongoing queue items: four RNG PRs (#3241~#3244) still open with no updates since 09-24; #3275 (SIMT im2col early exposure) and #3279 (async data staging unification and transfer-aware lowering) pending review with no activity; #3267 (magic division) and #3247 (TileIR backend, draft) not updated.
- Issue side: two new issues opened in the window — #3283 (SM120 scale fragment conflict, fixed and closed by #3284) and #3287 (warp-specialized pipeline releasing stages too early), the latter followed up by #3288.
1.4 TileOPs: Attention Consolidation — SM90 Prefill Trilogy and Decode Unification, Legacy Path Removal (09-25~09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: #2178 / #2208 / #2198 / #2207 / #2202 / #2197 / #2273 / #2281
- SM90 varlen GQA prefill trilogy: warp-specialized prefill kernel merged (#2178, 09-25 19:12) → converted to persistent kernel (#2208, 09-26 00:50) → D64 head dimension support (#2198, 09-27 23:19); concurrently restores the TMA path for the sliding-window varlen kernel (#2207, 09-25 18:37).
- Decode-side convergence: paged GQA decode migrated to the unified operator with legacy decode operator removed (#2202, 09-27 21:32); legacy varlen GQA path removed (#2197, 09-27 20:50).
- Correctness fixes: causal paged decode aligned to cache end, non-causal NSA truncated at sequence end, FusedTopK priced by score (#2273, 09-27 21:38).
- New follow-up: SM90 warp-specialized kernel for MHA backward service-ization and separate FA3 backward timing (#2281, draft in the early hours of 09-28) — a prelude to the backward path entering the same kernel lineage.
1.5 TileOPs: Linear Attention, FFT Rewrite, and Operator Misc Merges (09-25~09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: #2199 / #2196 / #2187 / #2040 / #2283 / #2242 / #2240 / #2280
- Linear attention: GLA decode integrated into
GLAInferenceFwdOp(#2199, merged 09-28 06:20, H200 dense decode service-ization); unowned Gated DeltaNet kernel removed (#2196, 09-27 22:57) — continuing last week’s “consolidation” approach. - FFT: #2187 replaces the LUT implementation with a computational kernel, extending power-of-two coverage to 2^28 (merged 09-25 13:03; closes the previous issue’s “under review” item).
- GEMM and element-wise:
GemmFp8FwdOpsupports 1D2D FP8 block scaling (#2040, 09-28 06:19; #2283 follows up with calibration matched bycalibration_key); element-wise NaN propagation removes clamp-family slowdowns and restores binary benchmark shapes (#2242); float32 reduction/softmax switched to single-kernel read correction (#2240). - Architecture maintenance: Engram decode compilation boundary moved to the operator layer (#2280, 09-28 06:10).
1.6 TileOPs: Test and Benchmark Infrastructure Refactor, Manifest Consistency Cases Landed (09-26/09-27)
Date: 2026-09-26, 2026-09-27 Source: #2239 / #2226 / #2243 / #2244 / #2213 / #2215
- Structural reorganization: merge fragmented tests/ops and workloads files (#2239); merge fragmented benchmark files into a single-import module (#2226); single source of truth for test and workload devices (#2243); unified markers for in-tree kernel tests, benchmarks gain target and device options (#2244); op tests stop importing benchmarks, unified standard tolerances, trust-model.md renamed to layer-boundaries (#2213).
- Contract tests: build target consistency cases invoked from the manifest (#2215) — paired with today’s focus on the Manifest migration.
- Note: This refactor group is the direct backdrop for the nightly snapshot case structure changes; the snapshot pipeline remains zero-failure (see 1.8).
1.7 TileFoundry: AtomSched — Instructionized Tensor Scheduling and Staged Memory Analysis, Seven PRs (09-25~09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: #189 / #191 / #187 / #186
- Seven merges (#185~#191): shard layout declarations (#185), AtomSched operation declarations (#186), CUDA MMA atom and async copy instruction declarations (#187), structured regex matching for layouts (#188), instructionized tensor scheduling (#189), TIR operand contracts (#190), staged memory allocation model (#191).
- #189 highlights (merged 09-27 21:08): Adds the
tf.scheduleoperation (five author parameters: operands/op/repeat/order/buffers) and 21 scheduling fixtures; operand contracts and result types derived from the selected instruction; cost accounting by hierarchy bandwidth, MMA work, and buffer residency; fail-closed when the evaluator is missing a handler (no fallback to identity). Scale: 64 files, +3445/-220. - #191 highlights (merged 09-28 03:31): Completes AtomSched analysis migration stage 3 — loop-aware residency, region-level register peaks, global/shared memory placement solving with addresses (requiring aliasing to hold); 21/21 shared-memory golden matches; source-mode validation 1241 passing, 10 skipped.
1.8 Nightly Snapshots: All 5 in the Window Zero-Failure, Latest 1093 Correctness + 1318 Benchmark (Early 09-28)
Date: 2026-09-25 ~ 2026-09-28 Source: Latest snapshot commit / Snapshot environment metadata
- 5 snapshots in the window (from the evening of 09-25 to the early hours of 09-28); the latest
dab9b1ab(corresponding to TileOPsee3ebaf247, i.e., the #2198 merge point): 1093 correctness tests, 2 skipped, zero failures and zero errors; 1318 benchmarks, 1 skipped, zero failures and zero errors. - Environment (meta.json): NVIDIA H200, CUDA 13.2 / driver 595.71.05, tilelang 0.1.11+cu132 (git afcebed1), torch 2.13.0+cu132, power limit 700W.
II. Multi-Backend Adaptation (Ascend / MLIR Ascend / MetaX / Hygon / Moore Threads)
2.1 All Five Repos: Zero Pushes, Zero PR Updates in the Window (09-28 Check)
Date: 2026-09-28 Source: tilelang-ascend / tilelang-metax / tilelang-hygon / tilelang-musa / tilelang-mlir-ascend
- Over the 3-day window, all five repos had no pushes and no PR updates: Ascend (last push 09-24), MLIR Ascend (09-24), MetaX (09-24), Hygon (09-24), Moore Threads (09-17) — checking each repo’s recent PR list individually, the update count within the window is 0 for all.
- Compared with the previous issue: no new merges in the Ascend documentation verification queue; no new PRs after MetaX sub-warp guard (#159) and Hygon #12; Moore Threads silent for 11 consecutive days since 09-17. Substantive hardware-oriented progress in the window is concentrated in the main repo’s NVIDIA line (SM120, see 1.1); none of the four domestic adaptations saw activity this period.
III. Ecosystem and Adopters
3.1 Adopters: FlashQLA fixes kernel synchronization on the Blackwell path (09-26)
Date: 2026-09-26 Source: FlashQLA fix merged / FlashQLA repo
- Qwen’s FlashQLA (a TileLang-based GDN linear attention kernel library supporting SM90/100/103/120/121) merged a fix: the
prepare_handfused_fwdkernels now add synchronization points on Blackwell (3 files, +21/-5, submitted by Chengruidong Zhang). This is the latest move by an adopter to keep polishing on the Blackwell platform; per the repo’s own description, its v0.1.2 already provides forward support for SM120 and serves as the GDN backend for flash-linear-attention.
3.2 Adopters and community: no pushes to TileKernels, no in-window updates to SoftHier extension (09-28 check)
Date: 2026-09-28 Source: TileKernels repo / tilelang-softhier repo
- DeepSeek’s TileKernels has had no pushes since 2026-04; ETH Zurich’s TileLang SoftHier extension (targeting RISC-V many-core accelerators) was last updated on 09-24 (reported in the previous issue), with no new commits in the window; community materials such as the TVM-TileLang cookbook have no new content.
3.3 Release cadence: main repo v0.1.14 now 26 days old, TileRT release page lags by 4 days (09-28)
Date: 2026-09-28 Source: tilelang releases / TileRT releases / TileFoundry releases
- Main repo tilelang: v0.1.14 (09-02) has now gone 26 days without a new version; TileOPs has no standalone release; TileFoundry’s latest remains v0.0.2 (09-10, 18 days).
- TileRT: the v0.1.6 release PR was merged on 09-24, but GitHub Releases and tags remain at v0.1.5.post2 — the public version has lagged for 4 days, and the repo saw no new commits in the window.
IV. Community, Tutorials, and Events
4.1 Documentation sites: TileOPs docs site lands three updates following the Manifest migration; main repo docs site gets routine update (09-26/09-27)
Date: 2026-09-26, 2026-09-27 Source: TileOPs docs site / tilelang docs site
- TileOPs docs site, three commits: trust-model.md rename follow-up (#52), Manifest migration follow-up (#53), benchmark page grouped by workload while preserving manifest order (#54) — the docs side stays in sync with the upstream contract refactor.
- Main repo docs site, one routine
Update docs(early morning 09-27), an automated regeneration with no content-level changes.
4.2 Media and academia: zero hits on Google News, no new arXiv preprints (09-28 check)
Date: 2026-09-28 Source: Google News / Hacker News / arXiv
- Multiple Chinese and English Google News RSS queries (via proxy) returned zero hits within the 72-hour window; Hacker News items in the window are unrelated to this topic; a full arXiv search for “tilelang” still shows the latest paper as TileSight (a tile-centric GPU performance model) from 07-24, with no new preprints in the window.
V. Trend Observations
5.1 Institutionalizing “contract first”: from roofline self-checks to release plans
The three-step path is fully on display in this window: the 09-20 roofline checklist self-audit (which formulas have never been run) → the 09-26/09-27 full migration of parameterized signatures (the contract becomes the single source for checks and generation) → the 09-27 200-operator release plan (spec-only contracts land first, with counts and coverage publicly tracked). Engineering implication: TileOPs is shifting from “implementation-driven” to “contract-driven” release governance — only once a checklist can answer “what exists, what’s missing, and how soon it can be merged” does version planning become possible.
5.2 Small, rapid steps on the SM120 line and a lowering-only test strategy
Requirement ticket (#3256) → patch (#3257) → hardening (#3283/#3284) → follow-up (#3286/#3290), completing four stages within a week; and regression tests run without hardware by explicitly lowering to sm_120a. Coverage of block-scaled low-precision GEMM for consumer/workstation Blackwell is starting to take shape, forming a dual track with the data-center line (measured FP8 speedups on B300/sm_103a).
5.3 Attention kernels: from parallel multi-path to a single lineage
Three consecutive SM90 varlen GQA prefill commits (including a persistent-kernel overhaul), unified paged decode, and removal of the legacy varlen path — the TileOPs attention layer completed “establish the new + retire the old” within a single window; linear attention (GLA/GDN) was consolidated over the same period, and a follow-up draft for the backward path (MHA backward) has been opened. Together with the MLA and DeepSeek sparse attention entries in the 200-operator plan, the intent to converge the attention operator lineage toward “one canonical path per operator” is clear.
5.4 Gaps and risks
First, the four domestic backends and MLIR Ascend saw zero activity in the window, with Moore Threads silent for 11 days — the uneven distribution of activity remains unchanged; second, the TileRT release page lags its release PR by 4 days, leaving blind spots in version tracking; third, the main repo’s RNG queue has gone 4 days without updates, and the slow review pattern for large PRs (TileIR backend #3247, async staging #3279) persists; fourth, the main repo merged 3 PRs in 72 hours (including the weekend) against 10 newly opened, so queue throughput remains to be seen.
Appendix: Sources and Verification Notes
| Source | Verification Result |
|---|---|
| tile-ai organization (29 repos) | 6 repos had pushes within the window: tilelang, TileOPs, TileOPs-nightly, TileFoundry, tilelang.github.io, TileOPs.github.io |
| Main repo tilelang | 3 merges (#3276, #3257, #3284); 10 new PRs opened; bug #3283 closed, #3287 newly opened; RNG queue and #3275, #3279, #3267, #3247 still open |
| TileOPs | 35 merges (including 17 in the Manifest series); 30 new PRs opened; #2210, #2211 closed without merge |
| TileRT | Releases page and tags still at v0.1.5.post2 (v0.1.6 PR merged 4 days ago); no commits within the window |
| TileOPs-nightly | 5 snapshots: latest dab9b1ab correctness 1093 items, zero failures, zero errors (2 skipped); benchmark 1318 items, zero failures, zero errors (1 skipped) |
| TileFoundry | 7 merges (#185–#191) |
| Five domestic backend repos | tilelang-ascend / tilelang-metax / tilelang-hygon / tilelang-musa / tilelang-mlir-ascend: zero pushes, zero PR updates within the window |
| Adopters and community | FlashQLA merged Blackwell sync fix; TileKernels no pushes; SoftHier extension no updates within the window |
| Google News / Hacker News / arXiv | Zero hits for Chinese and English queries; no relevant HN entries; no new arXiv preprints (latest still 07-24 TileSight) |
Full Source List
- [1] Main repo FP8 vector copy merge (#3276) — https://github.com/tile-ai/tilelang/pull/3276
- [2] Main repo SM120 block-scaled fragment support (#3257) — https://github.com/tile-ai/tilelang/pull/3257
- [3] Main repo SM120 layout conflict rejection (#3284) — https://github.com/tile-ai/tilelang/pull/3284
- [4] Main repo SM120 fragment conflict bug (#3283) — https://github.com/tile-ai/tilelang/issues/3283
- [5] Main repo SM120 pilot requirement (#3256) — https://github.com/tile-ai/tilelang/issues/3256
- [6] Main repo SM120 register A K atomic PR (#3286) — https://github.com/tile-ai/tilelang/pull/3286
- [7] Main repo SM120 W4A16 NVFP4 example PR (#3290) — https://github.com/tile-ai/tilelang/pull/3290
- [8] Main repo mxf4nvf4 block-scaled MMA PR (#3081) — https://github.com/tile-ai/tilelang/pull/3081
- [9] Main repo warp specialization staging bug (#3287) — https://github.com/tile-ai/tilelang/issues/3287
- [10] Main repo warp specialization phase fix PR (#3288) — https://github.com/tile-ai/tilelang/pull/3288
- [11] Main repo NVRTC device compilation PR (#3285) — https://github.com/tile-ai/tilelang/pull/3285
- [12] Main repo Windows JIT fix PR (#3281) — https://github.com/tile-ai/tilelang/pull/3281
- [13] Main repo deepseek_v32 topk fix PR (#3280) — https://github.com/tile-ai/tilelang/pull/3280
- [14] Main repo TIRX source range visualization PR (#3282) — https://github.com/tile-ai/tilelang/pull/3282
- [15] Main repo z3 thread inference PR (#3291) — https://github.com/tile-ai/tilelang/pull/3291
- [16] Main repo RNG sequence PR (#3242, see also #3241/#3243/#3244) — https://github.com/tile-ai/tilelang/pull/3242
- [17] Main repo async staging unification PR (#3279) — https://github.com/tile-ai/tilelang/pull/3279
- [18] Main repo SIMT im2col PR (#3275) — https://github.com/tile-ai/tilelang/pull/3275
- [19] Main repo magic division PR (#3267) — https://github.com/tile-ai/tilelang/pull/3267
- [20] Main repo TileIR backend PR (#3247) — https://github.com/tile-ai/tilelang/pull/3247
- [21] Main repo releases page — https://github.com/tile-ai/tilelang/releases
- [22] TileOPs Manifest design PR (#2209) — https://github.com/tile-ai/TileOPs/pull/2209
- [23] TileOPs signature check and generation PR (#2212) — https://github.com/tile-ai/TileOPs/pull/2212
- [24] TileOPs elementwise conversion (#2216) — https://github.com/tile-ai/TileOPs/pull/2216
- [25] TileOPs moe conversion (#2218) — https://github.com/tile-ai/TileOPs/pull/2218
- [26] TileOPs reduction/normalization conversion (#2221) — https://github.com/tile-ai/TileOPs/pull/2221
- [27] TileOPs linear_attention conversion (#2222) — https://github.com/tile-ai/TileOPs/pull/2222
- [28] TileOPs gemm/quantization/convolution conversion (#2223) — https://github.com/tile-ai/TileOPs/pull/2223
- [29] TileOPs attention conversion (#2224) — https://github.com/tile-ai/TileOPs/pull/2224
- [30] TileOPs mamba conversion (#2214) — https://github.com/tile-ai/TileOPs/pull/2214
- [31] TileOPs remaining family conversions (#2217) — https://github.com/tile-ai/TileOPs/pull/2217
- [32] TileOPs wrap-up and legacy path removal (#2225) — https://github.com/tile-ai/TileOPs/pull/2225
- [33] TileOPs manifest migration into spec/ (#2227) — https://github.com/tile-ai/TileOPs/pull/2227
- [34] TileOPs device rejection behavior (#2219) — https://github.com/tile-ai/TileOPs/pull/2219
- [35] TileOPs 200-operator tracking issue (#2271) — https://github.com/tile-ai/TileOPs/issues/2271
- [36] TileOPs 24 spec-only entries (#2279) — https://github.com/tile-ai/TileOPs/pull/2279
- [37] TileOPs SM90 WS prefill (#2178) — https://github.com/tile-ai/TileOPs/pull/2178
- [38] TileOPs prefill persistence (#2208) — https://github.com/tile-ai/TileOPs/pull/2208
- [39] TileOPs D64 support (#2198) — https://github.com/tile-ai/TileOPs/pull/2198
- [40] TileOPs sliding window TMA (#2207) — https://github.com/tile-ai/TileOPs/pull/2207
- [41] TileOPs paged decode unification (#2202) — https://github.com/tile-ai/TileOPs/pull/2202
- [42] TileOPs legacy varlen removal (#2197) — https://github.com/tile-ai/TileOPs/pull/2197
- [43] TileOPs decode fix (#2273) — https://github.com/tile-ai/TileOPs/pull/2273
- [44] TileOPs MHA backward WS kernel (#2281) — https://github.com/tile-ai/TileOPs/pull/2281
- [45] TileOPs GLA decode serving (#2199) — https://github.com/tile-ai/TileOPs/pull/2199
- [46] TileOPs GDN kernel removal (#2196) — https://github.com/tile-ai/TileOPs/pull/2196
- [47] TileOPs FFT rewrite (#2187) — https://github.com/tile-ai/TileOPs/pull/2187
- [48] TileOPs FP8 block scaling (#2040) — https://github.com/tile-ai/TileOPs/pull/2040
- [49] TileOPs calibration matching (#2283) — https://github.com/tile-ai/TileOPs/pull/2283
- [50] TileOPs NaN speedup (#2242) — https://github.com/tile-ai/TileOPs/pull/2242
- [51] TileOPs reduction fix (#2240) — https://github.com/tile-ai/TileOPs/pull/2240
- [52] TileOPs Engram boundary (#2280) — https://github.com/tile-ai/TileOPs/pull/2280
- [53] TileOPs test decoupling (#2213) — https://github.com/tile-ai/TileOPs/pull/2213
- [54] TileOPs test merge (#2239) — https://github.com/tile-ai/TileOPs/pull/2239
- [55] TileOPs benchmark merge (#2226) — https://github.com/tile-ai/TileOPs/pull/2226
- [56] TileOPs device single source (#2243) — https://github.com/tile-ai/TileOPs/pull/2243
- [57] TileOPs benchmark options (#2244) — https://github.com/tile-ai/TileOPs/pull/2244
- [58] TileOPs manifest consistency cases (#2215) — https://github.com/tile-ai/TileOPs/pull/2215
- [59] TileFoundry instruction-based scheduling (#189) — https://github.com/tile-ai/TileFoundry/pull/189
- [60] TileFoundry phased memory (#191) — https://github.com/tile-ai/TileFoundry/pull/191
- [61] TileFoundry CUDA atomic declarations (#187) — https://github.com/tile-ai/TileFoundry/pull/187
- [62] TileFoundry operation declarations (#186) — https://github.com/tile-ai/TileFoundry/pull/186
- [63] Nightly snapshot commit (dab9b1ab) — https://github.com/tile-ai/TileOPs-nightly/commit/dab9b1ab535a8116480198369641128c563aacc8
- [64] Snapshot environment metadata — https://github.com/tile-ai/TileOPs-nightly/blob/snapshots/meta.json
- [65] FlashQLA Blackwell fix commit — https://github.com/QwenLM/FlashQLA/commit/cdfcb99061daa61d78aa07bb254a6cf9ddfa8450
- [66] TileRT releases page — https://github.com/tile-ai/TileRT/releases
- [67] TileFoundry releases page — https://github.com/tile-ai/TileFoundry/releases
- [68] TileOPs docs site — https://github.com/tile-ai/TileOPs.github.io
- [69] tilelang docs site — https://github.com/tile-ai/tilelang.github.io
- [70] TileKernels repo — https://github.com/deepseek-ai/TileKernels
- [71] FlashQLA repo — https://github.com/QwenLM/FlashQLA
- [72] tilelang-softhier repo — https://github.com/HaozeG/tilelang-softhier
- [73] Google News RSS (Chinese and English queries, via proxy) — https://news.google.com/
- [74] Hacker News search — https://hn.algolia.com/
- [75] arXiv search — https://arxiv.org/