Conversation
Cojo01
force-pushed
the
512-linalg-rewrites
branch
3 times, most recently
from
July 12, 2026 21:42
8f538af to
bd752de
Compare
Cojo01
force-pushed
the
512-linalg-rewrites
branch
from
July 14, 2026 12:58
db5677a to
cc78404
Compare
Cojo01
force-pushed
the
512-linalg-rewrites
branch
2 times, most recently
from
July 18, 2026 21:19
a1fd6b8 to
cb20bd6
Compare
Cojo01
force-pushed
the
512-linalg-rewrites
branch
from
July 19, 2026 12:24
087fbbe to
4d5786d
Compare
Cojo01
force-pushed
the
512-linalg-rewrites
branch
from
July 19, 2026 13:14
4d5786d to
e857aee
Compare
Cojo01
marked this pull request as ready for review
July 19, 2026 16:23
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fourth of four PRs for #512. It adds a set of linear-algebra rewrites, building on the earlier three PRs.
The four PRs:
ScalarConstantFoldableinterface and the algebraic-trait vocabulary; moving existing rewrites onto it and adding new trait-based rewrites.The rewrites
All hand-written
OpRewritePatterns in a newsrc/ir/daphneir/LinearAlgebraRewrites.{cpp,h}, registered into the existing--daphne-algebraic-simplifypass. No new pass, CLI flag, orPasses.tdchange.sum(diagVector(X @ Y))→sum(X * t(Y)). DAPHNE has notracebuiltin, so users write the diagonal-of-a-product form, which builds the fulln x nproduct just to keep its diagonal and drop the rest. The element-wise form gives the same sum without ever building the product.diag(v) @ X→X * v, and column scalingX @ diag(v)→X * t(v). Multiplying by a diagonal matrix just scales each row (or column), which an element-wise multiply does directly; so the big diagonal matrix and the matmul both go away.sum(s * X)→s * sum(X). Multiply once after the sum instead of once per element before it.rowAgg(X)/colAgg(X)when the axis being reduced has length 1. There is nothing to combine, so the aggregate returns its input and is removed. Covers thesum/min/maxrow and column variants.X + X→X * 2, plus(X * c) + X→X * (c + 1)and(X * c1) + (X * c2)→X * (c1 + c2). These chain, soX + X + XbecomesX * 3.(M1 + s1) + (M2 + s2)→(M1 + M2) + (s1 + s2). Add the two scalars once instead of spreading each one across a whole matrix.The change also marks
DiagVectorOpPure(its siblingDiagMatrixOpalready was) so dead-code elimination can remove the leftovermatMul/diagVectorafter the trace rewrite fires. The other rewrites already relied onTransposeOp/EwMulOpbeingPure.Soundness notes
Each rewrite does nothing unless its guards hold. They read shapes through the fail-closed accessors from #1015 and give up on anything unknown. The ones worth calling out:
sum(s * X)adds up in the product's type,s * sum(X)inX's type. Ifsis wider thanX(sayf64timessi64), the rewritten integer sum can overflow where the original would not, so this rewrite only fires whenX's type already matches the sum's result type.X + X → X * 2is the one exception: doubling a float is exact, so it stays bit-for-bit identical (NaN/Inf/-0.0included). Integer overflow in the regrouped scalar sums is fine; DAPHNE's integer kernels wrap at2^n, so the result is identical even when an intermediate wraps.EwMulcanonicalizer only swaps a scalar left operand, so that order sticks.Correctness
Each rewrite has both a LIT test (
test/codegen/rewrite_linalg_simplify.mlir: a positive case, plus the negatives that must not fire: unknown shapes, transposed matmul flags, and the type cases above) and an end-to-end numerical test (test/api/cli/operations/) that runs a triggering script and an equivalent reference script written a way the rewrite can't match, and checks the output matches exactly. The repeated-add and shifted-matrix pairs deliberately overflowsi64/si32, so identical output confirms the wrapping arithmetic is preserved.Those numerical tests paid off. The first version of the trace rewrite produced an unknown-typed product that compiled on its own but failed to lower on a real script, because the type was never resolved and kernel dispatch had nothing to bind to. A test that only checks whether the rewrite fires would have missed it. The fix, giving the new ops a real result type, is included here.
Full
[operations]suite green (488 assertions / 67 cases). Otherwise the suite matches baseline. The failures the local wrapper filters out are two pre-existing aarch64-only crashes (#1011, #1012) that don't reproduce on x86_64 CI.Deliberately out of scope
wsloss/wsigmoid/wdivmm/wcemm); no ops or kernels for them in DAPHNE yet, so they are follow-up work.ShapeFromArgmakes the column check always pass, so it would silently drop partial writes. Zero-absorption based on sparsity: a compile-time sparsity estimate isn't a guarantee about the actual values. Thesum(t(X))/sum(reverse(X))fold an earlier version carried now lives in feat(ir): Algebraic trait vocabulary + shared scalar-fold interface [#512] #1014, via theOnlyReordersElements/OrderAgnosticAggregatetraits.On the commit range
This sits on top of #1015. The base branches live on my fork and can't be a PR base upstream, so the diff currently also includes the commits from #1013, #1014, and #1015. Once those merge I'll rebase and force-push, leaving only this PR's own commits.
Refs #512.