Repository navigation
fix: address embedding correctness and worker configuration - #766
Closed
RAMZI0TO99 wants to merge 1 commit into
Closed
RAMZI0TO99 wants to merge 1 commit into
RAMZI0TO99 wants to merge 1 commit into
Conversation
Member
|
Hey @RAMZI0TO99 |
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This consolidates six fixes for sparse embedding edge cases, parallel worker configuration, and ConvNeXT preprocessing. Each fix includes regression tests and docstrings for the affected functions.
SparseEmbedding.from_dict({})used for NumPy indexingad1u66picuda=Falseand no device IDsdo_resize=False1e-8log(1+x)rounds its weight to zero and drops the dimensionlog1p(x)preserves the small positive weightkitchens kitchens prostaglandins, reorderedThe hash function and existing token IDs are preserved. The resize flag still defaults to enabled when omitted. SPLADE masking, maximum pooling, and nonzero filtering are unchanged.
This is a draft for collision-policy review. Related: #369.
kitchensandprostaglandinshave signed hashes 358434922 and -358434922, so their absolute IDs collide. This proposal sums their separately calculated lexical BM25 weights. Aggregating their counts before applying the nonlinear BM25 formula is another policy and would produce 1.9936283185840709 for the example. The collision regressions check order independence within floating-point tolerance without prescribing either policy. Maintainer input is requested on which policy belongs here.Corrected collision weights require recomputing affected stored document vectors for consistent use. Preserving IDs does not eliminate hash aliases. No retrieval-quality improvement or bitwise order invariance is claimed.
Validation against the actual checkout, FastEmbed 0.8.1 on Windows 11 build 10.0.26200, Python 3.13.5, NumPy 2.3.5, and ONNX Runtime 1.30.0 CPU:
600ce23fd7611e71694fcf724f46778f1ad24c1c, covering batch/single reference embeddings, parallel processing, lazy loading, session options, and token counts. Canonical expectations were unchanged. The single-embedding test used its existing model-selection hook to select SPLADE.The full model suite, GPU execution, retrieval benchmarks, and full CI matrix were not run locally. The model-test runner emitted one non-failing anyio assertion-rewrite warning.
This combined submission supersedes #762, #763, #764, and #765. The UTF-8 stopword candidate is excluded following duplicate review because #684 already contains that production change.