Repository navigation
feat: kipoiseq2, kipoiseq without the Kipoi dataloaders and kipoi dependencies - #1
Merged
Merged
Conversation
…endencies kipoiseq2 continues kipoiseq as its own package, so kipoiseq 0.7.1 and its dependents keep working unchanged. The Kipoi model zoo is deprecated, so kipoiseq2 drops the Kipoi dataloaders and the kipoi dependencies. Extractors, transforms, Interval and Variant keep their API under the import name kipoiseq2. - rename the distribution and import package to kipoiseq2, starting at version 0.0.0 so that release-please cuts 0.1.0 - remove kipoiseq.dataloaders with its tests, notebooks and docs page, and the kipoi, kipoi-utils, kipoi-conda and gffutils dependencies - declare deprecation, which the extractors import and which only arrived through kipoi before - require Python >= 3.10 and drop six and the __future__ imports - move setup.py, setup.cfg, requirements.txt and bumpversion to pyproject.toml (hatchling, uv, tox) - replace CircleCI, .dependabot and pep8speaks with GitHub Actions: CI on Python 3.10 to 3.14, PR title lint, release-please and PyPI trusted publishing - format the code with ruff and fix the lint and mypy findings fix(extractors): infer UTRs from the CDS with pandas 2 and later
filter() stored the Interval class instead of each interval, so iter_intervals() and filter_range() saw the class after a filter call.
Interval is 0-based and half-open everywhere else: the FASTA fetch, the cyvcf2 region and width() all treat `end` as exclusive. Three places treated it as inclusive and so dropped the last base of a chromosome: - Interval.is_valid rejected an interval with end == chrom_len - Interval.truncate clamped the end to chrom_len - 1 - VariantSeqExtractor.extract clamped the fetch end to chrom_len - 1, so it filled the last base with N (is_padding=True) or raised Document Interval as 0-based, half-open and Variant.end as exclusive.
kipoiseq2 keeps the sequence, VCF and matching extractors that MMSplice, AbSplice and kipoi_enformer use. The dropped modules have no user among them and pulled in pandas-heavy GTF parsing. - remove extractors.gtf, extractors.protein, extractors.multi_interval and extractors.variant_combinations with their tests, the protein notebook, the protein benchmark and the test data only they used - remove Interval.from_pybedtools and Interval.to_pybedtools - remove the tqdm progress bar (the progress argument of MultiSampleVCF.query_variants, query_all and VariantIntervalQueryable) - replace the deprecation package with warnings.warn
VariantSeqExtractor used pyfaidx.Sequence only to carry coordinates through slices, and pyfaidx.complement for the reverse complement. pyfaidx.Sequence guesses 0- or 1-based coordinates from the sequence length, so a fetch that is one base short silently shifts the slices. - add Subsequence, a frozen dataclass with 0-based, half-open coordinates that stay in sync on slicing - add reverse_complement, which uses str.translate with the pyfaidx IUPAC table and raises ValueError on other characters On 3000 random intervals with SNVs, MNVs, insertions and deletions the output equals that of the pyfaidx version.
tl;dr: SingleVariantMatcher drops pandas and pyranges. It takes the intervals as a polars DataFrame and finds all pairs with one polars-bio overlap join. kipoiseq2 now needs only numpy and pyfaidx. - add overlap_variants: one pb.overlap join in 0-based, half-open coordinates, set per frame instead of through the process-wide pb.set_option; the pairs are sorted by variant, then by interval - add SingleVariantMatcher.pairs(), which returns all pairs as one DataFrame, and iter_batches(); iteration still yields (Interval, Variant), one join per batch of variant_batch_size - replace the pranges argument with regions (polars DataFrame) and drop gtf_path and bed_path; arguments after variant_fetcher are keyword-only - replace variants_to_pyranges, intervals_to_pyranges and pyranges_to_intervals with variants_to_polars and intervals_to_polars, and PyrangesVariantFetcher with VariantListFetcher - move cyvcf2 to the extra "vcf" and polars and polars-bio to the extra "ranges"; require Python >= 3.12, which polars-bio needs - add boundary tests: variants at the first and last base, deletions across the edges, insertions at the edges and empty REF alleles The boundary cases and 2000 random variants on 300 random intervals give the same pairs as the pyranges matcher of kipoiseq 0.7.1.
Follow the reloftee layout: the package lives in src/kipoiseq2, so the tests import the installed package and not the checkout.
The function assigned the read-only start and end attributes, so it raised AttributeError on every Interval. It now uses Interval.slop.
MultiSampleVCF already raises an ImportError that names the extra. to_vcf imported cyvcf2.Writer directly and raised a bare ModuleNotFoundError. Both paths now have a test.
The protein extractors were removed.
release-please.yml called publish.yml via workflow_call. PyPI rejects such uploads with "400 Invalid attestations": the attestation names the top-level workflow, release-please.yml, but the trusted publisher on PyPI is publish.yml. The publish job now starts publish.yml with workflow_dispatch. That is the one event GITHUB_TOKEN may use to start another workflow, so no PAT is needed. publish.yml keeps its manual trigger and drops workflow_call.
The GitHub default branch is main.
…argument The variant matchers now take one required `intervals` argument: a polars DataFrame or a sequence of Interval objects. The matcher picks the input type with isinstance. `BaseVariantMatcher.regions` is now the private `_interval_frame`.
The intervals were converted to a polars DataFrame first, which used up a generator. The list kept for the Interval objects was then empty.
_regions_from_variants compared with a hard-coded 150. The docstring of get_variants also named a strategy argument that does not exist.
polars-bio indexes the second input of an overlap join in memory and streams the first one. With the variants as the second input, peak memory grew with the number of variants.
scan_vcf_variants returns one row per ALT allele, with the coordinates and the ALT filter of MultiSampleVCF. The polars floor moves to 1.37.1, the version that polars-bio needs and that has explode(empty_as_null=).
The carrier check works per ALT allele: a sample with GT 0/2 carries the second ALT allele of the record, but not the first one.
The matchers now read a VCF file with scan_vcf_variants and also take a polars frame as variants. SingleVariantMatcher gains scan_pairs, and iter_batches yields the pair DataFrames with a global variant_idx. MultiVariantsMatcher joins once instead of querying per interval. variant_fetcher, vcf_lazy and VariantListFetcher are removed.
scan_vcf_variants and scan_vcf_genotypes replace MultiSampleVCF, so MultiSampleVCF, vcf_query, VariantFetcher, the two VCF seq extractors, Variant.from_cyvcf and the vcf extra go. The tests keep the cyvcf2 results as fixed expected values.
Filters on QUAL and FILTER replace query_all().filter(...) and FilterVariantQuery. The matchers keep only the coordinates, ref, alt and allele_idx of a VCF file, so the pairs do not grow.
With use_zero_based=False, the frame kept the polars-bio metadata of 1-based coordinates, although start and end are 0-based. polars-bio range operations on the frame then shifted the coordinates.
The matchers expose their intervals again, as the polars DataFrame `intervals`. Merging `regions` into `intervals` had made the table private, so callers that read `matcher.regions` had no public replacement.
…rval start VariantSeqExtractor.extract with fixed_len=True returned one base too many when the anchor was at or before the interval start and a deletion lay before the interval. The cut `down_str[-down_len:]` keeps the whole string for down_len == 0. kipoiseq 0.7.1 guarded this case. Commit 1bfbeb9 in kipoiseq master dropped the guard and was never released, but kipoiseq2 inherited it from master.
The repository moved from kipoi to gagneurlab. GitHub redirects the old URLs, but the PyPI project page takes the Homepage link from pyproject.toml.
…able The table read every TGA as selenocysteine, so a nonsense variant that creates TGA translated to U instead of a stop. Its XXX and NNN entries were markers of kipoiseq's protein dataloader, which kipoiseq2 does not have.
…translation exceptions translate takes the NCBI transl_table (1, 2 or 11) and transl_except, a map from 1-based codon number to amino acid, as in cdot. So selenocysteine, non-AUG starts and mitochondrial stops completed by poly(A) translate as annotated, and a TGA elsewhere stays a stop.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
kipoiseq2 continues kipoiseq as its own package, so kipoiseq 0.7.1 and its dependents keep working unchanged. The Kipoi model zoo is deprecated, so kipoiseq2 drops the Kipoi dataloaders and the kipoi dependencies. Interval, Variant, the sequence extractors, the variant matchers and the transforms stay, under the import name kipoiseq2. VCF files are read as polars tables with polars-bio instead of cyvcf2, so large VCFs stream through the variant matching. The core depends on numpy and pyfaidx only. The README lists the API changes under "Migrating from kipoiseq".
progressargumentstranslate(seq, hg38=True)and TRANSLATION_TABLE_FOR_HG38, which read every TGA as selenocysteineintervalsargument, a polars DataFrame or Interval objects, instead of the pyrangespranges, losegtf_pathandbed_path, and yield the pairs in VCF order. The variants come fromvcf_fileor fromvariants, a polars frame or Variant objects;variant_fetcherand PyrangesVariantFetcher are gone. Both matchers gainpairs()andscan_pairs(), and SingleVariantMatcher gainsiter_batches(). All three return polars framesscan_vcf_variantsreturns one row per ALT allele, andscan_vcf_genotypesone row per ALT allele and sample, with a per-allele carrier flag. This replaces MultiSampleVCF, vcf_query, VariantFetcher, the two VCF seq extractors and Variant.from_cyvcf, and drops cyvcf2rangesfrom __future__importsfix: allow intervals that end at the chromosome end
fix(transforms): build a new Interval in resize_interval
fix(extractors): keep the fixed length when the anchor is at the interval start
feat(transforms): translate with the NCBI genetic code and annotated translation exceptions