Skip to content

feat: kipoiseq2, kipoiseq without the Kipoi dataloaders and kipoi dependencies - #1

Merged
Hoeze merged 28 commits into
mainfrom
refactor/drop-kipoi
Oct 4, 2026
Merged

Hoeze merged 28 commits into
mainfrom
refactor/drop-kipoi

Conversation

@Hoeze

@Hoeze Hoeze commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

kipoiseq2 continues kipoiseq as its own package, so kipoiseq 0.7.1 and its dependents keep working unchanged. The Kipoi model zoo is deprecated, so kipoiseq2 drops the Kipoi dataloaders and the kipoi dependencies. Interval, Variant, the sequence extractors, the variant matchers and the transforms stay, under the import name kipoiseq2. VCF files are read as polars tables with polars-bio instead of cyvcf2, so large VCFs stream through the variant matching. The core depends on numpy and pyfaidx only. The README lists the API changes under "Migrating from kipoiseq".

  • rename the distribution and import package to kipoiseq2, starting at version 0.0.0 so that release-please cuts 0.1.0
  • remove kipoiseq.dataloaders with its tests, notebooks and docs page, and the kipoi, kipoi-utils, kipoi-conda and gffutils dependencies
  • remove the GTF, protein and variant-combination extractors and multi_interval.py, which no downstream package uses. This drops the pandas, pyranges, pybedtools, tqdm and deprecation dependencies and the progress arguments
  • remove translate(seq, hg38=True) and TRANSLATION_TABLE_FOR_HG38, which read every TGA as selenocysteine
  • match variants to intervals with a polars-bio overlap join, with the variants as the streaming side. SingleVariantMatcher and MultiVariantsMatcher take one intervals argument, a polars DataFrame or Interval objects, instead of the pyranges pranges, lose gtf_path and bed_path, and yield the pairs in VCF order. The variants come from vcf_file or from variants, a polars frame or Variant objects; variant_fetcher and PyrangesVariantFetcher are gone. Both matchers gain pairs() and scan_pairs(), and SingleVariantMatcher gains iter_batches(). All three return polars frames
  • read VCF files with polars-bio: scan_vcf_variants returns one row per ALT allele, and scan_vcf_genotypes one row per ALT allele and sample, with a per-allele carrier flag. This replaces MultiSampleVCF, vcf_query, VariantFetcher, the two VCF seq extractors and Variant.from_cyvcf, and drops cyvcf2
  • apply variants in VariantSeqExtractor without pyfaidx.Sequence
  • move polars and polars-bio to the extra ranges
  • require Python >= 3.12, which polars-bio needs; drop six and the from __future__ imports
  • move the package to the src layout, and setup.py, setup.cfg, requirements.txt and bumpversion to pyproject.toml (hatchling, uv, tox)
  • replace CircleCI, .dependabot and pep8speaks with GitHub Actions: CI on Python 3.12 to 3.14, PR title lint, release-please and PyPI trusted publishing
  • format the code with ruff and fix the lint and mypy findings

fix: allow intervals that end at the chromosome end

fix(transforms): build a new Interval in resize_interval

fix(extractors): keep the fixed length when the anchor is at the interval start

feat(transforms): translate with the NCBI genetic code and annotated translation exceptions

…endencies

kipoiseq2 continues kipoiseq as its own package, so kipoiseq 0.7.1 and
its dependents keep working unchanged. The Kipoi model zoo is
deprecated, so kipoiseq2 drops the Kipoi dataloaders and the kipoi
dependencies. Extractors, transforms, Interval and Variant keep their
API under the import name kipoiseq2.

- rename the distribution and import package to kipoiseq2, starting at
  version 0.0.0 so that release-please cuts 0.1.0
- remove kipoiseq.dataloaders with its tests, notebooks and docs page,
  and the kipoi, kipoi-utils, kipoi-conda and gffutils dependencies
- declare deprecation, which the extractors import and which only
  arrived through kipoi before
- require Python >= 3.10 and drop six and the __future__ imports
- move setup.py, setup.cfg, requirements.txt and bumpversion to
  pyproject.toml (hatchling, uv, tox)
- replace CircleCI, .dependabot and pep8speaks with GitHub Actions: CI
  on Python 3.10 to 3.14, PR title lint, release-please and PyPI trusted
  publishing
- format the code with ruff and fix the lint and mypy findings

fix(extractors): infer UTRs from the CDS with pandas 2 and later
filter() stored the Interval class instead of each interval, so
iter_intervals() and filter_range() saw the class after a filter call.
Interval is 0-based and half-open everywhere else: the FASTA fetch,
the cyvcf2 region and width() all treat `end` as exclusive. Three
places treated it as inclusive and so dropped the last base of a
chromosome:

- Interval.is_valid rejected an interval with end == chrom_len
- Interval.truncate clamped the end to chrom_len - 1
- VariantSeqExtractor.extract clamped the fetch end to chrom_len - 1,
  so it filled the last base with N (is_padding=True) or raised

Document Interval as 0-based, half-open and Variant.end as exclusive.
kipoiseq2 keeps the sequence, VCF and matching extractors that MMSplice,
AbSplice and kipoi_enformer use. The dropped modules have no user among
them and pulled in pandas-heavy GTF parsing.

- remove extractors.gtf, extractors.protein, extractors.multi_interval
  and extractors.variant_combinations with their tests, the protein
  notebook, the protein benchmark and the test data only they used
- remove Interval.from_pybedtools and Interval.to_pybedtools
- remove the tqdm progress bar (the progress argument of
  MultiSampleVCF.query_variants, query_all and VariantIntervalQueryable)
- replace the deprecation package with warnings.warn
VariantSeqExtractor used pyfaidx.Sequence only to carry coordinates
through slices, and pyfaidx.complement for the reverse complement.
pyfaidx.Sequence guesses 0- or 1-based coordinates from the sequence
length, so a fetch that is one base short silently shifts the slices.

- add Subsequence, a frozen dataclass with 0-based, half-open
  coordinates that stay in sync on slicing
- add reverse_complement, which uses str.translate with the pyfaidx
  IUPAC table and raises ValueError on other characters

On 3000 random intervals with SNVs, MNVs, insertions and deletions the
output equals that of the pyfaidx version.
tl;dr: SingleVariantMatcher drops pandas and pyranges. It takes the
intervals as a polars DataFrame and finds all pairs with one
polars-bio overlap join. kipoiseq2 now needs only numpy and pyfaidx.

- add overlap_variants: one pb.overlap join in 0-based, half-open
  coordinates, set per frame instead of through the process-wide
  pb.set_option; the pairs are sorted by variant, then by interval
- add SingleVariantMatcher.pairs(), which returns all pairs as one
  DataFrame, and iter_batches(); iteration still yields
  (Interval, Variant), one join per batch of variant_batch_size
- replace the pranges argument with regions (polars DataFrame) and
  drop gtf_path and bed_path; arguments after variant_fetcher are
  keyword-only
- replace variants_to_pyranges, intervals_to_pyranges and
  pyranges_to_intervals with variants_to_polars and
  intervals_to_polars, and PyrangesVariantFetcher with
  VariantListFetcher
- move cyvcf2 to the extra "vcf" and polars and polars-bio to the
  extra "ranges"; require Python >= 3.12, which polars-bio needs
- add boundary tests: variants at the first and last base, deletions
  across the edges, insertions at the edges and empty REF alleles

The boundary cases and 2000 random variants on 300 random intervals
give the same pairs as the pyranges matcher of kipoiseq 0.7.1.
Follow the reloftee layout: the package lives in src/kipoiseq2, so the
tests import the installed package and not the checkout.
The function assigned the read-only start and end attributes, so it
raised AttributeError on every Interval. It now uses Interval.slop.
MultiSampleVCF already raises an ImportError that names the extra. to_vcf
imported cyvcf2.Writer directly and raised a bare ModuleNotFoundError.
Both paths now have a test.
The protein extractors were removed.
release-please.yml called publish.yml via workflow_call. PyPI rejects
such uploads with "400 Invalid attestations": the attestation names the
top-level workflow, release-please.yml, but the trusted publisher on
PyPI is publish.yml.

The publish job now starts publish.yml with workflow_dispatch. That is
the one event GITHUB_TOKEN may use to start another workflow, so no PAT
is needed. publish.yml keeps its manual trigger and drops workflow_call.
The GitHub default branch is main.
…argument

The variant matchers now take one required `intervals` argument: a polars DataFrame or a sequence of Interval objects. The matcher picks the input type with isinstance. `BaseVariantMatcher.regions` is now the private `_interval_frame`.
The intervals were converted to a polars DataFrame first, which used up
a generator. The list kept for the Interval objects was then empty.
_regions_from_variants compared with a hard-coded 150. The docstring of
get_variants also named a strategy argument that does not exist.
polars-bio indexes the second input of an overlap join in memory and
streams the first one. With the variants as the second input, peak
memory grew with the number of variants.
scan_vcf_variants returns one row per ALT allele, with the coordinates
and the ALT filter of MultiSampleVCF. The polars floor moves to 1.37.1,
the version that polars-bio needs and that has explode(empty_as_null=).
The carrier check works per ALT allele: a sample with GT 0/2 carries
the second ALT allele of the record, but not the first one.
The matchers now read a VCF file with scan_vcf_variants and also take
a polars frame as variants. SingleVariantMatcher gains scan_pairs, and
iter_batches yields the pair DataFrames with a global variant_idx.
MultiVariantsMatcher joins once instead of querying per interval.
variant_fetcher, vcf_lazy and VariantListFetcher are removed.
scan_vcf_variants and scan_vcf_genotypes replace MultiSampleVCF, so
MultiSampleVCF, vcf_query, VariantFetcher, the two VCF seq extractors,
Variant.from_cyvcf and the vcf extra go. The tests keep the cyvcf2
results as fixed expected values.
Filters on QUAL and FILTER replace query_all().filter(...) and
FilterVariantQuery. The matchers keep only the coordinates, ref, alt
and allele_idx of a VCF file, so the pairs do not grow.
With use_zero_based=False, the frame kept the polars-bio metadata of
1-based coordinates, although start and end are 0-based. polars-bio
range operations on the frame then shifted the coordinates.
The matchers expose their intervals again, as the polars DataFrame `intervals`.
Merging `regions` into `intervals` had made the table private, so callers that read `matcher.regions` had no public replacement.
…rval start

VariantSeqExtractor.extract with fixed_len=True returned one base too many when the anchor was at or before the interval start and a deletion lay before the interval.
The cut `down_str[-down_len:]` keeps the whole string for down_len == 0.
kipoiseq 0.7.1 guarded this case. Commit 1bfbeb9 in kipoiseq master dropped the guard and was never released, but kipoiseq2 inherited it from master.
The repository moved from kipoi to gagneurlab. GitHub redirects the old URLs, but the PyPI project page takes the Homepage link from pyproject.toml.
…able

The table read every TGA as selenocysteine, so a nonsense variant that creates TGA translated to U instead of a stop. Its XXX and NNN entries were markers of kipoiseq's protein dataloader, which kipoiseq2 does not have.
…translation exceptions

translate takes the NCBI transl_table (1, 2 or 11) and transl_except, a map from 1-based codon number to amino acid, as in cdot. So selenocysteine, non-AUG starts and mitochondrial stops completed by poly(A) translate as annotated, and a TGA elsewhere stays a stop.
@Hoeze
Hoeze merged commit 36c521f into main Oct 4, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant