Skip to content

[Docs] Document filetype-aware input-extension defaults in deduplication workflows #2127

Description

@lbliii

Context

Child of #2118. #2045 changed exact, fuzzy, semantic, and removal workflows so omitted input_file_extensions derives from input_filetype instead of generic partitioner defaults. This is useful behavior but is absent from the relevant workflow references.

Requirements

  • Document the default extension set for every supported input_filetype.
  • Explain the distinction between omitted, explicit list, and empty/invalid values.
  • Cover exact, fuzzy, semantic, and duplicate-removal workflows consistently.
  • Include the JSONL example from the PR and at least one Parquet example.
  • Verify the existing fuzzy input_blocksize clarification from docs: clarify fuzzy dedup input blocksize #2096 remains correct and nearby terminology is consistent.

Acceptance criteria

  • Relevant dedup workflow pages show the derived defaults
  • Examples demonstrate explicit override behavior
  • Unsupported combinations and recursive discovery behavior are clear
  • Cross-workflow terminology is consistent
  • Fern checks pass

Related PRs: #2045, #2096

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

documentationImprovements or additions to documentation

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions