Skip to content

Commit 4adfe88

Browse files
authored
docs: center scraping guidance on auto (#1255)
* docs(concepts): update server architecture and default http strategy * docs(guides): update strategy migration notes and direct tls deployment * docs(adapters): document default httpx strategy, cli flags, mcp schemas, and error handling * docs(infrastructure): sync environment variables and timeout hierarchy with falcon * docs: center scraping guidance on auto Deepen the public strategy interface around auto and default while keeping HTTP and server adapters behind their implementation seams. Remove migration narration and incomplete deployment side paths so end users can act on scrape outcomes instead of transport internals. * docs: correct strategy contract details State capture persistence and MCP discovery behavior exactly as implemented while keeping compatibility aliases out of recommended user guidance. Use the runtime UnknownStrategy name in troubleshooting.
1 parent a683638 commit 4adfe88

12 files changed

Lines changed: 55 additions & 63 deletions

File tree

‎src/content/docs/creating-custom-feeds.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -22,7 +22,7 @@ Prefer [Getting Started](/getting-started/) (URL paste) and the [Feed Directory]
2222
3. Validate: `html2rss validate your-config.yml`
2323
4. Live-check: `html2rss test your-config.yml`, then ship with `html2rss apply your-config.yml`
2424
5. Mount into `html2rss-web` or contribute to html2rss-configs
25-
6. Escalate to `strategy: botasaurus` (or `auto` with `BOTASAURUS_SCRAPER_URL`) only when Faraday is not enough
25+
6. Keep `strategy: auto`; if results are incomplete, confirm the companion scraper is configured before adding site-specific controls
2626

2727
`html2rss feed` is a Thor alias for `apply`. `html2rss auto` aliases `scrape` (one-shot, no YAML).
2828

‎src/content/docs/ruby-gem/guides/ai-agent-workflows.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -50,7 +50,7 @@ Put `BOTASAURUS_SCRAPER_URL` in the MCP `env` block — a shell export does not
5050

5151
## Golden loops
5252

53-
**Articles now:** `scrape` (or `batch_scrape`). `strategy: "auto"` already falls back to Botasaurus when configured — do not retry with explicit `faraday` after `auto`. Empty items can still be success; follow `next_step`.
53+
**Articles now:** `scrape` (or `batch_scrape`). Keep `strategy: "auto"` so html2rss can choose the available fetch path. Empty items can still be success; follow `next_step` instead of retrying with another strategy.
5454

5555
**Durable YAML:** optional `inspect` → `recon` → `capture` → `test` → `apply`.
5656

‎src/content/docs/ruby-gem/guides/capturing-feed-configs.mdx‎

Lines changed: 6 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -29,9 +29,6 @@ Print YAML to stdout:
2929
<Code
3030
code={`
3131
html2rss capture https://example.com/articles
32-
# Pin Botasaurus for JS-rendered listings
33-
BOTASAURUS_SCRAPER_URL="http://localhost:4010" \\
34-
html2rss capture https://example.com/articles --strategy botasaurus
3532
# Hint the item card when auto detection is weak
3633
html2rss capture https://example.com/articles --items_selector ".post-card"
3734
# Analyze a local HTML fixture
@@ -46,7 +43,7 @@ Print YAML to stdout:
4643

4744
Common options:
4845

49-
- `--strategy` — `auto`, `faraday`, `botasaurus`, or `local_file` (default `auto`)
46+
- `--strategy` — request strategy override; defaults to `auto` (see [Strategy](/ruby-gem/reference/strategy/))
5047
- `--items_selector` — CSS selector hint for item cards
5148
- `--limit` — maximum articles kept while deriving selectors (default `25`)
5249
- `--max-redirects` / `--max-requests` — request budget overrides
@@ -62,10 +59,9 @@ See the [CLI reference](/ruby-gem/reference/cli-reference/#capture) for the full
6259
require 'html2rss'
6360
# Derive a config hash (channel + items selector with enhance: true)
6461
config = Html2rss.capture('https://example.com/articles')
65-
# Pin strategy or provide an items hint
62+
# Provide an item-card hint when automatic detection is weak
6663
config = Html2rss.capture(
67-
'https://spa-site.com',
68-
strategy: :botasaurus,
64+
'https://example.com/articles',
6965
items_selector: '.article-card'
7066
)
7167
File.write('my-feed.yml', Html2rss::Config.to_yaml(config))
@@ -76,11 +72,11 @@ See the [CLI reference](/ruby-gem/reference/cli-reference/#capture) for the full
7672

7773
## How It Works
7874

79-
1. **Request** — `FeedPipeline` (AutoFallback when `:auto`)
75+
1. **Request** — fetch the page using the configured strategy (`auto` by default)
8076
2. **Discover** — AutoSource extracts admitted articles
8177
3. **Segment** — SST Segmenter strategies `:list` → `:cluster` → `:semantic`
8278
4. **Gate** — emit an items selector only when enough articles match
83-
5. **Assemble** — `{ items: { selector:, enhance: true } }` plus channel. When AutoFallback selects a concrete transport (or you pin one), Capture **stamps** `strategy:` into the YAML so later `html2rss apply` / `Html2rss.feed` replay the same transport.
79+
5. **Assemble** — `{ items: { selector:, enhance: true } }` plus channel. Capture records the concrete strategy that produced the draft so `html2rss apply` / `Html2rss.feed` can reproduce that fetch path.
8480

8581
When the quality gate fails, selectors are omitted (`has_selectors: false`) rather than inventing attribute selectors. Hint with `--items_selector` or refine by hand.
8682

@@ -99,7 +95,7 @@ MCP `capture` returns that YAML in `payload.yaml`. `validate` / `apply` accept t
9995

10096
1. Validate: `html2rss validate my-feed.yml`
10197
2. Render: `html2rss apply my-feed.yml`
102-
3. Tighten the items selector, strategy, or `request.botasaurus` options if needed
98+
3. Tighten the items selector if needed; keep `auto` unless diagnosis proves a fixed fetch mode is required
10399
4. For Feed Directory contributions, add `directory.topics` and keep `enhance: true` unless chrome leaks (see [Creating Custom Feeds](/creating-custom-feeds/#sharing-your-config))
104100

105101
## Related

‎src/content/docs/ruby-gem/guides/handling-dynamic-content.mdx‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -9,9 +9,9 @@ Some websites load their content dynamically using JavaScript. Static fetch path
99

1010
## Solution
1111

12-
Default `strategy: auto` automatically tries `faraday` first, then falls back to `botasaurus` when `BOTASAURUS_SCRAPER_URL` is configured. This handles many JS-rendered listing pages without needing custom configuration.
12+
Keep the default `strategy: auto` and configure `BOTASAURUS_SCRAPER_URL`. html2rss can then choose an available fetch path based on the scrape result, which handles many JavaScript-rendered listing pages without per-feed strategy configuration.
1313

14-
When a site requires browser rendering or anti-bot bypass by default, you can explicitly set `strategy: botasaurus` and configure request controls under `request.botasaurus`:
14+
If a site still needs browser-specific navigation, waits, or scrolling, pin `strategy: botasaurus` and configure those advanced controls under `request.botasaurus`:
1515

1616
<Code
1717
code={`
@@ -92,7 +92,7 @@ Configure browser actions under `request.botasaurus`:
9292

9393
### JSON Loaded Over XHR
9494

95-
When Botasaurus uses the browser tier, captured JSON XHR/fetch bodies feed AutoSource `xhr_articles` automatically (enabled by default). Prefer `strategy: botasaurus` (or `auto` with `BOTASAURUS_SCRAPER_URL`) for SPA listing pages that hydrate article lists over the network rather than embedding them in HTML. See [Auto Source](/ruby-gem/reference/auto-source/) and [Strategy](/ruby-gem/reference/strategy/#botasaurus).
95+
Browser-rendered fetches can pass captured JSON XHR/fetch bodies to AutoSource `xhr_articles` automatically (enabled by default). With the companion scraper configured, keep `auto` for SPA listing pages that hydrate article lists over the network; pin `botasaurus` only when you need its browser-specific controls. See [Auto Source](/ruby-gem/reference/auto-source/) and [Strategy](/ruby-gem/reference/strategy/#botasaurus).
9696

9797
## Performance Considerations
9898

@@ -102,7 +102,7 @@ Browser-based extraction uses more resources than static HTTP fetching because i
102102
- Executes JavaScript and handles DOM events
103103
- Manages browser pools and network emulation
104104

105-
Use static HTTP fetching (`faraday`) for static content, and lean on `auto` or explicit `botasaurus` strategies when browser rendering is required. See the [Strategy Reference](/ruby-gem/reference/strategy/) for details.
105+
Keep `auto` for normal use. Pin a concrete strategy only when diagnosing a site or configuring browser-specific behavior. See the [Strategy Reference](/ruby-gem/reference/strategy/) for details.
106106

107107
## Related Topics
108108

‎src/content/docs/ruby-gem/reference/auto-source.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -35,7 +35,7 @@ What each scraper does:
3535
- **`microdata`:** Extracts HTML Microdata annotations (`itemscope itemtype`).
3636
- **`microformats2`:** Parses Microformats2 `h-entry` markup, extracting `p-name`, `e-content`, `u-url`, `dt-published`, `p-author`, `p-category`, and `u-photo` / `u-featured` media.
3737
- **`json_state`:** Walks in-page JSON (`<script type="application/json">`, `window.__NEXT_DATA__`, `window.__NUXT__`, `window.STATE`) for arrays with `title`/`url` pairs.
38-
- **`xhr_articles`:** Reuses JSON XHR/fetch bodies captured during a Botasaurus **browser** scrape (no extra HTTP). Empty for Faraday and Botasaurus HTTP-request tiers.
38+
- **`xhr_articles`:** Reuses JSON XHR/fetch bodies captured during browser execution (no extra request). It is empty when the selected fetch path does not capture browser responses.
3939
- **`wordpress_api`:** Detects `<link rel="https://api.w.org/">` and pulls posts from the REST API. See [WordPress API](/ruby-gem/reference/wordpress-api/).
4040
- **`sitemap`:** Locates XML sitemaps (`<link rel="sitemap">`, `/sitemap.xml`, or `/robots.txt`), filtering by priority and recency, with Google News tags (`<news:news>`).
4141
- **`meta_oembed`:** OpenGraph/Twitter meta tags plus JSON oEmbed (`<link rel="alternate" type="application/json+oembed">`).

‎src/content/docs/ruby-gem/reference/cli-reference.mdx‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -37,7 +37,7 @@ Command: `html2rss scrape [URL]`
3737

3838
Available options:
3939

40-
- `--strategy`: Optional request strategy (`auto`, `faraday`, `botasaurus`, `local_file`). Defaults to `auto`, which tries `faraday` -> `botasaurus`.
40+
- `--strategy`: Optional request strategy (`auto`, `default`, `botasaurus`, `local_file`). Defaults to `auto`, which chooses an available fetch path based on scrape results. Use a concrete value only for diagnosis or strategy-specific controls.
4141
- `--format`: Output format for the auto-sourced feed (`rss` or `jsonfeed`). Defaults to `rss`.
4242
- `--limit`: Maximum number of articles to extract during discovery (defaults to `25`).
4343
- `--items_selector`: Optional CSS selector hint for item extraction.
@@ -76,11 +76,11 @@ When no extractable items are found, `scrape` classifies likely causes instead o
7676

7777
Known anti-bot interstitial responses (for example Cloudflare challenge pages) are surfaced explicitly as blocked-surface errors.
7878

79-
If all fallback tiers run but still extract zero items, html2rss raises:
79+
If `auto` exhausts the available fetch paths without items, html2rss raises:
8080

8181
- `No RSS feed items extracted after auto fallback ...`
8282

83-
If failures continue after URL/surface fixes, ensure `BOTASAURUS_SCRAPER_URL` is set so the `auto` Botasaurus tier can run.
83+
If failures continue after URL/surface fixes, ensure `BOTASAURUS_SCRAPER_URL` is set so `auto` can use the companion scraper.
8484

8585
Start by changing the input URL to a direct listing/update page, then move to explicit selectors if needed.
8686

@@ -123,7 +123,7 @@ Command: `html2rss apply YAML_FILE [feed_name]`
123123

124124
Available options:
125125

126-
- `--strategy`: Request strategy override (`auto`, `faraday`, `botasaurus`, `local_file`).
126+
- `--strategy`: Request strategy override (`auto`, `default`, `botasaurus`, `local_file`).
127127
- `--params`: Dynamic parameters passed as key-value pairs (e.g. `--params id:42 section:news`).
128128
- `--max-redirects`: Maximum redirects to follow per request.
129129
- `--max-requests`: Total request budget allowed for this feed build.
@@ -150,15 +150,15 @@ Command: `html2rss capture [URL]`
150150

151151
Available options:
152152

153-
- `--strategy`: Optional request strategy (`auto`, `faraday`, `botasaurus`, `local_file`). Defaults to `auto`.
153+
- `--strategy`: Optional request strategy (`auto`, `default`, `botasaurus`, `local_file`). Defaults to `auto`.
154154
- `--items_selector`: Optional CSS selector hint for item extraction.
155155
- `--limit`: Maximum number of articles to keep (defaults to `25`).
156156
- `--max-redirects`: Maximum redirects to follow per request.
157157
- `--max-requests`: Maximum requests to allow for this feed build.
158158
- `--input`: Local HTML file path to read input from without making network requests.
159159
- `--explain`: Print capture quality JSON to stderr (`articles_count`, `channel_title`, `has_selectors`, `segment_strategy`, `selected_strategy`, `admission_drops`). Stdout stays YAML.
160160

161-
When AutoFallback (or a pinned strategy) selects a concrete transport, the printed YAML includes a top-level `strategy:` so later `html2rss apply` uses the same hop.
161+
Capture records the concrete strategy that produced the draft as a top-level `strategy:` so later `html2rss apply` can reproduce that fetch path.
162162

163163
### MCP
164164

‎src/content/docs/ruby-gem/reference/configuration.mdx‎

Lines changed: 3 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -16,7 +16,6 @@ Global configuration is thread-safe and implements an RCU (Read-Copy-Update) mec
1616
Html2rss.configure do |config|
1717
config.log_level = :info
1818
config.min_ttl = 60
19-
config.default_strategy = :faraday
2019
config.headers = { 'User-Agent' => 'MyCustomUserAgent/1.0' }
2120
end
2221
`}
@@ -76,15 +75,15 @@ Defines HTTP headers that are globally appended/prepended to all requests. You c
7675

7776
### `default_strategy`
7877

79-
Sets the default scraper strategy name used when a feed configuration doesn't specify a `strategy`. The strategy name must correspond to a registered strategy.
78+
Overrides the strategy used when a feed configuration does not specify one. Keep the gem default (`auto`) unless every feed in the process requires a fixed fetch mode.
8079

8180
- **Type**: `Symbol`, `String`, or `nil`
82-
- **Default**: `nil` (falls back to the gem default strategy, usually `auto`)
81+
- **Default**: `nil` (uses `auto`)
8382

8483
<Code
8584
code={`
8685
Html2rss.configure do |config|
87-
config.default_strategy = :faraday
86+
config.default_strategy = :auto
8887
end
8988
`}
9089
lang="ruby"

‎src/content/docs/ruby-gem/reference/mcp-server.mdx‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -28,9 +28,9 @@ Daemon logs go to **stderr** (stdout is JSON-RPC). Default `LOG_LEVEL` for the M
2828

2929
## Strategy
3030

31-
`scrape` and `capture` with `strategy: "auto"` run Faraday → Botasaurus AutoFallback in one call. Prefer native RSS/Atom when present; weak homepage URLs may be rewritten via entry resolution.
31+
`scrape` and `capture` use `strategy: "auto"` by default so html2rss can choose an available fetch path from the scrape result. Prefer native RSS/Atom when present; weak homepage URLs may be rewritten via entry resolution.
3232

33-
`inspect` with `auto` stays on Faraday. Pin `strategy: "botasaurus"` when inspect needs browser rendering.
33+
Keep `auto` for normal use. Pin `strategy: "botasaurus"` only when a diagnostic requires browser rendering or browser-specific controls.
3434

3535
Read `html2rss://runtime` for `botasaurus_configured` (boolean only). Set `BOTASAURUS_SCRAPER_URL` on the **MCP process** env.
3636

@@ -62,7 +62,7 @@ Golden path for durable YAML: `inspect` (optional) → `recon` → `capture` →
6262

6363
One-shot article extraction as JSON Feed items (no saved config).
6464

65-
- **Parameters:** `url` (required); `strategy` (`auto` / `faraday` / `botasaurus`, default `auto`); `limit` (default `25`); optional `items_selector`
65+
- **Parameters:** `url` (required); `strategy` (`auto` / `default` / `botasaurus`, default `auto`); `limit` (default `25`); optional `items_selector`
6666
- **Payload:** `items`, plus totals / strategy fields when present
6767
- Empty items can still be `ok: true` — follow `next_step` / `guidance`
6868

@@ -120,7 +120,7 @@ Ship gate: build RSS from a config. Required `url` plus exactly one of `config`
120120
| ----------------------- | ------------------------------------------------------------------------------------------------------------------ |
121121
| `html2rss://schema` | Feed config JSON Schema |
122122
| `html2rss://extractors` | Extractor names |
123-
| `html2rss://strategies` | `auto`, `faraday`, `botasaurus` (not `local_file`) |
123+
| `html2rss://strategies` | Runtime strategy names for client discovery; prefer `auto`, `default`, or `botasaurus` |
124124
| `html2rss://runtime` | `version`, `mcp_contract_version`, `catalog_fingerprint`, `tools`, `botasaurus_configured` — never the scraper URL |
125125

126126
`validate` / `apply` reject `strategy: local_file` and `request.local_file_path`. Use CLI `--input` for fixtures.

0 commit comments

Comments
 (0)