You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
* docs(concepts): update server architecture and default http strategy
* docs(guides): update strategy migration notes and direct tls deployment
* docs(adapters): document default httpx strategy, cli flags, mcp schemas, and error handling
* docs(infrastructure): sync environment variables and timeout hierarchy with falcon
* docs: center scraping guidance on auto
Deepen the public strategy interface around auto and default while keeping HTTP and server adapters behind their implementation seams. Remove migration narration and incomplete deployment side paths so end users can act on scrape outcomes instead of transport internals.
* docs: correct strategy contract details
State capture persistence and MCP discovery behavior exactly as implemented while keeping compatibility aliases out of recommended user guidance. Use the runtime UnknownStrategy name in troubleshooting.
Copy file name to clipboardExpand all lines: src/content/docs/ruby-gem/guides/ai-agent-workflows.mdx
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -50,7 +50,7 @@ Put `BOTASAURUS_SCRAPER_URL` in the MCP `env` block — a shell export does not
50
50
51
51
## Golden loops
52
52
53
-
**Articles now:**`scrape` (or `batch_scrape`). `strategy: "auto"`already falls back to Botasaurus when configured — do not retry with explicit `faraday` after `auto`. Empty items can still be success; follow `next_step`.
53
+
**Articles now:**`scrape` (or `batch_scrape`). Keep `strategy: "auto"`so html2rss can choose the available fetch path. Empty items can still be success; follow `next_step` instead of retrying with another strategy.
4.**Gate** — emit an items selector only when enough articles match
83
-
5.**Assemble** — `{ items: { selector:, enhance: true } }` plus channel. When AutoFallback selects a concrete transport (or you pin one), Capture **stamps**`strategy:` into the YAML so later `html2rss apply` / `Html2rss.feed`replay the same transport.
79
+
5.**Assemble** — `{ items: { selector:, enhance: true } }` plus channel. Capture records the concrete strategy that produced the draft so `html2rss apply` / `Html2rss.feed`can reproduce that fetch path.
84
80
85
81
When the quality gate fails, selectors are omitted (`has_selectors: false`) rather than inventing attribute selectors. Hint with `--items_selector` or refine by hand.
86
82
@@ -99,7 +95,7 @@ MCP `capture` returns that YAML in `payload.yaml`. `validate` / `apply` accept t
99
95
100
96
1. Validate: `html2rss validate my-feed.yml`
101
97
2. Render: `html2rss apply my-feed.yml`
102
-
3. Tighten the items selector, strategy, or `request.botasaurus` options if needed
98
+
3. Tighten the items selector if needed; keep `auto` unless diagnosis proves a fixed fetch mode is required
103
99
4. For Feed Directory contributions, add `directory.topics` and keep `enhance: true` unless chrome leaks (see [Creating Custom Feeds](/creating-custom-feeds/#sharing-your-config))
Copy file name to clipboardExpand all lines: src/content/docs/ruby-gem/guides/handling-dynamic-content.mdx
+4-4Lines changed: 4 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -9,9 +9,9 @@ Some websites load their content dynamically using JavaScript. Static fetch path
9
9
10
10
## Solution
11
11
12
-
Default `strategy: auto`automatically tries `faraday` first, then falls back to `botasaurus` when `BOTASAURUS_SCRAPER_URL` is configured. This handles many JS-rendered listing pages without needing custom configuration.
12
+
Keep the default `strategy: auto`and configure `BOTASAURUS_SCRAPER_URL`. html2rss can then choose an available fetch path based on the scrape result, which handles many JavaScript-rendered listing pages without per-feed strategy configuration.
13
13
14
-
When a site requires browser rendering or anti-bot bypass by default, you can explicitly set `strategy: botasaurus` and configure request controls under `request.botasaurus`:
14
+
If a site still needs browser-specific navigation, waits, or scrolling, pin `strategy: botasaurus` and configure those advanced controls under `request.botasaurus`:
15
15
16
16
<Code
17
17
code={`
@@ -92,7 +92,7 @@ Configure browser actions under `request.botasaurus`:
92
92
93
93
### JSON Loaded Over XHR
94
94
95
-
When Botasaurus uses the browser tier, captured JSON XHR/fetch bodies feed AutoSource `xhr_articles` automatically (enabled by default). Prefer `strategy: botasaurus` (or `auto` with `BOTASAURUS_SCRAPER_URL`) for SPA listing pages that hydrate article lists over the network rather than embedding them in HTML. See [Auto Source](/ruby-gem/reference/auto-source/) and [Strategy](/ruby-gem/reference/strategy/#botasaurus).
95
+
Browser-rendered fetches can pass captured JSON XHR/fetch bodies to AutoSource `xhr_articles` automatically (enabled by default). With the companion scraper configured, keep `auto` for SPA listing pages that hydrate article lists over the network; pin `botasaurus` only when you need its browser-specific controls. See [Auto Source](/ruby-gem/reference/auto-source/) and [Strategy](/ruby-gem/reference/strategy/#botasaurus).
96
96
97
97
## Performance Considerations
98
98
@@ -102,7 +102,7 @@ Browser-based extraction uses more resources than static HTTP fetching because i
102
102
- Executes JavaScript and handles DOM events
103
103
- Manages browser pools and network emulation
104
104
105
-
Use static HTTP fetching (`faraday`) for static content, and lean on `auto` or explicit `botasaurus` strategies when browser rendering is required. See the [Strategy Reference](/ruby-gem/reference/strategy/) for details.
105
+
Keep `auto` for normal use. Pin a concrete strategy only when diagnosing a site or configuring browser-specific behavior. See the [Strategy Reference](/ruby-gem/reference/strategy/) for details.
-**`json_state`:** Walks in-page JSON (`<script type="application/json">`, `window.__NEXT_DATA__`, `window.__NUXT__`, `window.STATE`) for arrays with `title`/`url` pairs.
38
-
-**`xhr_articles`:** Reuses JSON XHR/fetch bodies captured during a Botasaurus **browser** scrape (no extra HTTP). Empty for Faraday and Botasaurus HTTP-request tiers.
38
+
-**`xhr_articles`:** Reuses JSON XHR/fetch bodies captured during browser execution (no extra request). It is empty when the selected fetch path does not capture browser responses.
39
39
-**`wordpress_api`:** Detects `<link rel="https://api.w.org/">` and pulls posts from the REST API. See [WordPress API](/ruby-gem/reference/wordpress-api/).
40
40
-**`sitemap`:** Locates XML sitemaps (`<link rel="sitemap">`, `/sitemap.xml`, or `/robots.txt`), filtering by priority and recency, with Google News tags (`<news:news>`).
41
41
-**`meta_oembed`:** OpenGraph/Twitter meta tags plus JSON oEmbed (`<link rel="alternate" type="application/json+oembed">`).
-`--strategy`: Optional request strategy (`auto`, `faraday`, `botasaurus`, `local_file`). Defaults to `auto`, which tries `faraday` -> `botasaurus`.
40
+
-`--strategy`: Optional request strategy (`auto`, `default`, `botasaurus`, `local_file`). Defaults to `auto`, which chooses an available fetch path based on scrape results. Use a concrete value only for diagnosis or strategy-specific controls.
41
41
-`--format`: Output format for the auto-sourced feed (`rss` or `jsonfeed`). Defaults to `rss`.
42
42
-`--limit`: Maximum number of articles to extract during discovery (defaults to `25`).
43
43
-`--items_selector`: Optional CSS selector hint for item extraction.
@@ -76,11 +76,11 @@ When no extractable items are found, `scrape` classifies likely causes instead o
76
76
77
77
Known anti-bot interstitial responses (for example Cloudflare challenge pages) are surfaced explicitly as blocked-surface errors.
78
78
79
-
If all fallback tiers run but still extract zero items, html2rss raises:
79
+
If `auto` exhausts the available fetch paths without items, html2rss raises:
80
80
81
81
-`No RSS feed items extracted after auto fallback ...`
82
82
83
-
If failures continue after URL/surface fixes, ensure `BOTASAURUS_SCRAPER_URL` is set so the `auto`Botasaurus tier can run.
83
+
If failures continue after URL/surface fixes, ensure `BOTASAURUS_SCRAPER_URL` is set so `auto`can use the companion scraper.
84
84
85
85
Start by changing the input URL to a direct listing/update page, then move to explicit selectors if needed.
When AutoFallback (or a pinned strategy) selects a concrete transport, the printed YAML includes a top-level `strategy:` so later `html2rss apply`uses the same hop.
161
+
Capture records the concrete strategy that produced the draft as a top-level `strategy:` so later `html2rss apply`can reproduce that fetch path.
@@ -76,15 +75,15 @@ Defines HTTP headers that are globally appended/prepended to all requests. You c
76
75
77
76
### `default_strategy`
78
77
79
-
Sets the default scraper strategy name used when a feed configuration doesn't specify a `strategy`. The strategy name must correspond to a registered strategy.
78
+
Overrides the strategy used when a feed configuration does not specify one. Keep the gem default (`auto`) unless every feed in the process requires a fixed fetch mode.
80
79
81
80
-**Type**: `Symbol`, `String`, or `nil`
82
-
-**Default**: `nil` (falls back to the gem default strategy, usually`auto`)
Copy file name to clipboardExpand all lines: src/content/docs/ruby-gem/reference/mcp-server.mdx
+4-4Lines changed: 4 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -28,9 +28,9 @@ Daemon logs go to **stderr** (stdout is JSON-RPC). Default `LOG_LEVEL` for the M
28
28
29
29
## Strategy
30
30
31
-
`scrape` and `capture`with`strategy: "auto"`run Faraday → Botasaurus AutoFallback in one call. Prefer native RSS/Atom when present; weak homepage URLs may be rewritten via entry resolution.
31
+
`scrape` and `capture`use`strategy: "auto"`by default so html2rss can choose an available fetch path from the scrape result. Prefer native RSS/Atom when present; weak homepage URLs may be rewritten via entry resolution.
32
32
33
-
`inspect` with `auto`stays on Faraday. Pin `strategy: "botasaurus"` when inspect needs browser rendering.
33
+
Keep `auto`for normal use. Pin `strategy: "botasaurus"`only when a diagnostic requires browser rendering or browser-specific controls.
34
34
35
35
Read `html2rss://runtime` for `botasaurus_configured` (boolean only). Set `BOTASAURUS_SCRAPER_URL` on the **MCP process** env.
36
36
@@ -62,7 +62,7 @@ Golden path for durable YAML: `inspect` (optional) → `recon` → `capture` →
62
62
63
63
One-shot article extraction as JSON Feed items (no saved config).
0 commit comments