You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
docs: reshape hero path into Getting Started funnel (#1254)
Point readers through Docker auto-source quickstart → deployment →
advanced feeds; fold MCP/CLI/skill docs onto the same path and redirect
legacy /web-application/getting-started/.
Copy file name to clipboardExpand all lines: AGENTS.md
+9-9Lines changed: 9 additions & 9 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -86,18 +86,18 @@ Preferred verification flow for docs/content changes:
86
86
87
87
### User Journey Funnel
88
88
89
-
Maintain a directed "funnel" for documentation to maximize user success and conversion:
89
+
Keep docs pointed along one success path:
90
90
91
-
1.**Phase 1: Quickstart (Local Demo)** — The primary entry point. Run `html2rss-web` with Docker and generate a feed from a page URL in minutes.
92
-
2.**Phase 2: Production (Deployment)** — The goal for invested users. Move to a stable, production-ready instance.
93
-
3.**Phase 3: Refinement (Custom Configs)** — Secondary optimization. Author custom YAML configs only when automatic generation needs precise control.
91
+
1.**Getting Started** — Run `html2rss-web` with Docker; paste a page URL; open the generated feed.
92
+
2.**Deployment** — Production compose, tokens, LAN HTTP vs HTTPS reverse proxy.
93
+
3.**Advanced Feeds** — Custom YAML only when auto-source needs precise control (escape hatch).
94
94
95
-
**Rules for Funnel Maintenance:**
95
+
**Rules:**
96
96
97
-
-Avoid branching paths in introductory pages; always point toward the next phase in the funnel.
98
-
-Define "html2rss-web" as the primary interface and "page-to-RSS" as the primary workflow.
99
-
-Use "Feed Directory" consistently to refer to the pre-built feed catalog; avoid terms like "catalog", "included feeds", or "packaged configs" in user-facing docs.
100
-
- Do not introduce new terminology (e.g., "toolkit") or unrelated infrastructure concepts (e.g., "custom domains") unless they are essential to a specific guide.
97
+
-Introductory pages hand off to the next step; do not fork the reader into parallel “primary” paths.
98
+
-`html2rss-web` is the primary interface; page-URL auto-source is the primary workflow.
99
+
-Say **Feed Directory** for the curated feed list; avoid “catalog”, “included feeds”, or “packaged configs” in user-facing copy.
100
+
- Do not invent product terms or infrastructure side quests unless a specific operator guide needs them.
When existing feeds or auto-sourcing are not enough, write a YAML config for the site you want to follow.
10
+
Prefer [Getting Started](/getting-started/) (URL paste) and the [Feed Directory](/feed-directory/) first. Use a custom YAML config when auto-source misses items you care about, or when you need reviewable selectors.
11
11
12
-
**Prerequisites:** You should be familiar with the [Getting Started](/getting-started/) guide before diving into custom configurations.
13
-
14
-
<Asidetype="tip"title="Use this guide when you need more control">
15
-
Reach for a custom config when you need stable, reviewable extraction rules or generated output misses
16
-
important content.
12
+
<Asidetype="tip"title="Escape hatch">
13
+
Agents can draft configs via MCP (`capture` → `test` → `apply`) or the
skill. You still mount or publish the YAML yourself.
17
16
</Aside>
18
17
19
-
---
20
-
21
-
## When to Use Custom Configs
22
-
23
-
**Use custom configs when:**
24
-
25
-
-**Auto-sourcing doesn't work** for the website you want to follow
26
-
-**Existing feeds are incomplete** or missing important content
27
-
-**You need specific formatting** or data extraction
28
-
-**The website has complex structure** that requires custom selectors
29
-
-**You want to combine data** from multiple sources
30
-
31
-
## Recommended Workflow
32
-
33
-
1.**Inspect the live page** in your browser developer tools
34
-
2.**Optionally draft with capture** — `html2rss capture https://example.com/articles > your-config.yml` (see [Capturing Feed Configs](/ruby-gem/guides/capturing-feed-configs/))
35
-
3.**Write or refine the smallest useful config** that extracts items, titles, and links
36
-
4.**Validate the config** with `html2rss validate your-config.yml`
37
-
5.**Render the feed** with `html2rss feed your-config.yml`
38
-
6.**Add it to `html2rss-web`** so you can use it through your normal instance
39
-
7.**Escalate request strategy when needed**: use Botasaurus (`strategy: botasaurus` or `auto` with `BOTASAURUS_SCRAPER_URL`) only when troubleshooting requires browser rendering
40
-
41
-
This order keeps iteration fast and makes it easier to see whether the problem is the page structure, your
42
-
selectors, or the fetch strategy.
43
-
44
-
---
45
-
46
-
## How It Works
18
+
## Recommended workflow
47
19
48
-
A config file is a simple "recipe" that tells html2rss:
4. Live-check: `html2rss test your-config.yml`, then ship with `html2rss apply your-config.yml`
24
+
5. Mount into `html2rss-web` or contribute to html2rss-configs
25
+
6. Escalate to `strategy: botasaurus` (or `auto` with `BOTASAURUS_SCRAPER_URL`) only when Faraday is not enough
49
26
50
-
1.**Which website** to look at
51
-
2.**What content** to find
52
-
3.**How to organize** it into an RSS feed
27
+
`html2rss feed` is a Thor alias for `apply`. `html2rss auto` aliases `scrape` (one-shot, no YAML).
53
28
54
-
### The `channel` Block
55
-
56
-
This tells html2rss basic information about your feed - like giving it a name and telling it which website to look at.
57
-
58
-
**Example:**
29
+
## Minimal config
59
30
60
31
<Code
61
32
code={`
62
33
channel:
63
34
url: https://example.com/blog
64
-
title: My Awesome Blog
65
-
`}
66
-
lang="yaml"
67
-
/>
68
-
69
-
This says: "Look at this website and call the feed 'My Awesome Blog'"
70
-
71
-
### The `selectors` Block
72
-
73
-
This is where you tell the html2rss engine exactly what to find on the page. You use CSS selectors (like you might use in web design) to point to specific parts of the webpage.
74
-
75
-
**Example:**
76
-
77
-
<Code
78
-
code={`
35
+
title: My Blog
79
36
selectors:
80
37
items:
81
38
selector: "article.post"
@@ -84,123 +41,56 @@ This is where you tell the html2rss engine exactly what to find on the page. You
84
41
url:
85
42
selector: "h2 a"
86
43
extractor: "href"
87
-
`}
44
+
`}
88
45
lang="yaml"
89
46
/>
90
47
91
-
This says: "Find each article, get the title from the h2 anchor, and get the link from the same h2 anchor's href attribute"
92
-
93
-
**Need more details?** Check our [complete guide to selectors](/ruby-gem/reference/selectors/) for all the options.
**Step 1:** Inspect the website you want to create a feed for. Start with your browser's developer tools to inspect the live DOM. "View Page Source" can still help, but it may miss JavaScript-rendered content.
100
-
101
-
**Step 2:** Create a file called `example.com.yml` with this basic structure:
Use `{Organization} — {Feed surface}` for `directory.title` (for example, `Anthropic — News`). The [Feed Directory](/feed-directory/) lists configs from a running `html2rss-web` instance.
225
-
226
-
**Need help?** See our [contribution guide](/get-involved/contributing/) for detailed instructions.
227
-
228
-
---
229
-
230
-
## Troubleshooting
231
-
232
-
**Common issues when writing configs:**
233
-
234
-
-**No items found?** Check your selectors with browser tools (F12) - the `items.selector` might not match the page structure
235
-
-**Invalid YAML?** Use spaces, not tabs, and ensure proper indentation
236
-
-**Website not loading?** Check the URL and try accessing it in your browser
237
-
-**Missing content?** Try a browser-based rendering strategy during troubleshooting
238
-
-**Wrong data extracted?** Verify your selectors are pointing to the right elements
239
-
240
-
**Need more help?** See our [comprehensive troubleshooting guide](/troubleshooting/troubleshooting/) or ask in [GitHub Discussions](https://github.com/orgs/html2rss/discussions).
241
-
242
-
---
243
-
244
-
## Next Steps
245
-
246
-
**🎉 Congratulations!** You've learned the basics of creating html2rss configuration files.
247
-
248
-
### What's Next?
249
-
250
-
**For Beginners:**
251
-
252
-
-**[Run html2rss-web with Docker](/web-application/getting-started/)** - Use the newest integrated behavior
253
-
-**[Learn more about selectors](/ruby-gem/reference/selectors/)** - Master CSS selectors
254
-
-**[Submit your config via GitHub Web](https://github.com/html2rss/html2rss-configs)** - No Git knowledge required!
0 commit comments