Skip to content

Add html_to_r() - #1812

Closed
stevecondylios wants to merge 1 commit into
yihui:masterfrom
stevecondylios:master
Closed

stevecondylios wants to merge 1 commit into
yihui:masterfrom
stevecondylios:master

Conversation

@stevecondylios

Copy link
Copy Markdown

Regards #1811

  • Adds html_to_r() to easily extract R code from .html files generated through R Markdown
  • Adds tests for html_to_r()

@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@cderv cderv linked an issue Jan 29, 2021 that may be closed by this pull request
@yihui

yihui commented Sep 17, 2026

Copy link
Copy Markdown
Owner

Thanks for this, @stevecondylios, and apologies for the very long delay in getting back to you.

Coming back to this now, I don't think we should merge it, for a few reasons:

  1. It doesn't match what the discussion in Extract R code from R Markdown HTML file #1811 converged on. There, @atusy proposed a much simpler pandoc-based approach (HTML → commonmark → purl()), and you agreed it was "a great simplification and improvement on the DIY solution." That thread also resolved the two open questions — pandoc handles the character-entity decoding for free, and inc_out was deemed unnecessary. But this PR is still the original regex-based implementation, not the agreed direction.

  2. The regex approach doesn't work on real R Markdown HTML. html_to_r() searches for Markdown code fences (```r, ```) inside <body>, but rmarkdown::render() produces <pre class="r"><code>...</code></pre>, not Markdown fences. The tests here pass because their input is hand-written HTML containing literal fences, which isn't what a knitted document looks like — so on the actual motivating example (the applied-ml .html files), nothing would be extracted.

  3. The bare ``` branch also captures non-R output blocks, which is the ambiguity @atusy flagged.

There's also a new stringr dependency, which we try to avoid (knitr leans on xfun).

Given that the working approach from #1811 is essentially a short pandoc round-trip that anyone can run ad hoc, and that reconstructing source from rendered HTML is a bit outside knitr's core scope, I'll close this and #1811. If you (or anyone) still wants a helper for this, the pandoc-based purloc() sketch in #1811 is the right starting point. Thanks again for the idea and the write-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Extract R code from R Markdown HTML file

3 participants