Skip to content

Commit 50a1bfc

Browse files
jeremymanningclaude
andcommitted
Clear the last mypy errors and store the resolvable DOI
Closes #18. extract_cv.py, the one finding that was a real shape problem rather than a stub gap: preprocess_content returned a bare string when extract_footnotes was false and a (footnote, content) pair when it was true. There is exactly one caller in the repo and it always passes true, so the string branch was dead code that its own docstring described as backwards compatibility. Dropped the flag and the function now always returns the pair. Losing the else branch also loses its blfootnote removal, which is equivalent: extract_footnote strips the footnote it extracts, and when there is no footnote there is nothing to strip. extract_people.py gets one attr_str helper rather than a cast per call site. BeautifulSoup types attribute access as str, AttributeValueList or None because a multi-valued attribute like class comes back as a list; src and href never do, so the helper joins a list and passes a string through. It applies at six sites, not the three mypy first reported -- fixing those unmasked two more that had been hidden behind them. onboard_member.py: the RGBA corner check now tests isinstance before indexing, which behaves as before for a non-tuple since the enclosing except already caught that; and the re.subn lambda became a named function so mypy can pin the AnyStr overload. The publications DOI goes back to 10.31234/osf.io/mhxtd_v1, which resolves. PsyArXiv has not minted the unversioned form, so the shorter one was a dead identifier. The column is still never rendered. Verified by rebuilding: JRM_CV.html is byte-identical after the parser change and both section-note footnotes still render, which is what actually proves it, and all four content pages are unchanged too. extract_people.py has no tests, so its extraction was exercised directly against the live page: 181 rows over 7 sections with every url and image field a string. mypy and ruff are both clean and the suite is 348 passed, 12 skipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QQUWxAFYESZCZe3DLyy1eB
1 parent 72cebb2 commit 50a1bfc

5 files changed

Lines changed: 44 additions & 27 deletions

File tree

‎README.md‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -456,16 +456,22 @@ The `extract_cv.py` script provides a custom LaTeX-to-HTML converter that handle
456456
### Key Functions in extract_cv.py
457457

458458
| Function | Purpose |
459-
|----------|---------|
459+
|-|-|
460460
| `extract_document_body()` | Extract content between `\begin{document}` and `\end{document}` |
461461
| `balanced_braces_extract()` | Parse nested LaTeX braces correctly |
462462
| `convert_latex_formatting()` | Convert LaTeX commands to HTML |
463463
| `parse_etaremune()` | Parse reverse-numbered publication lists |
464464
| `extract_header_info()` | Extract name and contact information |
465465
| `extract_sections()` | Split document into sections/subsections |
466+
| `extract_footnote()` | Pull `\blfootnote{}` out of a section, returning `(footnote_html, remaining_content)` |
467+
| `preprocess_content()` | Strip comments, the footnote, `\vspace`, and the "Last updated" line before parsing; returns `(footnote, cleaned_content)` |
466468
| `render_section_content()` | Convert section content based on type |
467469
| `generate_html()` | Assemble complete HTML document |
468470

471+
> `preprocess_content()` is the single entry point for footnote handling — it always
472+
> returns the `(footnote, content)` pair, and `extract_footnote()` strips the
473+
> `\blfootnote` it extracts, so a footnote never survives into the rendered body.
474+
469475
### CV Stylesheet (cv.css)
470476

471477
The stylesheet provides:

‎data/publications.xlsx‎

5 Bytes
Binary file not shown.

‎scripts/extract_cv.py‎

Lines changed: 9 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -8,7 +8,7 @@
88

99
import re
1010
from pathlib import Path
11-
from typing import Dict, List
11+
from typing import Any, Dict, List, Optional, Tuple
1212
from dataclasses import dataclass, field
1313

1414

@@ -280,7 +280,7 @@ def parse_labeled_lists(content: str) -> List[tuple]:
280280

281281
def extract_header_info(body: str) -> Dict[str, str]:
282282
"""Extract header information (name, title, contact)."""
283-
info = {}
283+
info: Dict[str, Any] = {}
284284

285285
# Find header section (before first section)
286286
header_match = re.search(r'^(.+?)\\section\*', body, re.DOTALL)
@@ -428,7 +428,7 @@ def render_labeled_lists(labeled_lists: List[tuple], use_two_column: bool = Fals
428428
return html
429429

430430

431-
def extract_footnote(content: str) -> tuple:
431+
def extract_footnote(content: str) -> Tuple[Optional[str], str]:
432432
"""Extract blfootnote content from text. Returns (footnote_text, cleaned_content)."""
433433
pattern = r'\\blfootnote\{'
434434
match = re.search(pattern, content)
@@ -446,22 +446,17 @@ def extract_footnote(content: str) -> tuple:
446446
return None, content
447447

448448

449-
def preprocess_content(content: str, extract_footnotes: bool = False) -> tuple:
449+
def preprocess_content(content: str) -> Tuple[Optional[str], str]:
450450
"""Preprocess content to handle problematic LaTeX commands before parsing.
451451
452-
If extract_footnotes is True, returns (footnote, cleaned_content).
453-
Otherwise returns just cleaned_content for backwards compatibility.
452+
Returns (footnote, cleaned_content). extract_footnote strips the
453+
\\blfootnote it extracts, so the footnote never survives into the body.
454454
"""
455455
# Remove comments first
456456
content = re.sub(r'^%.*$', '', content, flags=re.MULTILINE)
457457
content = re.sub(r'(?<!\\)%.*$', '', content, flags=re.MULTILINE)
458458

459-
footnote = None
460-
if extract_footnotes:
461-
footnote, content = extract_footnote(content)
462-
else:
463-
# Remove blfootnote (with nested braces)
464-
content = remove_command_with_braces(content, 'blfootnote')
459+
footnote, content = extract_footnote(content)
465460

466461
# Remove vspace
467462
content = remove_command_with_braces(content, 'vspace')
@@ -470,16 +465,14 @@ def preprocess_content(content: str, extract_footnotes: bool = False) -> tuple:
470465
content = re.sub(r'\{\\scriptsize\s*Last updated:.*?\\today\s*\}', '', content, flags=re.DOTALL)
471466
content = re.sub(r'Last updated:\s*$', '', content, flags=re.MULTILINE)
472467

473-
if extract_footnotes:
474-
return footnote, content
475-
return content
468+
return footnote, content
476469

477470

478471
def render_section_content(content: str, section_title: str) -> str:
479472
"""Render section content to HTML based on section type."""
480473

481474
# Extract footnote first (for display as note under section header)
482-
footnote, content = preprocess_content(content, extract_footnotes=True)
475+
footnote, content = preprocess_content(content)
483476

484477
# Build HTML with optional footnote note
485478
html_prefix = ''

‎scripts/extract_people.py‎

Lines changed: 20 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -33,7 +33,7 @@ def extract_people(html_path: Path) -> dict:
3333
role = name_parts[1].strip() if len(name_parts) > 1 else ''
3434

3535
data['director'] = [{
36-
'image': img.get('src', '').replace('images/people/', '') if img else '',
36+
'image': attr_str(img, 'src').replace('images/people/', '') if img else '',
3737
'name': name,
3838
'role': role,
3939
'bio': ps[0].get_text(strip=True) if ps else '',
@@ -95,7 +95,7 @@ def extract_people(html_path: Path) -> dict:
9595
if link and p.get_text(strip=True):
9696
collab = {
9797
'name': link.get_text(strip=True),
98-
'url': link.get('href', ''),
98+
'url': attr_str(link, 'href'),
9999
'description': p.get_text(strip=True)
100100
}
101101
collaborators.append(collab)
@@ -120,7 +120,7 @@ def extract_person_card(card) -> dict:
120120
if h3:
121121
link = h3.find('a')
122122
if link:
123-
name_url = link.get('href', '')
123+
name_url = attr_str(link, 'href')
124124

125125
text = h3.get_text()
126126
if '|' in text:
@@ -131,7 +131,7 @@ def extract_person_card(card) -> dict:
131131
name = text.strip()
132132

133133
return {
134-
'image': img.get('src', '').replace('images/people/', '') if img else '',
134+
'image': attr_str(img, 'src').replace('images/people/', '') if img else '',
135135
'name': name,
136136
'name_url': name_url,
137137
'role': role,
@@ -204,7 +204,7 @@ def extract_alumni_list(elem) -> list:
204204
# Get name - either from link or from text before parenthesis
205205
if name_link:
206206
alum['name'] = name_link.get_text(strip=True)
207-
alum['name_url'] = name_link.get('href', '')
207+
alum['name_url'] = attr_str(name_link, 'href')
208208
else:
209209
# No name link, name is text before (
210210
if paren_pos > 0:
@@ -214,7 +214,7 @@ def extract_alumni_list(elem) -> list:
214214

215215
# Get position link URL if available
216216
if position_link:
217-
alum['current_position_url'] = position_link.get('href', '')
217+
alum['current_position_url'] = attr_str(position_link, 'href')
218218

219219
# Parse years and current position from parenthesis
220220
paren_match = re.search(r'\(([^)]+)\)', full_text)
@@ -268,6 +268,20 @@ def extract_alumni_simple_list(elem) -> list:
268268
return alumni
269269

270270

271+
def attr_str(tag: Any, name: str) -> str:
272+
"""Read a single-valued tag attribute as a string.
273+
274+
BeautifulSoup types attribute access as str | AttributeValueList | None,
275+
because a multi-valued attribute such as class comes back as a list. The
276+
attributes read here (src, href) are always single-valued, so join a list
277+
rather than letting one leak into a field declared as text.
278+
"""
279+
value = tag.get(name, '')
280+
if isinstance(value, list):
281+
return ' '.join(value)
282+
return value or ''
283+
284+
271285
def get_inner_html(element) -> str:
272286
"""Get the inner HTML of an element as a string."""
273287
if not element:

‎scripts/onboard_member.py‎

Lines changed: 8 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -640,8 +640,10 @@ def photo_already_processed(photo_base: str, project_root: Path) -> bool:
640640
img.getpixel((0, h - 1)),
641641
img.getpixel((w - 1, h - 1)),
642642
]
643-
# All corners should be fully transparent (alpha == 0)
644-
if not all(c[3] == 0 for c in corners):
643+
# All corners should be fully transparent (alpha == 0). getpixel is
644+
# typed as float | tuple | None for single-band images; mode is
645+
# RGBA here, so each corner is a 4-tuple.
646+
if not all(isinstance(c, tuple) and c[3] == 0 for c in corners):
645647
return False
646648

647649
return True
@@ -1040,9 +1042,11 @@ def close_cv_entry(cv_path: Path, name: str, end_year: str) -> bool:
10401042
# count=1: close the first open entry only. Closing every one at a stroke
10411043
# would rewrite an unrelated second open entry for the same name, which is
10421044
# never what a single role change means.
1045+
def close_range(m: "re.Match[str]") -> str:
1046+
return f"{m.group(1)}{end_year})"
1047+
10431048
new_content, count = re.subn(
1044-
pattern, lambda m: f"{m.group(1)}{end_year})", content, count=1,
1045-
flags=re.IGNORECASE
1049+
pattern, close_range, content, count=1, flags=re.IGNORECASE
10461050
)
10471051
if count == 0:
10481052
return False

0 commit comments

Comments
 (0)