Content Extraction
Last updated
This page explains how Extractor pulls the title, author, and body out of a news page, and how the word count is calculated.
Extraction Flow
NewExtractor() creates an extractor with a 30-second HTTP timeout; Get(url) fetches the page and parses it with goquery:
| Step | Rule |
|---|---|
| Remove noise | Drop script, style, nav, aside, footer, .advertisement, .ads, .comment |
| Title | The first <h1>; falls back to <title> |
| Author | Text of the first element matching [rel='author'], .author, or [itemprop='author'] |
| Body | See the next section |
| Normalize | Collapse runs of whitespace into a single space |
It returns model.NewsContent{Title, Author, Content, WordCount}; on an HTTP or HTML parse error it returns the error and the caller falls back to the RSS description.
Body Selection
Tried in order; the first match wins:
| Order | Condition |
|---|---|
| 1 | <article> text longer than 128 bytes |
| 2 | <main> text longer than 128 bytes |
| 3 | Among all <div> elements, the longest one over 128 bytes whose link text is under 30% of its text |
The link ratio in step 3 rules out link-heavy blocks such as navigation bars and related-article lists. When none match, the body is an empty string.
Word Count
WordCount is the number of Han characters (\p{Han}) plus the number of English words ([a-zA-Z]+), counted on the body before normalization.
Limits
| Case | Behavior |
|---|---|
| JavaScript-rendered pages | Only the initial HTML is read; client-rendered bodies are missed |
| Paywalls | Returns whatever precedes the paywall |
| HTTP status codes | Not checked; error pages are parsed as the body |