# Content Extraction

This page explains how `Extractor` pulls the title, author, and body out of a news page, and how the word count is calculated.

## Extraction Flow

`NewExtractor()` creates an extractor with a 30-second HTTP timeout; `Get(url)` fetches the page and parses it with goquery:

| Step | Rule |
|------|------|
| Remove noise | Drop `script`, `style`, `nav`, `aside`, `footer`, `.advertisement`, `.ads`, `.comment` |
| Title | The first `<h1>`; falls back to `<title>` |
| Author | Text of the first element matching `[rel='author']`, `.author`, or `[itemprop='author']` |
| Body | See the next section |
| Normalize | Collapse runs of whitespace into a single space |

It returns `model.NewsContent{Title, Author, Content, WordCount}`; on an HTTP or HTML parse error it returns the error and the caller falls back to the RSS description.

## Body Selection

Tried in order; the first match wins:

| Order | Condition |
|-------|-----------|
| 1 | `<article>` text longer than 128 bytes |
| 2 | `<main>` text longer than 128 bytes |
| 3 | Among all `<div>` elements, the longest one over 128 bytes whose link text is under 30% of its text |

The link ratio in step 3 rules out link-heavy blocks such as navigation bars and related-article lists. When none match, the body is an empty string.

## Word Count

`WordCount` is the number of Han characters (`\p{Han}`) plus the number of English words (`[a-zA-Z]+`), counted on the body before normalization.

## Limits

| Case | Behavior |
|------|----------|
| JavaScript-rendered pages | Only the initial HTML is read; client-rendered bodies are missed |
| Paywalls | Returns whatever precedes the paywall |
| HTTP status codes | Not checked; error pages are parsed as the body |
