Optionally keep original title headers for main content extraction accuracy#1006
Optionally keep original title headers for main content extraction accuracy#1006mcPear wants to merge 3 commits into
Conversation
Keep original title headers for SEO reverse-engineering research accuracy
gijsk
left a comment
There was a problem hiding this comment.
Thanks for submitting a PR.
I'm confused about this being labeled a "chore". And I think that the relabeling of h1 to h2 and removing headers that match the titles should be split into different options. The h1-to-h2 replacement would fix #863 . I would be tempted to make not swapping it the default, though it would potentially break existing consumers - I can fix Firefox but not sure about others.
gijsk
left a comment
There was a problem hiding this comment.
Apologies for the very unreasonable delay here.
I'm fairly confused at this point. The updated patch here added a lot more code which is a bit hard to read, and also doesn't seem consistent with what I asked in my previous review. Instead, it seems to change the approach used to keep the headers. Why is that?
|
Sorry, I had to adjust the logic to fix corner cases we found in our app. I wasn't aware you will see it as kinda response to this conversation. I'm closing it as for now. |
Summary
Reader extraction currently rewrites all in-article
h1elements toh2so the article title can remain the sole top-level heading in the reader UI. Moreover, removes the first similar heading spotted after the title. These normalizations improve classic “reader mode” presentation but weaken the semantic outline of the page: crawlers, SEO tooling, and systems that infer structure from HTML (including retrieval and “reverse engineering” of how a page is organized) rely on stable heading levels that match the publisher’s markup.This change preserves the original heading tag names and levels in the extracted content wherever we are not explicitly removing noise, so the serialized article HTML stays closer to the source document’s hierarchy. All that is gated behind an option.
What changes (high level)
h1→h2replacement in article content