If soup.select_one(...).text raises AttributeError, the likely problem is that the selector found no element, not that .text itself is broken. select_one() returns None when nothing matches. Start by testing that boundary with a small HTML string.
from bs4 import BeautifulSoup
def read_title(html: str) -> str:
soup = BeautifulSoup(html, "html.parser")
title = soup.select_one("h1.article-title")
if title is None:
raise ValueError("Title element not found")
return title.get_text(strip=True)
assert read_title('<article><h1 class="article-title">New post</h1></article>') == "New post"h1.article-title selects an h1 element with class article-title. Call .text or get_text() only after you have an element. An input with a changed class should fail explicitly:
try:
read_title('<article><h1 class="headline">New post</h1></article>')
except ValueError as error:
assert "Title element" in str(error)
else:
raise AssertionError("A missing required title went unnoticed")
soup = BeautifulSoup('<article><h1 class="article-title">New post</h1></article>', "html.parser")
assert soup.select_one("h2.subtitle") is None # An optional subtitle may be absent.Inspect the HTML you actually received
With the same selector, a matching class produces an element and a changed class produces None. A required title should fail explicitly when missing. The ellipsis in the diagram stands for omitted HTML content.
What you see in a browser may differ from the HTML returned by an HTTP request. If JavaScript renders content later, a correct selector can still find no title in the response body. First check the response status and a small portion of the received HTML for article-title. If it is present, compare the selector's tag and class with the actual structure. Avoid logging the entire response, which may contain personal or session data.
Decide whether absence is an error
If some list items legitimately have no subtitle, keep that field as None and count omissions. If a title is required, fail as the example does so that a structural change is visible. Converting every missing field to an empty string hides the difference between an intentionally empty value and failed extraction.
Before collecting from a real site, check the allowed scope and request rate. These examples parse only local HTML strings and make no network requests, so they isolate selector behavior. When provenance matters, keep the data source and collection time alongside extracted values to help trace later page changes.
Key takeaways
select_one() returns None for no match. Separate required from optional fields before calling .text, and inspect both the received HTML and the selector. Observing missing fields helps prevent a crawler from silently collecting bad data after a page layout changes.

