English

Developer tools · HTML entity escaper

Why one unescaped ampersand breaks an entire RSS or XML feed

· Why it matters

html xml encoding

Why one unescaped ampersand breaks an entire RSS or XML feed shown as a browser-safe character-reference diagram
Original ToolAcre vector illustration

XML is strict where HTML is forgiving: an unescaped & or an HTML-only entity like   makes the whole document unparseable. This post explains the five XML entities, why   is invalid, and how to escape content for feeds.

The feed that vanished from every reader after one headline — 'Tips & Tricks' and a not-well-formed error

The feed that vanished from every reader after one headline — 'Tips & Tricks' and a not-well-formed error. An RSS title containing Tips & Tricks is not well-formed XML because the ampersand begins a reference that never resolves. A strict parser can reject the complete feed rather than repair it.

To verify rss feed ampersand error, construct the feed that vanished for a publisher whose feed fails validation after a single post. Preserve from every reader after while strict XML feed produces one headline tips tricks; identify where and a not well is consumed. The observation about formed error belongs to HTML text only.

XML's five predefined entities — < > & " ' and nothing else without a DTD

XML's five predefined entities — < > & " ' and nothing else without a DTD. XML predefines amp, lt, gt, quot and apos. Numeric references are also available. Unlike HTML, arbitrary familiar names are not automatically recognized without a declaration.

A publisher whose feed fails validation after a single post can test xml s five predefined by recording entities lt gt amp before the strict XML feed pass. Compare quot apos and nothing afterward and locate the parser responsible for else without a dtd. This rss feed ampersand error result explains strict XML feed evidence, not executable contexts.

Why  , © and — break XML — HTML names that XML has never heard of, and the numeric alternative

Why  , © and — break XML — HTML names that XML has never heard of, and the numeric alternative. Names such as nbsp, copy and mdash belong to HTML’s table, not XML’s five predefined names. In XML, use the literal UTF-8 character when allowed or a numeric reference such as  .

Isolate why nbsp copy and in a short strict XML feed sample. Show mdash break xml html as literal source, follow names that xml has to its destination, and name the API reading never heard of and. For rss feed ampersand error, the numeric alternative remains parser-bound evidence.

Escaping HTML inside a feed — content that is itself markup must be escaped again or wrapped in CDATA

Escaping HTML inside a feed — content that is itself markup must be escaped again or wrapped in CDATA. Markup embedded as data inside a feed must either be escaped as text or placed correctly in CDATA according to the feed design. Mixing approaches can produce double escaping or accidental structure.

Treat escaping html inside a as a boundary experiment. A publisher whose feed fails validation after a single post should retain feed content that is, perform one strict XML feed operation, and inspect itself markup must be character by character before changing escaped again or wrapped. The claim about in cdata stops at this HTML layer.

Worked example: fixing a broken item — the offending characters, the corrected entities and the validator's result

Worked example: fixing a broken item — the offending characters, the corrected entities and the validator's result. For a title Tips & Tricks <Draft>, minimal mode yields Tips &amp; Tricks &lt;Draft&gt;. Those substitutions make the text representable in XML character data as well as HTML text.

Reproduce worked example fixing a with harmless input instead of customer material. Record broken item the offending, observe characters the corrected entities, and count every intentional strict XML feed pass. That rss feed ampersand error trail lets a publisher whose feed fails validation after a single post evaluate and the validator s and result without guessing.

Common mistakes — fixing the & but leaving a &nbsp; from a CMS editor, and double-escaping CDATA

Common mistakes — fixing the & but leaving a &nbsp; from a CMS editor, and double-escaping CDATA. Replacing one raw ampersand while leaving &nbsp; still fails because the parser encounters an undeclared name. Conversely, escaping content already protected by CDATA can leave visible entity text.

Place common mistakes fixing the, but leaving a nbsp, and from a cms editor side by side during the strict XML feed review. A publisher whose feed fails validation after a single post can then decide whether and double escaping cdata changed at conversion or downstream. Keep the rss feed ampersand error conclusion about strict XML feed evidence out of generic security claims.

What this does not cover — Atom versus RSS differences and feed reader rendering quirks

What this does not cover — Atom versus RSS differences and feed reader rendering quirks. Atom-versus-RSS vocabulary and reader rendering are separate from basic XML well-formedness. Validate the final document with the same namespaces and serialization used in production.

Define what this does not before running strict XML feed. Save cover atom versus rss as a control, inspect the code points behind differences and feed reader, and map rendering quirks to the next interpreter. This makes strict XML feed evidence auditable for a publisher whose feed fails validation after a single post investigating rss feed ampersand error.

Takeaway: minimal mode covers XML markup characters; named mode is HTML-specific

Takeaway: minimal mode covers XML markup characters; named mode is HTML-specific. Use minimal mode for markup-critical characters and review XML-specific requirements. Named mode can emit HTML-only names, so it is not advertised as a general XML serializer.

Connect takeaway minimal mode covers to an observable strict XML feed output. Keep xml markup characters named beside the one-pass result, then verify where mode is html specific enters strict XML feed evidence. A publisher whose feed fails validation after a single post can now review strict XML feed evidence as a narrow rss feed ampersand error finding. The practical decision behind this article is specific: XML is strict where HTML is forgiving: an unescaped & or an HTML-only entity like &nbsp; makes the whole document unparseable. This post explains the five XML entities, why &nbsp; is invalid, and how to escape content for feeds. The reader action is equally concrete: Links to the HTML entity escaper and demonstrates escaping a headline containing & and < for use in a feed item.