HTML Decode
Turn HTML entities back into the characters they represent.
Text that arrives full of & and ' is usually a symptom rather than a formatting choice. Something in the pipeline encoded it twice, and decoding tells you what it was meant to say.
Decoding entities
- Paste the text containing HTML entities.
- Decode it.
- Read the plain-text result.
- If entities remain, decode again — that confirms double encoding.
- Trace back to whichever step is encoding an already-encoded value.
How double encoding happens
A CMS escapes user input on the way into the database. The template then escapes it again on the way out, because the template is doing the right thing and does not know the value was already escaped. The result is & appearing on the page where an ampersand belongs.
Seeing & means it happened three times. The fix is never to decode at display time as a workaround — that reintroduces the injection risk the escaping was preventing. Fix the layer that is escaping too early, which is almost always the one writing to storage.
Where you will need to decode
- Cleaning up content imported from another CMS or an RSS feed.
- Reading API responses from services that escape their JSON string values.
- Fixing product titles and descriptions after a migration.
- Reading email HTML source to see what a message actually said.
- Debugging a template that is escaping something it should not.
Named versus numeric entities
Both forms mean the same thing. < and < both produce a less-than sign; the first is a named reference and the second points at the character code directly. Numeric references can express any Unicode character, which is why you see them for symbols and emoji that have no convenient name.
HTML decoding questions
Why does my text show & instead of &?
Because it was HTML-encoded twice. The original ampersand became &, and a second pass encoded that ampersand again into &amp;, which renders as & on the page.
Is it safe to decode HTML entities?
Decoding as a diagnostic step is fine. Decoding untrusted content and then inserting it into a page is not — that is precisely how cross-site scripting works. Keep escaped content escaped when you render it.
What is the difference between named and numeric entities?
None in effect. Named entities such as < are easier to read; numeric ones such as < can represent any Unicode code point, including characters with no assigned name.
Why do I see in content pasted from Word?
Word and many rich-text editors insert non-breaking spaces liberally. They are legitimate characters but cause uneven spacing, so they are usually worth replacing with ordinary spaces.
Can I decode a whole HTML document?
You can, but it would destroy the markup by turning escaped examples into real tags. Decode only the text fragments you are inspecting, not an entire page.