HTML Entities
Escaped characters like & and <.
Also known as: character entities, &, escaping HTML
HTML entities are codes that represent characters which would otherwise be read as part of the markup, or that are hard to type. They start with & and end with ;.
| Entity | Shows | Why |
|---|---|---|
< | < | otherwise starts a tag |
> | > | |
& | & | otherwise starts an entity |
" | " | inside attribute values |
' or ' | ' | |
| a non-breaking space | keeps words together |
© | © | easy-to-type symbols |
€ or € | € | numeric or named forms |
<p>5 < 10 && 10 > 5</p> <!-- displays: 5 < 10 && 10 > 5 -->
<p>Use the <strong> tag for emphasis.</p>
Why it matters most: escaping user input
If you put text into a page without escaping it, any < or & in it is treated as markup. A user who enters <script>...</script> could run code on your page: that’s cross-site scripting (XSS).
user types: <img src=x onerror=steal()>
unescaped: runs the code
escaped: <img src=x onerror=steal()> shows as harmless text
The fix is to escape output for the HTML context (output encoding):
- Modern frameworks and template engines do it automatically. Use
{{ value }}in a template,textContentin the DOM, JSX braces. - The danger is bypassing it:
innerHTML,dangerouslySetInnerHTML,|safefilters.
import html
html.escape("<b>Tom & Jerry</b>") # '<b>Tom & Jerry</b>'
Other notes
- Escaping for HTML text and for attributes, URLs and JavaScript are different contexts with different rules.
- With UTF-8 you can usually type characters like © or € directly, instead of entities (UTF-8).
- Entities aren’t the same as URL encoding.