# Spec Issues: (Character) Entity References

**URL:** <https://talk.commonmark.org/t/spec-issues-character-entity-references/2306>\
**Category:** Spec\
**Created:** [November 29, 2016, 12:08pm UTC](https://talk.commonmark.org/t/spec-issues-character-entity-references/2306 "2016-11-29T12:08:03Z")\
**Posts on this page:** 1\
**Showing post:** 12

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [November 30, 2016, 10:16pm UTC](https://talk.commonmark.org/t/spec-issues-character-entity-references/2306/12 "2016-11-30T22:16:01Z")

</div>

I forgot in the above to point out one more good reason why a _CommonMark_ (or _Markdown_) processor should **not** be required to undiscriminatingly replace all _character entity references_ and not even all _numeric character references_, and this has nothing to do with encodings, fonts, or “exotic” Unicode characters:

When a _CommonMark_ author writes, for example, `&verbar;` or `&#124;` instead of `|` for the U+007C&nbsp;VERTICAL&nbsp;LINE in some text, he probably has a reason for making this distinction: Very likely, the literal `|` character will have some significance in a post-processing tool (maybe to separate columns of some sort), which the `&#124;` will not entail. (Note that is the very same distinction between `[` and `\[` in _CommonMark_.)

And it seems quite unhelpful to **mandate** that a _CommonMark_ processor, when generating some sort of XML/HTML/SGML/DocBook etc. output, should obliterate this distinction, hereby making such post-processing much harder if not impossible.

So a simple recommendation would be that at least _entity references_ and _character references_ to ASCII characters **should** be preserved and reproduced in the processor’s output (at the user’s option, maybe?).

---

_[View the full topic](https://talk.commonmark.org/t/spec-issues-character-entity-references/2306)._
