# Ordered vs Unordered List

**URL:** <https://talk.commonmark.org/t/ordered-vs-unordered-list/1945>\
**Category:** Spec\
**Created:** [December 24, 2015, 5:33am UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945 "2015-12-24T05:33:25Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [December 24, 2015, 5:33am UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/1 "2015-12-24T05:33:25Z")

</div>

Contrary to _[emphasis](http://talk.commonmark.org/t/weak-vs-strong-emphasis/1943)_, the term _bullet list_ bears no ambiguity. However, similarly to _[horizontal rule](http://talk.commonmark.org/t/horizontal-rule-or-thematic-break/912/11)_ RIP, it implies certain presentational semantics.

I therefore propose to replace all occurrences of _bullet list_ with _unordered list_.

---

<div class="post-metadata">

**Author:** ![Crissov](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/crissov/32/1455_2.png) [@Crissov](https://talk.commonmark.org/u/Crissov)\
**Post date:** [December 24, 2015, 8:06pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/2 "2015-12-24T20:06:25Z")

</div>

Nah, “unordered” isn’t really any better, because often list items are in fact pre-ordered, but often in an arbitrary, non-immanent way (e.g. alphabetic collation). Likewise, many “ordered” lists are enumerated without any particular reason.

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [December 24, 2015, 10:08pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/3 "2015-12-24T22:08:46Z")

</div>

According to [W3C](http://www.w3.org/TR/html-markup/ul.html):

> The `ul` element represents an unordered list of items; that is, a list in which changing the order of the items would not change the meaning of list.

---

<div class="post-metadata">

**Author:** ![Crissov](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/crissov/32/1455_2.png) [@Crissov](https://talk.commonmark.org/u/Crissov)\
**Post date:** [December 28, 2015, 4:54pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/4 "2015-12-28T16:54:32Z")

</div>

Sure, that’s `ul` in HTML, but I doubt people are actually using bullet lists in Markdown that way at all times.

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [December 29, 2015, 1:09pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/5 "2015-12-29T13:09:42Z")

</div>

I believe 99.9% of the people use emphases to make text underlined and bold, and yet we call those _emphasis_ ([ambiguously](http://talk.commonmark.org/t/weak-vs-strong-emphasis/1943)) and _strong emphasis_.

The only valid counter-argument I’ve seen so far came from my browser’s spellchecker, which persistently claims that _unordered_ is not an English word (neither is _inline_). As I don’t feel full confidence in this regard, any comments from native speakers would be herzlich wilkommen.

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [December 31, 2015, 12:16pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/6 "2015-12-31T12:16:53Z")

</div>

## The `UL` and `OL` lists

> […] “unordered” isn’t really any better, because […]

We can certainly try to come up with better names for the `UL` and `OL` list types, but I really can’t see any merit in it: this seems rather like an attempt to rewrite history.

As far as I can reconstruct, the GIs `LI`, `UL`, and `OL` (where `UL` means _“unordered list”_ and `OL` means “_ordered list_”) have been around for **more than 30 years now** (probably rather 40 years or more), and the history goes like something this:

* * *

### The 70’s – Before SGML

Even before SGML, the precursor [GML](https://en.wikipedia.org/wiki/Generalized_Markup_Language) at IBM had “unordered” and “ordered” lists, and used `UL` and `OL` for the corresponding “tags”.

A hands-on description of the difference in the IBM documentation [reads like this](http://publibfp.dhe.ibm.com/cgi-bin/bookmgr/BOOKS/dsm04m00/2.2.1?DT=19910507103243):

> Unordered lists are similar to simple lists, except that each item in an unordered list is preceded by a special symbol. (The special symbol used depends on the device on which you are having your output printed.) You would use an unordered list when the items in the list are fairly long, maybe even many paragraphs, but don’t need to be in a specific sequence.

And for `OL`:

> And then there’s the ordered list. Use an ordered list when the items you’re listing need to be in a specific sequence. An example of an ordered list follows. […]  
> When you create an ordered list, you don’t have to number the items yourself. The starter set does it for you. This saves you a lot of work when you decide to insert, delete, or rearrange items.

(The _“starter set”_ is the _“GML starter set (GMLSS)”_ product.)

### The 80’s – SGML

Many of the GMLSS elements were adopted into the _“general document”_ DTD example in annex&nbsp;E of [ISO&nbsp;8879:1986](http://www.iso.org/iso/catalogue_detail.htm?csnumber=16387) (the very first SGML standard). As this DTD is in turn the very first “official”, publicly available, “document type”, one can truly say that `LI`, `UL`, and `OL` were already in use when “angle-bracket-tags” were invented and standardized.

The _“general document”_ DTD has actually _six kinds_ of list, of which only three (`OL`, `UL`, and `DL`—but not `SL`, `NL`, and `GL`) later made it into HTML. Regarding the “ordered” vs “unordered” meaning, the “user manual” for SGML, [ISO/IEC/TR&nbsp;9573:1988](http://www.iso.org/iso/home/store/catalogue_tc/catalogue_detail.htm?csnumber=17319), describes the difference like this (in section&nbsp;5.3.10):

> The “general” document includes six types of lists:
> 
> - “ordered” list, where there may be a need to refer to each item in the list. Typically the list items are numberred by the text-formatter.
> 
> - “unordered” list, where there is no need to refer to any item in the list, but where each item should be clearly standing out. Typically the list items are indicated by bullets, stars, or dashes by the text-formatter.
> 
> […]

I find these descriptions perfectly reasonable and not too presentation-biased. (The DTD has a `LIREF` element type dedicated to referencing list items, which is _declared empty_, so the intended meaning of _“where there may be a need to refer to each item”_ seems to be that `LIREF` can be used to refer to `OL` list items, but should _not_ or even _can not_ be used to refer to `UL` items—although these can have an `ID` attribute too.)

### The 90’s – Early WWW and HTML

These “types” of lists were then _already inherited_ (via the [AAP tag set](http://info.cern.ch/hypertext/WWW/MarkUp/AAP.html)) in the [very first HTML sketches](http://www.w3.org/History/1991-WWW-NeXT/Implementation/test.html) by Tim Barners-Lee. Note that in this 1992 specimen there is also the `XMP` element type from the _“general document”_ DTD, with roughly the use and purpose of the later `PRE` and `CODE` element types. And the tags `HP1`, `HP2`, etc—also from the _“general document”_ DTD—for _highlighted phrases_ are also mentioned (qualified as “not currently used”) [in early descriptions](http://info.cern.ch/hypertext/WWW/MarkUp/Tags.html) of HTML.

The HTML&nbsp;2 specification dated 1995-09-22 used a [different wording](http://www.w3.org/MarkUp/html-spec/html-spec_5.html#SEC5.6.1), leaning more to the “presentational” side, for the two list types. (I could not find [earlier HTML descriptions](http://www.w3.org/MarkUp/html-spec) than version 2.0;)

### The 90’s – Stable HTML

Then [HTML&nbsp;3](http://www.w3.org/MarkUp/html3/bulletlists.html) added “customization” attributes for list elements, so that one [can choose](http://www.w3.org/TR/REC-html32#ul) the “marker” style for list items. One more move in the “presentational markup” direction, in the absence of CSS or similar techniques. Consequently, the specification does not even bother to explicitly define any “semantic” differences between “ordered” or “unordered” lists, but only explains the _rendering_ differences instead.

The later [HTML&nbsp;4.01](http://www.w3.org/TR/html401/struct/lists.html) specification is again more explicit, and gives a distinction which is not _purely_ presentational:

> An ordered list, created using the `OL` element, should contain information where order should be emphasized, as in a recipe: […]

But the only “real” difference is again presentational:

> Ordered and unordered lists are rendered in an identical manner except that visual user agents number ordered list items. User agents may present those numbers in a variety of ways. Unordered list items are not numbered.

### The 2010’s – Shiny new HTML 5

It was only in [HTML&nbsp;5](http://www.w3.org/TR/html5/grouping-content.html#the-ul-element) that the—IMO silly—attempt to define a _semantic_ difference was made in this manner:

> The `OL` element represents a list of items, where the items have been intentionally ordered, such that changing the order would change the meaning of the document.

And accordingly, for `UL`:

> The `UL` element represents a list of items, where the order of the items is not important — that is, where changing the order would not materially change the meaning of the document.

Taking these descriptions seriously would IMO mean that a user agent or any other application is free to re-order `UL` items, say alphabetically. This has always been the case for the order or _attribute specification lists_, but I doubt that many authors would be happy if their `UL` items were randomly re-ordered with the excuse that this “does not change the meaning of the document” …

* * *

To conclude this little excursion:

- The terms “ordered list” and “unordered list” have been in use for literally decades,

- but with quite some variation in explicitly or implicitly defined meaning.

- But the pseudo-precise distinction made in HTML&nbsp;5 based on “re-ordering does/does not change the _meaning of the document_” is IMO the _least_ useful definition for these terms (what is the “meaning of a document”, after all?).

So probably the best thing to do in a _CommonMark_ specification is to just talk about “ordered list” and “unordered list”, and leave the meaning of these terms implicitly defined by the fourty-year long common usage of these words …

---

<div class="post-metadata">

**Author:** ![Crissov](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/crissov/32/1455_2.png) [@Crissov](https://talk.commonmark.org/u/Crissov)\
**Post date:** [December 31, 2015, 2:14pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/7 "2015-12-31T14:14:38Z")

</div>

Latex uses `enumerate` and `itemize` (and `description`) – so what?  
Why should HTML or \*gasp\* SGML terminology matter so much (more) for CM?

Since CM authors need to type either a bullet marker `+`/`*`/`-` or a number followed by an ordinal marker `.`/`)`, I think _bullet list item_ and _enumerated_ or _enumeration list item_ make a lot of inherent sense. The lists consisting of such items would naturally be called _bullet lists_ and _enumerated/enumeration lists_.

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [December 31, 2015, 6:53pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/8 "2015-12-31T18:53:40Z")

</div>

> Why should HTML or _gasp_ SGML terminology matter so much (more) for CM?

Mostly because XML/SGML **“technology”** (note that “HTML” is a different _category_!) **does** matter so much for CM.

I think this and related questions boil down to the decision on how much the _CommonMark_ specification and syntax should be tied to:

1. one or another version of HTML and the terminology used in the particular HTML specification, and

2. the markup syntax and structural concepts of XML and SGML in general.

I certainly do agree that _CommonMark_ should be specified in the most general terms and as independently from any _particular document type_ (like HTML) as possible, in order to allow a wide range of use scenarios and “target” _document types_; and in fact the specification right now has IMO already (or still) too much “HTML bias”.

I would even prefer to _define_ _CommonMark_ completely in terms of its own, “native”, document model, for example using the _CommonMark_ DTD. Mapping this into HTML in a “canonical” way is trivial to specify anyway.

But I just see no good reason for nor a reasonable way of making _CommonMark_ independent from the the XML/SGML markup languages1) in general, for the simple reason that the sole purpose of _Markdown_ and hence _CommonMark_ was and is to provide an “author-friendly” syntax for creating HTML/XML/SGML content.

Just note how much the specification is concerned with _character references_ and “raw HTML”, including _comment declaration_, _processing instructions_, _markup declarations_! Without these “loopholes”, you’d end up with either a very limited syntax subset with very restricted expressive power, or with using whatever _other_ syntax (maybe LATEX, maybe `eqn` and `tbl`, or whatever) for the “missing features” of _CommonMark_.

And it was indeed an _explicit_ design choice (by Gruber) for _Markdown_ to allow the use of “raw” HTML syntax, which means from my point of view, and by extension: allowing to use “raw” XML/SGML syntax for these “missing features”.

It sure would be interesting to design a _Markdown_ syntax variant without _any_ recurrence to even XML/SGML, but _CommonMark_ is not it; and this would basically require to come up with a complete new (extensible!) markup language: which amounts to a complete alternative syntax for XML.

As always, you can process the content authored in _CommonMark_ syntax in any way and transform it into any format you like, even without technically and explicitly generating a marked-up XML document in between: but this still would not make _CommonMark_ unrelated to XML/SGML, wouldn’t you agree?

So why should the _CommonMark_ specification suddenly avoid terms—in particular names for document elements which have been in common use even long before HTML was inventend—from precisely this field of markup languages, and start borrowing eg from LATEX, terminology? Or `nroff`? Or RTF? Or SCRIBE, or _OpenOffice_, or _MS Word_ and so on?

That said, I don’t think the choice between “ordered list” and “enumerated list” rsp “unordered list” and “itemized list” is very important, or would really matter much or would enhance or degrade the spec in any significant way—beyond that your preferred terms feel rather arbitrary and “foreign” in this context, in my opinion. (I’m not saying that one choice is “better” than the other!)

But I do think the fundamental question discussed above is in fact an important one for _CommonMark_.

* * *

1. More precisely: the _syntax_ of XML respectively the _reference concrete syntax_ of SGML, and the corresponding content model (a small subset of the XML Infoset rsp SGML ESIS).

---

<div class="post-metadata">

**Author:** ![jgm](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/jgm/32/1360_2.png) [@jgm](https://talk.commonmark.org/u/jgm)\
**Post date:** [December 31, 2015, 7:20pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/9 "2015-12-31T19:20:08Z")

</div>

> [@tin-pot](#):
>
> for the simple reason that the sole purpose of Markdown and hence CommonMark was and is to provide an “author-friendly” syntax for creating HTML/XML/SGML content.

This was, I think, the sole purpose of Markdown originally, as John Gruber conceived of it. But it is certainly not the sole purpose of Markdown variants today, and it’s certainly not the sole purpose of CommonMark.

Indeed, I think one of the nicest things about Markdown/CommonMark is the ability to have a single source document that can be rendered with perfect accuracy into many different formats. I regularly convert Markdown to PDF (via LaTeX), man pages, HTML, EPUBs, and Word documents. I know others who convert Markdown to ICML for import into page layout software. Markdown is rooted in an HTML-generating script, and the ability to include raw HTML reflects that initial purpose. But the idea that Markdown is primarily useful as a way of generating HTML is something we’ve grown out of.

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [December 31, 2015, 8:15pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/10 "2015-12-31T20:15:30Z")

</div>

> [@jgm](#):
>
> I regularly convert Markdown to PDF (via LaTeX), man pages, HTML, EPUBs, and Word documents. I know others who convert Markdown to ICML for import into page layout software.

How many of these (excluding Markdown) have _bullet lists_? I know Word does (since 1983), but how about the others?

I believe CommonMark should use the most accurate terminology, not the most common (no pun intended) or, for that matter, the oldest. A project I’m working on has a format (not invented by me) that defines an _unbulleted list_. Is that a special case of bullet list? Should the project’s documentation refer to **unbulleted bullet lists**?

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [December 31, 2015, 9:47pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/11 "2015-12-31T21:47:22Z")

</div>

> Indeed, I think one of the nicest things about Markdown/CommonMark is the ability to have a single source document that can be rendered with perfect accuracy into many different formats.

I completely agree, and isn’t this also the nicest thing about XML (and SGML) too? And doesn’t _Markdown_/_CommonMark_ basically _inherit_ this feat from there? You don’t generate ePUB (XHMTL) from `cmark`’s LaTeX output, do you?

> But the idea that Markdown is primarily useful as a way of generating HTML is something we’ve grown out of.

Absolutely right: that’s why I’d prefer to minimize _CommonMark_’s reliance on HTML specific _character names_, _element type names_, syntactic peculiarities and so on.

But even if you convert _CommonMark_ “directly” into LaTeX (via `cmark` say): this is simply a technical shortcut and an alternative to producing first an (for example “native” _CommonMark_ DTD) XML document from the input text, and then transforming this XML document (or parsed content thereof) to LaTeX or whatever target format you have.

Consider for example the _character reference_ syntax in _Markdown_ and _CommonMark_: it is **not** simply a “reflection of the initial purpose” of _Markdown_ to generate HTML, but rather essential for representing and entering Unicode characters in a portable way.

And similarly, the easiest way to support the indeed “nice thing” of single-source publication from _CommonMark_ input text is IMO to use application-specific _entities_ and _tags_ directly in _CommonMark_ for stuff outside the feature set of the syntax proper (just like using “raw HTML”, but in no way restricted to HTML!). Much better at least than introducing and using “foreign syntax” for which the transformation into PDF, ePUB or what target format you have takes a whole different route—while the latter is a reasonable alternative if one (post- or pre-)processes this “foreign syntax” into XML, and in this way again transforms “extended” _CommonMark_ into XML.

Both approaches assume a target _document type_ or “target format” which is at least _representable_ in XML (which is practically _always_ the case); and the first approach obviously depends on the _Markdown_ and _CommonMark_ property to support the “raw markup loophole”.

Again, this has **nothing** to do with HTML or “vestigial” support for it, but is IMO essential for _Markdown’s_ and _CommonMark’s_ generality and usefulness: There simply is no other widely supported, standardized, generalized and extensible markup language than—well, you know the acronyms 😉

I’m a bit puzzled by you seemingly asserting that the relationship between _CommonMark_ and SGML/XML is basically a historical accident (of John Gruber’s whim) from which we “have grown out of”: on the contrary, I see the **core reason** for the flexibility and general usefulness of _Markdown_ and _CommonMark_ precisely in that relationship … (Note that I’m **not** talking about HTML here!)

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [December 31, 2015, 10:20pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/12 "2015-12-31T22:20:55Z")

</div>

There are two different things or concepts for which terminology is needed, and which are easy to confuse:

1. The syntactic construct in the _CommonMark_ input text: here the “list item” starts with say a “`-`” HYPHEN-MINUS character, typically typed in by our esteemed author him- or herself. Technically, this is a non-terminal symbol of a _concrete syntax_, and the specification will certainly need to talk about this, and could do so in freely chosen (and hopefully easy to understand) terms;

2. The “meaning”, or interpretation, or translation result which corresponds to this input text: that one is a more abstract thing, which could be described as a node in some kind of _Abstract Syntax Tree_, or equivalently as a _non-terminal_ of an _abstract syntax_, or as (an instance of) an element type from the _CommonMark_ DTD, or as the “parsed content” of such an element, and so on: **this** is where the “prior art” of having `UL` and `OL` list element types applies, and **here** the only thing that matters is the element’s attributes and content (and it’s “intended meaning”, however fuzzy), but **not** whether it is rendered with a “bullet” or a “pointing hand” dingbat or a “hypen-minus” or whatever style of marker, or in what font style.

My remarks were all pertinent to the **second** concept only, and I freely admit that for the **first** concept there may be better, easier-to-understand, more fitting terms.

* * *

Btw, the _“unbulleted list”_ of your project would be the “simple list”, or `<SL>`, element type in the mentioned “general document” `<!DOCTYPE general PUBLIC "ISO 8879:1986//DTD General Document//EN">`: also one of the more than 30 years old list types, and also already provided back then in GML on the IBM/360.

If you ask me, this is a good name to use in the documentation. Just sayin’ … 😉

[_Disclaimer:_ I’m not (quite) as old as I possibly seem here, and I wasn’t around at the time of GML etc, or probably still shat into my diapers during those days …]

---

<div class="post-metadata">

**Author:** ![jgm](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/jgm/32/1360_2.png) [@jgm](https://talk.commonmark.org/u/jgm)\
**Post date:** [January 1, 2016, 4:52am UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/13 "2016-01-01T04:52:27Z")

</div>

@tin-pot, yes, XML is flexible enough for representing the structure of CommonMark documents. But that doesn’t give CommonMark any closer relation to XML than to any of a multitude of other general markup languages that are equally capable of representing this abstract structure.

Anyway, this whole thread is about a simple matter of terminology: should we call these things bullet list or unordered lists or itemized lists or something else? I don’t care too much. If people strongly prefer “unordered list,” I can go with that. But I don’t want to get distracted from the many more substantive issues that still need to be resolved.

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [January 1, 2016, 1:24pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/14 "2016-01-01T13:24:20Z")

</div>

> [@jgm](#):
>
> Anyway, this whole thread is about a simple matter of terminology: should we call these things bullet list or unordered lists or itemized lists or something else? I don’t care too much.

Most people on this forum seem to share your sentiment.

Let me present the (hopefully) ultimate argument. The spec does not refer to _ordered lists_ as _number lists_ (or _numbered lists_, or _numeric lists_), although they are marked exclusively (in core CommonMark) by ASCII decimal digits **in both input and output**. On the other hand, _unordered list_ items are only **rendered** with bullet markers (in most formats, at that), but only have `-`, `+` or `*` markers in the input (at least until [Unicode bullets](http://talk.commonmark.org/t/unicode-character-bullet-u-2022/) are part of the spec). Why _bullet lists_ then?

@tin-pot, I hope my “old terminology” remark hasn’t offended you, and I do apologize if it has. I never meant to imply `old == bad` (especially since my proposal seems to be in complete accord with SGML), only that terms should be judged on merit alone ([proof](http://talk.commonmark.org/t/weak-vs-strong-emphasis/1943/3)). Anyway, the older an engineer, the more spec term changes he proposes 😉

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [January 1, 2016, 1:36pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/15 "2016-01-01T13:36:38Z")

</div>

@jgm:

First of all, I’d like to wish you and everybody around here a **happy new year!**

* * *

> [@jgm](#):
>
> But that doesn’t give CommonMark any closer relation to XML than to any of a multitude of other general markup languages that are equally capable of representing this abstract structure.

If I understand your remark correctly, you do include LaTeX, RTF, nroff etc among the “multitude of other general markup languages” here? If that’s the case, then I have a pretty different view on _Markdown_ in general and _CommonMark_ in particular. Mine is based for the most part on the extent to which the _Markdown_ description and _CommonMark_ specification **do** support/include/refer to/borrow from/rely on “HTML” or generally XML syntax, structure, and notions—which you seem to regard as merely inherited baggage, so to say? After all, none of these “XML support features” are useful for creating any of these “other general markup languages” directly from _CommonMark_ input, right?

* * *

But anyway, what really matters though is the _definition_ of _CommonMark_, as expressed in the specification: and I see no reason why the same specification (more or less as it is right now) couldn’t be “compatible” with both points of view.

[_To be precise:_  
As I said, I’m primarily referring to the _document content model_ of XML/SGML anyway—formally the [XML Infoset](http://www.w3.org/TR/xml-infoset/#infoitem.element) rsp [SGML ESIS](http://xml.coverpages.org/WG8-n931a.html#n931-esis) as “upper bounds”1) of what _CommonMark_ can express, but the _far, far simpler_ [“data model” of µXML](https://dvcs.w3.org/hg/microxml/raw-file/tip/spec/microxml.html#data-model) would probably suffice too (being isomorphic to what you refer to as the “AST”)—and I’m not so much concerned with the syntax: if this is what you mean by “representing this abstract structure”, then we **do** largely agree again after all … 😉 And be assured that I certainly don’t want to “require” every _CommonMark_ processor to generate output in XML syntax, ie produce XML documents or fragments!

On the contrary, I think that the _CommonMark_ specification should **not prescribe any syntax at all** for the representation of the parsing result (the _parsed content_, the _AST_, the _content model instance_, the _abstract structure_, whatever you call it). For example, a _CommonMark_ parser could well “represent” the result only as a sequence of invocations of SAX-style _callback functions_, and the manner in which this is done is **clearly outside the scope** of any _CommonMark_ specification. But still the specification must allow to test and decide of a parser of this kind does in fact parse correctly. The same considerations are behind the XML Infoset and ESIS specifications, which was the reason I pointed to RAST and to [canonical XML](http://www.w3.org/TR/xml-c14n) in the following related [discussion](http://talk.commonmark.org/t/use-xml-for-the-spec-examples-and-tests-comments-welcome/994):

But it sure would be useful to have a simple but precise **notation** for the “abstract” parsing result for use in the specification text (and for denoting test case results): that discussion is [over there](http://talk.commonmark.org/t/use-xml-for-the-spec-examples-and-tests-comments-welcome/994)

\_\_\_\_\_\_\_\_

1. The term “upper bound” is used here in an only “half-sloppy way”, because (any reasonable definition of) an _“is-lossless-translatable-to”_ relation between content models would obviously induce a pre-order among content models (and thus document types in general), in which “upper bound” would have the usual, precise [meaning](http://mathworld.wolfram.com/BoundedfromAbove.html).  
]

* * *

> [@jgm](#):
>
> Anyway, this whole thread is about a simple matter of terminology: should we call these things bullet list or unordered lists or itemized lists or something else? I don’t care too much.

Personally, I’d prefer “itemized list” or—equally good—“unordered list” over “bullet list”, because “bullet” refers to a particular glyph (or family of glyphs), thus this term is in my view too much tied to a presentation style. And also because there is relevant “[prior art](http://docbook.org/tdg/en/html/itemizedlist.html)” (and [in LaTeX](https://en.wikibooks.org/wiki/LaTeX/List_Structures#Itemize) too) for the term “itemized list”, while the term “bullet list” seems to be used primarily in and around [MS Office](https://support.office.com/en-us/article/Create-a-vertical-bullet-list-00516CF6-867D-4C4A-94BB-33D0E00BE245) documentation.

* * *

But all in all I don’t care _that_ much either.

However, and more generally, I still think that a distinction in terminology between

- _syntactic parts of input text_ on the one side and
- _output elements/nodes/subtrees/structural parts of the parsed content (or “AST”)_ on the other

could be occasionally helpful in the specification, if only to avoid confusion.

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [January 1, 2016, 3:26pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/16 "2016-01-01T15:26:56Z")

</div>

@Dmitry:

> [@Dmitry](#):
>
> I hope my “old terminology” remark hasn’t offended you, and I do apologize if it has. I never meant to imply old == bad (especially since my proposal seems to be in complete accord with SGML), […]

Haha, don’t worry: rest assured that I didn’t feel offended in the slightest way! But very nice of you to care!

In fact I had a bit of a feeling of being the “old fart” (among youngsters?), when referring to standards and developments from about four decades ago [that’s roughly as old as I am, being around since 1968 …] 😉

* * *

> […] only that terms should be judged on merit alone ([proof](http://talk.commonmark.org/t/weak-vs-strong-emphasis/1943/2)).

Would you accept as “merit” of a term if that term itself is _standardized_, in our case for example in [ISO&nbsp;2382-23 _Vocabulary – Part 32: Text processing_](http://www.iso.org/iso/home/store/catalogue_ics/catalogue_detail_ics.htm?csnumber=63598), or if the term is defined in an International Standard related to the subject matter in question, in our case for example ISO&nbsp;10646?

(I ask just out of curiosity, not that “unordered list” is such a term—unless one counts the mentioned “general document” example DTD as kind-of authoritative, that is—nor any of the proposed alternatives …)

* * *

Regarding the remark in the post you refer to above:

> However, if consistency with HTML5 terminology is deemed necessary, I’d reluctantly prefer […]

Allow me to add a **little rant** that in my not-so-humble opinion—from what I have seen and studied of “HTML&nbsp;5” so far—it rather seems that

&nbsp;&nbsp;&nbsp;&nbsp; **consistency with “HTML 5 terminology” is very much a thing to avoid**

because it looks like the designers of HTML5 got carried away by an intense desire to use “new, abstracter, terminology” just for the sake of it. For example, re-naming `<HR>` to _“thematic break”_ (did they **really** _eliminate_ the phrase “horizontal rule” from the HTML spec?): this is IMO incredibly silly, if not to say stupid, and there are a many other examples like this (just think of the whole “bogus comment” concept and terminology!).

As you can probably guess, I do **not** hold HTML5 in high regard, and I would certainly **not** want to use it as an example of —shudder!—consistent use of terminology!

[I have to admit though that replacing _Adobe Flash_ by whatever kind of _HTML5 streaming video element_ may in fact be worth the otherwise horribly misguided effort that HTML&nbsp;5 is in my opinion so far 😉 …]

* * *

> Anyway, the older an engineer, the more spec term changes he proposes.

This may well be because an “older engineer” tends to have seen more _needless_ changes in terminology for the same old concept (introduced out of cluelessness, or out of vanity, or because it’s fashionable), and thus tends to prefer using the same (“old”) terms for the same concepts—a principle which is an instance of [Occam’s razor](https://en.wikipedia.org/wiki/Occam%27s_razor), if you think about it. (Which explains part of my opinion about HTML5.)

* * *

We seem to agree on the “issue” of terminology though; I would be (pretty much equally) happy with either

- “ordered list” and “unordered list” (for said “historical” or “traditional” reasons), or

- “itemized list” instead of “unordered list” (for LaTeX and DocBook precedence), and maybe even “enumerated list” (for LaTeX precedence) instead of “ordered list”.

Both “bullet list” and “numbered list” are very much tied to a specific presentation style, and are thus not good choices

- either for naming the _CommonMark_ **input syntax** construct (ie the text of a list item that starts with “`-`” or “`+`” or “`*`”, or “`1.`” etc: this is _“just markup”_, and for example using “`-`” and “`*`” produces the **same** result anyway (as do for “ordered list items” the input markups “`1.`”, “`2.`”, or “`3)`”, IIRC);

- or for naming the **parsing result** : these lists are “typically” transformed into `<UL>` or `<OL>` elements (in HTML), or maybe into `<ItemizedList>` or `<OrderedList>` elements (in DocBook), and so forth: but in all these target document types, the **rendering style** of these lists is **solely** determined by some kind of applicable style sheet (or the whims of a _user agent_’s defaults, apart from some fuzzy recommendations for “default styles”). And there are [_lots_ of ways](http://www.w3.org/TR/CSS21/generate.html#propdef-list-style-type) to “number” the [items in a list](http://www.w3.org/TR/xsl/#d0e12377)!

So in the presentation of the result document,

- neither a [U+2022](http://www.fileformat.info/info/unicode/char/2022/index.htm) “bullet” character (in “unordered list” items)
- nor a number (in “ordered list” items)

needs to be present, which makes the names “bullet list” and “numbered list” rather obviously not very appropriate for “the output side”, and IMO hence inappropriate for the “input side” too.

* * *

To put my point of view bluntly into a simple slogan:

> _CommonMark_ is **not** a style-sheet or formatting language!

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [January 1, 2016, 4:30pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/17 "2016-01-01T16:30:18Z")

</div>

> [@tin-pot](#):
>
> Would you accept as “merit” of a term if that term itself is standardized, in our case for example in ISO 2382-23 Vocabulary – Part 32: Text processing, or if the term is defined in an International Standard related to the subject matter in question, in our case for example ISO 10646?

If you put a gun to my head, I’d pick “most widely used” over “standardized”. HTML 5 happens to be a standard (or [two](https://whatwg.org/html)), and yet

> [@tin-pot](#):
>
> consistency with “HTML 5 terminology” is very much a thing to avoid

The more familiar I become with the W3C standard, the more I tend agree with this statement.

> [@tin-pot](#):
>
> For example, re-naming \<HR\> to “thematic break” (did they really eliminate the phrase “horizontal rule” from the HTML spec?): this is IMO incredibly silly, if not to say stupid

Didn’t you get the [memo](http://talk.commonmark.org/t/horizontal-rule-or-thematic-break/912/9)? 😉

> [@tin-pot](#):
>
> “itemized list” instead of “unordered list”

Apparently, my command of English is insufficient for me to grasp this contraption. Doesn’t _itemized_ mean _consisting of items_, or simply a _list_?

> [@tin-pot](#):
>
> “enumerated list” […] instead of “ordered list”

I couldn’t imagine anyone having an issue with the term “ordered list”. Until now. 😄

> [@tin-pot](#):
>
> the text of a list item that starts with “-” or “+” or “_", or “1.” etc: this is “just markup”, and for example using “-” and "_” produces the same result anyway (as do for “ordered list items” the input markups “1.”, “2.”, or “3)”, IIRC

In CM the different unordered list markers determine whether an item belongs to an existing list (i.e.

```
- foo
+ bar

```

result in two distinct lists), whereas in ordered lists the first marker determines the start index.

As for presentation, the following is a perfectly standard HTML ordered list (the Arabic numerals are only provided for illustration):

甲、1  
乙、2  
丙、3  
丁、4  
戊、5  
己、6  
庚、7  
辛、8  
壬、9  
癸、10

---

<div class="post-metadata">

**Author:** ![tin-pot](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/tin-pot/32/810_2.png) [@tin-pot](https://talk.commonmark.org/u/tin-pot)\
**Post date:** [January 1, 2016, 6:24pm UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/18 "2016-01-01T18:24:25Z")

</div>

> [@Dmitry](#):
>
> If you put a gun to my head, I’d pick “most widely used” over “standardized”. HTML 5 happens to be a standard (or two),

It’s true that “most widely used” is often different from what is “standardized”. And because the whole _CommonMark_ effort is an exercise in _standardization_ (alas without any “officially” sanctioned status, of course), the question turns out to be whether it is a good idea

- to use the terminology of related standards (widely used or not), or

- rather try to “elevate” a widely-used term to a pseudo-standardized level, and use it in the _CommonMark_ specification _instead_ of existing terminology.

* * *

> [@Dmitry](#):
>
> The more familiar I become with the W3C standard, the more I tend agree with this statement.

I haven’t dug very deep into the HTML 5 specifications yet, but _what_ I’ve seen so far did not exactly fill me with awe nor made me [shiver with an–tici—pation](https://www.youtube.com/watch?v=wlwnbcxBuzI) …

Right now I still have the hope that I have severely misunderstood the whole HTML&nbsp;5 approach, and that I will one day see through and beyond my puzzlement the crystal-clear, well thought-out, forward-looking, both versatile and upwards-compatible specification that HTML&nbsp;5 is supposed to be.

But don’t hold your breath, I certainly don’t 😉

* * *

> Didn’t you get the memo?

I did, and I shook my head in disbelief.

* * *

> [@Dmitry](#):
>
> Doesn’t itemized mean consisting of items, or simply a list?

English is a foreign language for me too, but according to what I can deduce from context, **“itemizing”** here means something like _“visually marking [the beginning of] each item in a list”_, ie with some “item marker” like a “bullet” (this is consistent with the use of this term in LaTeX and in OASIS DocBook).

But if each item is marked _individually_, typically with an item number or letter in place of the “marker”, this would constitute an **“ordered list”** , or “enumerated” list, and would not be called an “itemized” list.

Finally, a “plain” list, where each “list item” is just a paragraph (even if indented) without any sort of “item marker”, would then be a **“simple list”** (consistent with the use of that term in the ISO&nbsp;8879 “general document” DTD, and in OASIS DocBook), and would not fall under the definition of “itemized list” I attempted here.

**By the way:** _Please note that it’s not **me** having an issue with “unordered” and “ordered” list, nor with “itemized” and “enumerated” list._ [_I’m having “an issue” with “bullet list” and “numbered list” instead_ 🙂]

* * *

> In CM the different unordered list markers determine whether an item belongs to an existing list

You’re right of course: the specification explicitly says so, and even has an [example](http://spec.commonmark.org/0.23/#example-251). I somehow forgot that case (but wouldn’t rely on it being commonly implemented anyway), and consider it more of a kludge (useful to enforce consistent input markup) than being a generally useful feature: how do you “split” a numbered list? Change between "`1. `", "`2. `", then "`1) `" and "`2) `"? Yuk!

A list-related feature which is _missing_ and _would_ be useful IMO would provide a way to _continue_ an “ordered” list, after some intervening stuff “interrupted” the list, like in this faked example:

> &nbsp;1. First Item in list.
> 
> Some paragraph.
> 
> &nbsp;2. Second item in list.

But that’s of course a rather different topic.

* * *

> [@Dmitry](#):
>
> As for presentation, the following is a perfectly standard HTML ordered list (the Arabic numerals are only provided for illustration):

I’m not sure what you mean by “standard HTML” here, but the HTML markup of your example—from what I can see in my browser—looks like this:

```
<p>甲、1<br>乙、2<br>丙、3<br>丁、4<br>戊、5<br>己、6<br>庚、7<br>辛、8<br>壬、9<br>癸、10</p>

```

Do you mean that a _user agent_ would be allowed to present an “ordered list” in the style of your example, that is in a manner where in the “marker box” of each item of the list is a _chinese celestial stem_1) instead of a _decimal_ or _roman_ or _alphabetic_ numeral?

Of course that’s “allowed” in “perfectly standard” HTML: to start with, for the simple reason that the HTML specification (or for that matter: the DocBook or any other2) _document type_ specification) does not and reasonably **can not** require a particular rendering style beyond rather general hints about the “meaning” of the various element types.

This is of course **in stark contrast** to specifications of _document formats_ or _formatting languages_ and so on, where for example a particular LaTeX style or the RTF specification as a whole **does** constrain the rendering of governed document instances pretty narrowly (I think the same does hold for the OpenOffice XML document format)—right up to _page description languages_ like PDF, SPDL, and PostScript or [Microsoft’s](https://msdn.microsoft.com/en-us/library/windows/hardware/dn614032%28v=vs.85%29.aspx) XPS / [ECMA](http://www.ecma-international.org/memento/TC46.htm) OpenXPS, which do **nothing else** but fixing and constraining the “presentation” or rendering of documents.  
\_\_\_\_\_\_

1. I freely admit that I had to use Google for that one. Hooray for Unicode! 😉
2. Except of course _document types_ which are explicitly designed to also convey presentation information (like the mentioned OpenOffice XML, or simpy HTML with embedded CSS `style` attributes) or are even dedicated to this purpose (like XPS and OpenXPS, which are _page description languages_).

* * *

But what has this to do with the difference between “ordered” lists (where each item is “marked” with an individual “marker”, taken from an _ordered_ set, be it decimal or roman numerals or celestial stems or [counting rods](http://unicode-table.com/en/blocks/counting-rod-numerals/) or [cuneiform numerals](http://unicode-table.com/en/blocks/cuneiform-numbers-and-punctuation/) or whatnot) and “unordered” (or “itemized”) lists (where each item is “marked” in the same way, say with a “bullet” character)?

---

<div class="post-metadata">

**Author:** ![Dmitry](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/dmitry/32/849_2.png) [@Dmitry](https://talk.commonmark.org/u/Dmitry)\
**Post date:** [January 2, 2016, 4:25am UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/19 "2016-01-02T04:25:02Z")

</div>

> [@tin-pot](#):
>
> I did, and I shook my head in disbelief.

I find _thematic break_ a poor choice on at least three levels:

1. _Theme_ is overloaded with an unrealed meaning in other technologies (WPF/Silverlight/XAML themes are analogous to CSS).
2. _Break_ is overloaded with a different meaning in both HTML and CommonMark. How does one produce a soft thematic break?
3. Too many words are used to denote a simple unambiguous concept.

If I were asked to propose an alternative term, my suggestion would be **boundary**.

Having said that, I do agree that _horizontal rule_ should have been replaced due to its presentational semantics. In fact, an abovementioned project of mine had used `***` (or `---` or `___`) for topic boundaries before CM’s sudden (for me) change in terminology.

I simply fail to see any good reason for rejecting _unordered lists_ while embracing _thematic breaks_.

---

<div class="post-metadata">

**Author:** ![Crissov](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/crissov/32/1455_2.png) [@Crissov](https://talk.commonmark.org/u/Crissov)\
**Post date:** [January 2, 2016, 9:55am UTC](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945/20 "2016-01-02T09:55:26Z")

</div>

> [@Dmitry](#):
>
> … unordered list items are only **rendered** with bullet markers (…), but only have `-`, `+` or `*` markers in the input (…). Why _bullet lists_ then?

Maybe it’s my non-native level of English, but I have no problem with calling `-`, `+` and `*` (ASCII) _bullets_ when used as line markers for list items.

[Next page](https://talk.commonmark.org/t/ordered-vs-unordered-list/1945.md?page=2)
