# Is a reverse conversion (HTML to Markdown) possible?

**URL:** <https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899>\
**Category:** Implementation\
**Created:** [November 11, 2014, 9:40pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899 "2014-11-11T21:40:06Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Uwe\_Keim](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/uwe_keim/32/2086_2.png) [@Uwe\_Keim](https://talk.commonmark.org/u/Uwe_Keim)\
**Post date:** [November 11, 2014, 9:40pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/1 "2014-11-11T21:40:06Z")

</div>

Currently, I’m investigating whether it would be possible to not convert Markdown to HTML but the other way around: HTML to Markdown.

I found [this JavaScript implementation](http://domchristie.github.io/to-markdown/) which doesn’t work very well.

**So my question is:**

Are you aware of a working HTML to Markdown conversion?

---

<div class="post-metadata">

**Author:** ![mb21](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/mb21/32/162_2.png) [@mb21](https://talk.commonmark.org/u/mb21)\
**Post date:** [November 11, 2014, 9:55pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/2 "2014-11-11T21:55:48Z")

</div>

[Pandoc](http://johnmacfarlane.net/pandoc/) works quite well. The command would be:

```
$ pandoc file.html -o file.md
```

---

<div class="post-metadata">

**Author:** ![Uwe\_Keim](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/uwe_keim/32/2086_2.png) [@Uwe\_Keim](https://talk.commonmark.org/u/Uwe_Keim)\
**Post date:** [November 11, 2014, 9:57pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/3 "2014-11-11T21:57:28Z")

</div>

Thanks! You probably mean [johnmacfarlane.net/pandoc](http://johnmacfarlane.net/pandoc) - right?

---

<div class="post-metadata">

**Author:** ![mb21](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/mb21/32/162_2.png) [@mb21](https://talk.commonmark.org/u/mb21)\
**Post date:** [November 11, 2014, 9:58pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/4 "2014-11-11T21:58:48Z")

</div>

Exactly. (link fixed)

---

<div class="post-metadata">

**Author:** ![Uwe\_Keim](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/uwe_keim/32/2086_2.png) [@Uwe\_Keim](https://talk.commonmark.org/u/Uwe_Keim)\
**Post date:** [November 11, 2014, 10:03pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/5 "2014-11-11T22:03:55Z")

</div>

I even found [this one](http://stackoverflow.com/a/6228889/107625) to make it work in C#.

Unfortunately starting a process seems way to slow for what I want to use it.

I’ll try to see whether I can migrate the related parts to C#…

---

<div class="post-metadata">

**Author:** ![mb21](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/mb21/32/162_2.png) [@mb21](https://talk.commonmark.org/u/mb21)\
**Post date:** [November 11, 2014, 10:09pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/6 "2014-11-11T22:09:50Z")

</div>

I don’t know your use case, but starting a process is really not as heavy as it used to be on old hardware (test it). If you work in Haskell, you can of course use Pandoc as a library as well, without starting a process.

---

<div class="post-metadata">

**Author:** ![Uwe\_Keim](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/uwe_keim/32/2086_2.png) [@Uwe\_Keim](https://talk.commonmark.org/u/Uwe_Keim)\
**Post date:** [November 11, 2014, 10:23pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/7 "2014-11-11T22:23:47Z")

</div>

Just saw that it is released under GPL, so it would not work in my case anyway (using it in a commercial software).

Too bad…

---

<div class="post-metadata">

**Author:** ![lu\_zero](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/lu_zero/32/557_2.png) [@lu\_zero](https://talk.commonmark.org/u/lu_zero)\
**Post date:** [November 12, 2014, 2:21pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/8 "2014-11-12T14:21:01Z")

</div>

kramdown works as well.

Porting either of those on C# might be a major task, wrapping them just to call them seems better.

---

<div class="post-metadata">

**Author:** ![Mike](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/mike/32/453_2.png) [@Mike](https://talk.commonmark.org/u/Mike)\
**Post date:** [November 12, 2014, 2:24pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/9 "2014-11-12T14:24:42Z")

</div>

You can use a GPL tool with a **process call** in a commercial application, just not with a library call.

---

<div class="post-metadata">

**Author:** ![Uwe\_Keim](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/uwe_keim/32/2086_2.png) [@Uwe\_Keim](https://talk.commonmark.org/u/Uwe_Keim)\
**Post date:** [November 12, 2014, 7:13pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/10 "2014-11-12T19:13:35Z")

</div>

Are you sure? What I understand is that I also are not allowed to bundle it within my installer.

---

<div class="post-metadata">

**Author:** ![Mike](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/mike/32/453_2.png) [@Mike](https://talk.commonmark.org/u/Mike)\
**Post date:** [November 12, 2014, 7:36pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/11 "2014-11-12T19:36:11Z")

</div>

I’m not sure about the installer, but calling a GPL process is ok as far as I know (I’m no lawyer). See here: [http://programmers.stackexchange.com/questions/50118/avoid-gpl-violation-by-moving-library-out-of-process](http://programmers.stackexchange.com/questions/50118/avoid-gpl-violation-by-moving-library-out-of-process)

---

<div class="post-metadata">

**Author:** ![philsturgeon](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/philsturgeon/32/533_2.png) [@philsturgeon](https://talk.commonmark.org/u/philsturgeon)\
**Post date:** [December 1, 2014, 12:33pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/12 "2014-12-01T12:33:32Z")

</div>

I’ve also recently had success with [Reverse Markdown](https://github.com/xijo/reverse_markdown) for Ruby.

It attacked all of my 2008-2011 TinyMCE-based WYSIWYG content, which was garbled and horrendous, and output some really basic Markdown. IT was about 99% spot on when parsed by Jekyll + kramdown.

---

<div class="post-metadata">

**Author:** ![pjt33](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/pjt33/32/534_2.png) [@pjt33](https://talk.commonmark.org/u/pjt33)\
**Post date:** [December 2, 2014, 10:38am UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/13 "2014-12-02T10:38:24Z")

</div>

I use [ittyeditor](http://code.google.com/p/ittyeditor/)’s JavaScript implementation. It does both MD (one flavour thereof, of course) to HTML and vice versa.

---

<div class="post-metadata">

**Author:** ![baynezy](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/baynezy/32/835_2.png) [@baynezy](https://talk.commonmark.org/u/baynezy)\
**Post date:** [November 14, 2015, 7:56pm UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/14 "2015-11-14T19:56:58Z")

</div>

I maintain a library written in C# to achieve this.

> **[baynezy/Html2Markdown](https://github.com/baynezy/Html2Markdown)**
>
> A library for converting HTML to markdown syntax in C# - baynezy/Html2Markdown

It is actively maintained and I happily accept contributions.

---

<div class="post-metadata">

**Author:** ![jsphpl](https://cdn.commonmark.org/user_avatar/talk.commonmark.org/jsphpl/32/998_2.png) [@jsphpl](https://talk.commonmark.org/u/jsphpl)\
**Post date:** [July 15, 2016, 7:35am UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/15 "2016-07-15T07:35:06Z")

</div>

There is this css file called [Markdown.css](http://mrcoles.com/demo/markdown-css/) that you can apply to your html to render it to markdown.

---

<div class="post-metadata">

**Author:** ![Charlie](https://cdn.commonmark.org/letter_avatar_proxy/v2/letter/c/d6d6ee/32.png) [@Charlie](https://talk.commonmark.org/u/Charlie)\
**Post date:** [March 16, 2017, 10:48am UTC](https://talk.commonmark.org/t/is-a-reverse-conversion-html-to-markdown-possible/899/16 "2017-03-16T10:48:24Z")

</div>

[Automate That Shit](https://automatethatshit.com/lab/html-to-markdown) works pretty well in my experience
