Plain text is still one of the best technologies we have

(deadparrotbbs.com)

204 points | by speckx 1 day ago

30 comments

  • devy 1 day ago
    Graydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.

    [1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)

    [2] https://hn.algolia.com/?q=always+bet+on+text

    [3] https://news.ycombinator.com/item?id=26164001

    [4] https://news.ycombinator.com/item?id=8451271

    [5] https://news.ycombinator.com/item?id=10284202

    [6] https://news.ycombinator.com/item?id=12815829

  • boomlinde 1 day ago
    The article briefly addresses the problem, but it's pretty fun how different "plain text" looks throughout history and in different domains.

    For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).

    There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.

    Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.

    • satiated_grue 10 hours ago
      I'm surprised by the existence of \n\r line terminations!

      \r\n was how lines were terminated on the wire for many early devices (for example, the Teletype ASR-33 seen in the famous photograph of Thompson and Ritchie at Bell Labs).

      If the carriage is very far to the right, and an LF is sent followed by the CR, then the carriage will still be returning when the next character arrives, which will print at some random location along the return path.

      CR followed by the mechanically much faster LF operation gives enough time for the carriage to be fully homed and in proper position for the printable character to be struck.

    • adiabatichottub 1 day ago
      Text is wonderful, but note that the one keyword not found in this text is 'parse'. How much extra work is generated writing parers for text that would have been so much easier to deal with if it had just been generated in a structured binary format? You still get the pleasure of designing a new wheel with every format.
      • cranky908canuck 22 hours ago
        It's not about writing a parser, it's about figuring out what the content is supposed to be in the first place. With plain text you can see what's a number, what's a string, what's an expression, etc... I'd rather try to figure out how to parse "2 7 + 3 -" than see a file with

        0x0 0x0 0x0 0x0 0x2 0x0 0x0 0x0 0x0 0x7 0x1 0x1 0x0 0x0 0x0 0x0 0x3 0x1 0x2

        and try to figure out what that is all about.

        • breakwaterlabs 7 hours ago
          If it were easy to see what's a number and whats a string, type confusions a la YAML's fun wouldn't exist: Is "no" a string, a country, a boolean? Is 0 a boolean, a number, or char[48]?

          This is specifically an area that "text" is quite bad and is the source of much woe. Your provided example would be parsed by most parsers today as a string, since it mixes spaces (non-numeric) and surrounds it with quotes.

          In fact, given well-defined binary, it is rather simpler to turn it into "text" than the reverse. It seems like you've forgotten that "0x32 0x20 0x37 0x20 0x2A..." is literally what is in that file you would view. Plain text is, after all, just one example of a binary format.

          • cranky908canuck 3 hours ago
            This misunderstands what I was trying to get across. The string is the file content, not a string in a file (so: a short snippet of source code), in text format; what followed was (an attempt at) a plausible binary tokenization.

            I agree that the string "no" is ambiguous (though infamous as a somewhat pathological case), but at least it can be seen in context without having to spelunk a binary definition.

            is_bright_and_shiny: "no"

            vs

            curt_negative_response: "no"

            vs

            country_code: "no"

        • eviks 11 hours ago
          Unless it's a different encoding from 50 years ago so you can't figure out how to parse 27+3 And you can easily make the opposite example: you'd rather figure out that .ext means format X that you already have an app for, where date is unambiguous courtesy of the format designers rather that make some common mistake trying to figure out where month is in 5/6/79
      • zelphirkalt 1 day ago
        It does not have to be a binary format. Just a structured format, for which we can reuse an existing parser.
    • gwbas1c 1 day ago
      This morning I had co-pilot generate a LINQPad script that fixes a lot of that.

      Granted, I knew that all files are UTF-8, so I didn't need to have it "guess" what the encoding was without the BOM.

    • Liquid_Fire 1 day ago
      Sorry, can you clarify about Microsoft having adopted \n? This is the first time I hear of it, and I can't find anything online about it.
      • eska 8 hours ago
        It’s not a consistent change, but for example programs like notepad couldn’t handle \n, but do accept it now. Visual Studio expects certain file types to use \r\n, but accepts others like .gitignore using \n.

        Same thing for \ and / in path separators: Microsoft never made a consistent push for it, and it depends on the application.

        • Liquid_Fire 7 hours ago
          I'm not sure I would call that "adoption", so much as "compatibility with", since the default line ending produced by Windows software is still almost always \r\n for files created from scratch.
          • boomlinde 6 hours ago
            FWIW I think it's a fair point. My view of this is as an outsider only seeing what comes in for review from Windows-using colleagues so maybe I overstated the level of adoption.
            • Liquid_Fire 6 hours ago
              I believe git has a setting to auto-convert to CRLF on checkout and back to LF on commit, so that could also be partially to blame.
  • drhagen 1 day ago
    This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].

    [0] Windows newlines not withstanding.

    • bruce511 1 day ago
      Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better.

      As it is programs have to guess the following;

      A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?

      B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?

      C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?

      D) what human-language is it in?

      E) CSV? Don't get me started...

      Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.

      Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.

      But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...

      • hnlmorg 1 day ago
        ASCII is just a raw binary format. It isn't just for text.

        The line endings problem really isn't a problem. Pretty much every text editor out there can handle different line endings.

        I don’t think number nor date formats are relevant here. For example you could have that same problem entering text into MS Word. That’s really more of an issue if you want to use text as a database rather than a document format, which isn’t really something that even plain text advocates would generally recommend.

        As for human language detection, that’s a much easier problem to solve than decoding a proprietary binary blob.

        CSV definitely has its warts. But it’s not like that’s the only plain text option for serialising data. (JSON, jsonlines, YAML, XML, etc). Or you could use the actual ASCII codes reserved for records, if you really wanted something that didn’t require quoting and escaping in plain text. It’s actually a pity nobody does this.

      • ghusto 3 hours ago
        > But sure, ASCII files in English with US date formats

        LoL, I love the irony of going with US date formats as the one of the obvious defaults to that variable.

      • ninalanyon 14 hours ago
        > C) number formats? 100.000 or 100,000?

        Most national standards now call for a space (preferably a half width space, Unicode: U+2009) as the thousands separator and then either comma or full stop as the decimal separator.

        The half width space was standardised in SI in 1948. See, i.a., https://snl.no/tusenskille

        • eviks 11 hours ago
          Why is it not 2008 punctuation space to match the width of alternative thousands ,. separator?
      • dwaltrip 1 day ago
        The entire point of plain text is that it doesn’t have number formats, date formats, and so on.

        This is the eternal tradeoff of specificity vs. generality.

        • Izkata 10 hours ago
          More generally, it's not rich text. The two terms are a pair, where rich text defines additional structure around the characters in order to specify things like that, and plain text has none of it.
      • conartist6 1 day ago
        I've made a language called CSTML to do some of this: https://docs.bablr.org/guides/cstml

        Basically it's a text-based language for embedding semantic metadata into some kind of underlying text stream. Is this interesting to you?

      • jimbokun 1 day ago
        [flagged]
    • layer8 1 day ago
      > Windows newlines not withstanding

      “Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard.

      I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports.

      Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign.

      • hnlmorg 1 day ago
        UNIX (and Linux) is even more annoying because pseudo TTYs will require CRLF when in raw mode but requires only LF when in line mode.
    • tyromaniac 1 day ago
      And file endings..
    • Analemma_ 1 day ago
      If you write files in ASCII you’re already writing in UTF-8.
      • sillysaurusx 1 day ago
        That’s true, though if you’re writing UTF-8 you have to handle the corner cases when reading UTF-8, of which there are many.

        Fortunately they’re easy to test for and most languages have standard libraries that make this painless.

      • bell-cot 1 day ago
        ASCII has, in principal, infinitely many superset.

        And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII

  • kerblang 1 day ago
    What, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcastic
    • adiabatichottub 1 day ago
      And then you add RFC2045 headers so you know what the encoding is...
      • kerblang 1 day ago
        Text headers, mind you! Ah, text text text

        Don't ask me what my point is

  • spcebar 1 day ago
    Fun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.

    With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.

    • dunham 1 day ago
      John August has also talked about using plain text for movie/television screenplays, specifying a format based on markdown and hollywood standard: https://fountain.io/

      Sharing aside, I imagine a plain text format would also be helpful for version control.

      • spcebar 1 day ago
        Haha, yes, as I was finishing writing the parser, I stumbled upon Fountain and realized my genius idea for a markdown based playwriting format was not so original. Ended up adding some Fountain syntax to my parser and ultimately I built a very functional but not feature complete implementation of Fountain by accident.

        I built sharing and (sort of) versioning into the web app. Everything is saved in localStorage and on a MySQL server. You can Save or Save As, and every time you Save As it inserts a new database row and gives you a new, sharable URL.

        I wanted an experience where you didn't have to login to use it, but ultimately, having a way to keep track of all those URLs would actually be a pretty decent versioning system--add diff checking between script versions and you end up with something pretty handy.

    • someonebaggy 1 day ago
      [flagged]
      • spcebar 1 day ago
        XML is a good candidate, but actually overkill for the purposes I was looking to achieve. Basically, all I needed was character names are all caps, scene headings have hashes, and stage directions are free text. Put a handful of other niceties like title pages, casts, and page breaks, but never anything so complex it needed additional markup.

        # INT. JOE'S BEDROOM, DAY.

        JOE

        This is my dialogue, yeehaw.

        JOE exits, and we know JOE is exiting because this is a stage direction, and we know this is a stage direction because it is free text and doesn't satisfy any other formatting requirement.

  • shmolyneaux 1 day ago
    I love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.

    It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.

    I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).

  • andsoitis 1 day ago
    The more easily you can access meaning from symbols without intermediate tools, the more durable and resilient.

    Symbols on paper is best.

    Plain text in a file that can be opened by any computer comes close, but needs a tool.

    Fancy file formats that need not only hardware but also special software are the worst.

    However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.

  • anvuong 1 day ago
    ASCII was amazingly efficient for conveying English text. Then we needed to encode multi languages and emojis, the resulting Unicode is just a mess.
    • Perseids 1 day ago
      Human languages are a mess. Some mix left-to-right and right-to-left. Some have variable amount of diacritics. Some have non-well-defined character sets where just trying to map it to a manageable complexity still loses cultural heritage. When you are fondly remembering the good old times of ASCII, you just wish back to be able to ignore anyone not speaking English. (Which is in your right to do, if you so desire.)
    • weinzierl 1 day ago
      Unicode was mostly OK when it wanted to encode multi languages. It started to get a mess when thought it needed to do more.
      • Dylan16807 17 hours ago
        The main complicating factors in Unicode that could have been avoided are around CJK unification and the overflow past 16 bits. Both are very much problems of encoding normal languages, not doing more.
      • frollogaston 1 day ago
        Yeah the multi-unicode-char grapheme clusters are basically only for emojis once they started adding stuff like every ethnicity/gender combo of 4-person family. And that one historical Korean script.
        • Dylan16807 17 hours ago
          Those clusters are reusing technology that was needed for actual languages. Combining characters in general are used all over.
      • OkayPhysicist 1 day ago
        What do you think doesn't belong in Unicode? It turns out that expressing all the different ways humans have used text in history has some necessary complexity, but assigning each character and modifier a number seems like a perfectly reasonable approach.
        • weinzierl 1 day ago
          "have used text in history"

          That's how Unicode started but nowadays Unicode expresses things that are only used because Unicode itself introduced them.

          And this while many important things in the "have used text in history" category have been unfinished or not tackled at all.

          • OkayPhysicist 1 day ago
            If you're bellyaching about emojis, you should be treating them as a litmus test: If your application can't handle emojis, it's because you're doing something you're not supposed to (specifically, making false assumption about the nature of text), and thus you will run into situations with text that break your shit, too.

            The way I see it, emojis are practically single-handedly responsible for near-universal adoption of Unicode in the Anglosphere. Otherwise, it would be all too easy for lazy developers to just use ASCII and then shove their fingers in their ears about any other writing systems.

            • weinzierl 12 hours ago
              I'm rather indifferent to emojis. They don't bother me and they are not my point.

              My point is that Unicode's own mission statement says: "The Unicode Consortium enables people around the world to use computers in any language."

              I think it strayed from that. And I'd argue the mission was only ever fulfilled when seen through an anglospheric lens. You don't even have to summon Han Unification, the elephant in the room. Documents from Western Europe from a couple of decades ago can't be represented properly.

              Example: you cannot properly represent a German telephone book today. Take "Müller" and the French surname "Haüy". Both contain the same Unicode character, but in "Müller" it is an umlaut and sorts like "ue" (Mueller), while in "Haüy" it is a trema and sorts like a plain u (Hauy). For all intents and purposes these are different letters. That they look similar (not the same, in proper typography) is a mere coincidence. Yet Unicode encodes them identically, so the information needed to sort correctly is simply not in the text anymore.

              Han Unification is the other sore point, and a much bigger one.

              I'd like Unicode to focus on its core mission, encoding the languages that exist, before it starts shaping new ones. For that I'd propose a grace period: a new character should have to be in significant active use for 10 to 20 years (not merely implemented on a popular device) before it is accepted into the standard.

    • eviks 13 hours ago
      How it efficient when it wastes a quarter of its budget on control nonsense?
  • real_faxenoff 10 hours ago
    Text is a great format, no doubt about it. But it's not structured. Maybe we should surround certain parts of the text with special characters to create some structure? But wait a minute...
  • wpollock 1 day ago
    I too love plain text data, but absent a universally accepted standard, meta data is needed to avoid guessing a document's charset, encoding, and line ending convention.

    Associating such meta data with text files wasn't simple in the beginning, but that changed in the 1990s. Nearly all filesystems since then allow meta data: POSIX compliant filesystems support extended attributes, mostly used for security meta data. Even ancient HFS (Apple Mac) supports meta data. NTFS supports it with alternative file streams.

    While possible, no system I know of attempts to use any such meta data, probably due to a lack of standards. A pity.

    • Dylan16807 17 hours ago
      Every day that goes by, it's more likely that it's UTF-8 and "\r and/or \n (greedy)(in that order)" works to parse the line endings. So I'm not very bothered.
  • frollogaston 1 day ago
    I greatly prefer plaintext in most cases where Markdown ends up being used. WYSIWYG is more useful than formatting in these cases. Most of the formatting makes it harder to read anyway. Some say you can ignore MD and treat it like plaintext... not so when you use a newline.

    Oh and now you could have LLMs write the Markdown, but chances are no human reads that, or even if you do, the formatting is going to be insane. I have to keep telling Claude to give me a .txt instead of .md when asking for a context dump. Maybe .md is popular for agent readmes because they optimize around its headings.

  • njarboe 1 day ago
    I work on scientific data repositories [1] and pushed for plain text files as our archiving file format two decades ago as the best long term format. We implemented it on our data repositories. Very human readable also. The datasets are small and table based, so this works well. We latter started archiving 2-D image data and that creates very large files. We are still looking for an elegant solutions for those datasets.

    [1] https://earthref.org/FIESTA/

  • CrimsonCape 1 day ago
    Hijacking the plain text discussion to ask what the best tools are to convert plain text to lexed/parsed output?

    I'm assuming any tool in this regard would expect the user to write an EBNF grammar.

    I found ANTLR to be nice but it's way too convoluted to use as a tool with Java dependencies and seems to be stagnating. And tree sitter just is too convoluted and requires the added C/C++ overhead to understand how to use it.

  • olexsmir 1 day ago
    I feel obligated to mention ledger[0] and hledger[1], those are plain text accounting software, well, as you can guess from their names, they allow you to do personal accounting in plain text.

    0: https://ledger-cli.org

    1: https://hledger.org

  • rurban 17 hours ago
    Still, new users tend to prefer videos over text. Even newspapers dont write anymore, they create video content.

    Why say anything in 2000 bytes, which you can read in 10 seconds, when you can express it in 4MB, and watch it in 7 minutes?

  • Grimeton 1 day ago
    ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.

    The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.

    Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.

    • Perseids 1 day ago
      You seemed to be deeply confused about encodings and character sets.

      > The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.

      That is only true for the mapping of Unicode character (code point to be exact) to UTF-8, which is an encoding of Unicode characters.

      > Unicode are multibyte characters with variable byte length and endianess at play.

      That is only true of the UTF-16 encoding. UTF-8 does not have endianess, UTF-32 does not have variable length per code point.

      None of that is true for Unicode, because it is abstracted away from any byte representation. Furthermore, getting back to the start:

      > ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.

      ASCII also depends on how you try to decode your text. If you interpret two bytes as one character, ASCII will never be correct. There is nothing magical about one byte mapping to one character. I'd even argue that the only reason ASCII support is so universal internationally is because of UTF-8. Otherwise many countries would default to encodings where ASCII is not a subset (as they did before UTF-8 became common). So IMHO, UTF-8 and Unicode are the only widely understood encoding and character set.

      • rerdavies 23 hours ago
        And you seem to be a bit confused about code points and graphemes clusters.

        UTF-32 grapheme clusters can easily be six or seven code-points long. Examples: modifier characters, flag/regional-indicator sequences, (and to a lesser extent, combining characters).

        Why does it matter? Because you can't insert a linebreak in the middle of a grapheme cluster.

        So even after converting all of your UTF-8 text to UTF-32, you STILL have to deal with multi-character code-point sequences.

        Madness!

        • Dylan16807 17 hours ago
          Just accept that text always has a variable number of bytes. It's not madness.
          • rerdavies 5 hours ago
            I'm fine with characters having a variable number of bytes. The problem comes when I have to interpret individual characters and/or codepoints and/or glyphs in a UTF string when trying to correctly lay out text on a Linux Terminal.

            I just find it disconcerting that a single character in UTF-8 can be of infinite length (an infinite chain of modifiers, followed by an emoji; or an infinite stack of composing accents). And that the Unicode consortium has chosen to break the one-code-point-per-glyph convention.

            An example of a truly weird outlier: the "family" emoji, which is composed as ManEmoji+ZeroWidthJoiner+WomanEmoji+ZeroWidthJoiner+ChildEmoji, which produces a single glyph.

            And then how to deal with modifiers? e.g.

               MonochromeEmojiModifier+
               AsianSkinToneModifier+
               ManEmoji+
               ZeroWidthJoiner+
               WomanEmoji+
               ZeroWidthJoiner+
               ChildEmoji. 
            
            (70 UTF-8 characters, 14 UTF-16 characters, 7 UTF-32 characters! How many glyphs? From one to three depending on what typeface you're using!). Unlike composing accents, which have a combined code-point, the Family emoji does not have a combined code-point, so it is typically implemented using combining in glyphs in the output TrueType/OpenType font, since TT/OT fonts can combine glyphs into private codepoints visible only inside the font.

            or

               AsianSkinToneModifier+
               ManEmoji+
               ZeroWidthJoiner+
               AsianSkinToneModifier+
               WomanEmoji+
               ZeroWidthJoiner+
               AsianSkinToneModifier+
               ChildEmoji. 
            
            (Is that even legal?)

            And then an appendix of special rules that apply mostly to obscure scripts/locales that have to be dealt with almost on a case-by-case basis (fortunately, of those, German and Turkish are the only locales with odd special cases that Linux supports).

            ALL of which are encoded in about 500,000 lines of code, documentation and data files by the Unicode consortium, that are updated annually. Unicode 18.0 includes 18 new emoji (including the infamous Pickle emoji, and skin-tone modifiers for left- and right-thumb), and three new scripts including proto-cuneiform. Yay! One assumes that Linux will implement the Pickle emoji on an urgent basis, and that proto-cuneiform will never be supported as a system locale.

            The use-case that almost broke me: line-wrapping and displaying and editing arbitrary user-entered UTF-8 paragraph text on a Linux terminal -- both graphical (almost full Unicode support but very out-of-date emoji), AND true text mode (up to 512 locale-dependent glyphs, each of which may or may not be double-width) that has to be supported to allow use on machines without desktops. The IBM Unicode libraries get you part of the way there, but far from all of the way there; and the few system APIs that Linux do provide are buggy as heck.

            • Dylan16807 4 hours ago
              > I just find it disconcerting that a single character in UTF-8 can be of infinite length (an infinite chain of modifiers, followed by an emoji; or an infinite stack of composing accents).

              One of the Unicode guidelines suggests it's not very important to handle sequences that are more than 31 code points long.

              > And that the Unicode consortium has chosen to break the one-code-point-per-glyph convention.

              The what convention? Combining accent marks were in there from the beginning, and those came from previous text formats.

              If you're dealing with a terminal you should be tracking one grapheme clusters per cell (or per two cells if you allow wide characters) and taking notice and control over when they merge. Store them in a way that putting line breaks in between the grapheme clusters is easy.

              > And there are the truly weird outliers, like "family" emoji, which is composed as ManEmoji+ZeroWidthJoiner+WomanEmoji+ZeroWidthJoiner+ChildEmoji, which produces a single glyph. And then how to deal with modifiers?

              That's the thing, it's not actually any weirder than the hangul jamo encoding for Korean text. And the flags also work like that but without needing ZWJ. They didn't add this complicated stuff for the sake of emojis, and like another comment said emoji acts as a shiny lure to get everyone to implement the spec properly.

              > 70 UTF-8 characters

              If you want to conflate characters and code points I'm not going to yelp too loud but please don't conflate characters and UTF-8 bytes, that's not how it works.

              Also I want to know how you calculated that number because it's wrong. It's 28 bytes in all three of those encodings. Even a broken conversion from UTF-16 into UTF-8 that bloats the number of bytes is 42.

              > And then an appendix of special rules that apply mostly to obscure scripts/locales that have to be dealt with almost on a case-by-case basis (fortunately, of those, German and Turkish are the only locales with odd special cases that Linux supports).

              That sucks but humans made text hard long before computers were around.

              > Unicode 18.0 includes 18 new emoji (including the infamous Pickle emoji, and skin-tone modifiers for left- and right-thumb), and three new scripts including proto-cuneiform. Yay! One assumes that Linux will implement the Pickle emoji on an urgent basis, and that proto-cuneiform will never be supported as a system locale.

              That seems fine? Adding another emoji is really easy and adding a locale is hard and will only be done if there's users.

              > The use-case that almost broke me: line-wrapping and displaying and editing arbitrary user-entered UTF-8 paragraph text on a Linux terminal -- both graphical (almost full Unicode support but very out-of-date emoji), AND true text mode (up to 512 locale-dependent glyphs, each of which may or may not be double-width) that has to be supported to allow use on machines without desktops. The IBM Unicode libraries get you part of the way there, but far from all of the way there; and the few system APIs that Linux do provide are buggy as heck.

              It does suck that these libraries and APIs are a mess.

  • ollien 1 day ago
    Tangential to the actual point of the post, but the talk "Plain text? Really?" by Dylan Beattie[1] is one of my favorite talks. It does a great job capturing the problems with something that "seems" so simple. I think the author of the blog is well aware of these, though :)

    [1] https://www.youtube.com/watch?v=_mZBa3sqTrI

    • s1mon 1 day ago
      Many things I knew, and a bunch I didn't. Thanks for that.

      I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.

      • radiator 1 day ago
        I notice that you have answered in considerably less time than the duration of the video.
  • pratikdeoghare 1 day ago
    I like plain text. I tolerate markdown. Markdown was designed as a shorthand for html. It gets stretched to do other stuff.

    I created Brashtag [1]. It is simpler than markdown.

    [1] https://github.com/PratikDeoghare/brashtag

    • rerdavies 23 hours ago
      This isn't a competitor to markdown. It is a competitor to JSON. Not sure what brashtag does that JSON does not. I'm not at all convinced that brashtag is simpler than JSON.
      • pratikdeoghare 19 hours ago
        > It is a competitor to JSON. Not sure what brashtag does that JSON does not. I'm not at all convinced that brashtag is simpler than JSON.

        As you have noticed it is poor substitute for json. It would be cumbersome to write config file in brashtag.

        > This isn't a competitor to markdown.

        You can write blog post in markdown. You can write it just as elegantly in brashtag. Only brashtag does not convert the doc to html for you. It just gives you tree. You write a program that processes the tree and produces html.

        Example,

        Markdown: # Lorem ipsum dolor sit amet

        Consectetur _adipiscing_ elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam.

        ---

        Brashtag:

        #h1{Lorem ipsum dolor sit amet}

        Consectetur #u{adipiscing} elit, sed do eiusmod tempor incididunt #b{ut labore et dolore} magna aliqua. Ut enim ad minim veniam.

        OR you could write

        #really big title{Lorem ipsum dolor sit amet}

        Consectetur #super duper underline{adipiscing} elit, sed do eiusmod tempor incididunt #emphsize{ut labore et dolore} magna aliqua. Ut enim ad minim veniam.

        • rerdavies 4 hours ago
          Not sure you're aware of this... most markdown engines allow you to also use a very constrained set of HTML tags inline.

          e.g.:

          ``` ## <p style="margin-left=4em;font-family: 'Courier,mono"> Life is death's way of giving what it kills time to grow in. &mdash;w.s. burroughs. </p> ```

          Necessarily, inline HTML tags, attributes and CSS is carefully sanitized by engines that support this. And, unfortunately the extent of sanitization, and actual supported tags vary from engine to engine. But it's there, and it is widely supported, and relatively predictable for basic HTML formatting. And mostly manages to excise ever-evolving ways to inject javascript into standard CSS.

          HNs markdown engine doesn't support inline HTML; but Github's markdown engine does.

    • CrimsonCape 1 day ago
      Can you think a little more about your decision to strongly type code blocks. Some domains have no need for code escaping but do need blob text escaping (which you already have defined elsewhere.) Technically you could treat code via 'this is a bag named code with a blob, therefore it is code."
      • pratikdeoghare 1 day ago
        I wanted to keep the famaliar things. We write code blocks with backticks in markdown and slack and everywhere so I used those. In fact I used $ signs earliar then switched to backticks.

        > 'this is a bag named code with a blob, therefore it is code."

        This would make #code{} a special bag. Right now no bag is special. Also it would be hard to find the closing } if somebody wrote unbalanced paren code in the bag like #code{ func main() { }.

        • CrimsonCape 1 day ago
          Fair point on the unbalanced braces. However,

          You could assume that { and } as a structural delimiter only applies to node types that can nest. Since blobs do not nest, then a tag followed by a backtick (or multiple backticks) could be a valid delimiter in lieu of braces. In this case, tags would extend to blobs.

          In my opinion, pairing minimal node types with differing delimiters helps legibility and is not overly mentally taxing to learn when node types are few.

          • pratikdeoghare 19 hours ago
            If I understand correctly, you mean something like this.

            #foo`fmt.Println("Hello world!")`

            instead of current

            #foo{`fmt.Println("Hello world!")`} .

            This would mean code can have tags too just like bags. Right now only bags can have tags and children. If you want to tag something, you have to put it in a bag. This works for tagging blogs, code, other bags.

            I think the extra braces are worth suffering for uniformity like the trailing commas.

    • jrm4 23 hours ago
      You're doing a different thing. The killer feature of Markdown isn't code and (my beloved) nerd stuff.

      It's that it can be a Microsoft Word, etc., killer.

      • pratikdeoghare 19 hours ago
        I think brashtag is more effective than markdown at killing those things.

        Because it stops after making you a tree. Markdown takes your document all the way to html. Making it do anything else requires much work.

        file.md -> markdown -> file.html

        file.bt -> brashtag -> tree -> to-html -> file.html

        file.bt -> brashtag -> tree -> todo-list-maker -> todo list

        ...

        file.bt -> brashtag -> tree -> ... -> ...

        Brashtag is not about code at all. Most text is already a valid brashtag. It is optimized for human writing. For example, this comment is a valid brashtag document if you surround it with #{} which can be done by a program easily.

        • jrm4 6 hours ago
          Respectfully, you are completely wrong.

          I can put Markdown in front of anyone who has ever used a computer for anything and they generally get it.

          You're doing things that look like code.

    • gandreani 1 day ago
      Why did you create it? What were you looking to solve address?
      • pratikdeoghare 1 day ago
        I had multiple frustrating experiences. I kept thinking I could just write things in some format in text and have some programs process it to do different things. So I created it.

        See how easy it is to write programs to process brashtag [1].

        I can chat with llm, interact with jupyter kernel and do literate programming from any text editor all in the same document [3].

        It took me like 10 minutes to build this mermaid like thing [2].

        List of frustrations:

        - One day Anki crashed. --- You can just write #card{....} in a text file. Some program would read it and show flashcards in browser.

        - Another day JabRef crashed. --- You can just write #bib{....} put link to a paper and its citation info there and have a program download the paper and copy the citation to bib file.

        - Markdown parser processed mathjax wrong. Why keep trying different markdown parsers? Just write #h1{}, #b{} #i{} and just convert them to html tags.

        - Literate programming tool I was using failed at some corner case.

        - My notes were scattered across files and tools. I couldn't find anything. I tried to have a text file where sections were separated by four dashes. But that didn't allow nesting. Also, the parser I wrote for that hit corner cases. Why not just put #note{} #JIRA-334{} etc. in text file.

        [1] https://github.com/PratikDeoghare/brashtag#some-programs [2] https://github.com/PratikDeoghare/brashtag/tree/master/cmd/m... [3] https://www.youtube.com/watch?v=IMXgIE0Vljg

    • someonebaggy 1 day ago
      [flagged]
      • pratikdeoghare 1 day ago
        What if you wanted say footnotes? html does not have it.

        Imagine you could just say #footnote{This will become footnote} and have a little program that would produce html that will show it at the bottom of the page.

  • vatsachak 1 day ago
    Probably because plain text can encode any form of distilled data. Technically our DNA can be plain text lol
  • mythrilkey30 1 day ago
    I dropped the first 3 paragraphs into pangram and it is 100% AI generated. Not reading and moving on.
    • jrm4 23 hours ago
      I genuinely believe that doing this is a terrible idea, as long as the underlying idea is good and the writing is decent. Good ideas are good ideas, and the more they are spread the better.
  • dimiprasakis 1 day ago
    Prediction: RSS is coming back stronger than ever
    • slowin 1 day ago
      RSS was removed from most sites (unfortunately) due to business, not technical reasons. Just like Web 2.0 era "mashup" friendly APIs, business started locking down their data despite it being an extremely user-hostile move.
    • DonHopkins 1 day ago
      .plan is really simpler syndication.
  • jaekwon 1 day ago
    that's why gno.land contracts render to markdown as the standard. you can browse the world through your terminal.

    imagine the world wide web but markdown (and decentralized). gno.land is that.

  • m463 1 day ago
    I think text is getting a big boost from the age of ai.
  • leftnode 1 day ago
    Because no single person or entity can own it.
    • rerdavies 22 hours ago
      ANSI: the American National Standards Institute, a private non-profit group that oversees the creation of voluntary safety and product guidelines in the United States. Including ASCII. Pretty sure they own plain text, in the sense that OP uses it.

      And then ISO, which owns a bunch of extended ASCII encodings that provide essential support for "plain text" in languages other than English (and an ugly historical hostility toward non-European languages).

      And then Unicode Consortium, which owns Unicode, its ever-expanding list of complexities and exceptions and an ugly history of hostility toward Asian languages.

  • ummonk 1 day ago
    > The Unicode Standard defines plain text essentially as a sequence of character codes, without the additional formatting information associated with rich text. Fonts, colors, layout, and similar presentation details belong somewhere else.

    Nope. Skintones are part of Unicode. It also has 33 control characters from ASCII including one that rings a bell... There are also numerous characters added by Unicode that are literally called "layout controls".

    If you want a format that provides purely semantic information, then "plaintext" doesn't fit the bill.

  • Koshkin 1 day ago
    "A picture is worth a thousand words"
    • smalltorch 1 day ago
      There was a challenge awhile back to describe a picture in 1000 words can't remember where I saw it. The winner just made a encoder/decoder to turn jpeg into words. The image was somewhat compressed but contained more visual data then you could accurately represent in the same way a raw description could. I thought it was so cool.
    • mohamedkoubaa 1 day ago
      Sometimes I'd rather have ten unambiguous words.
  • mythrilkey30 1 day ago
    I pasted the first 3 paragraphs of this into pangram and it's 100% AI generated. Not reading, moving on. We really shouldn't be promoting slop on hacker news
  • charcircuit 1 day ago
    >There are not many computer file formats I would trust to still be readable fifty years from now, but plain text is one of them.

    Just the other day I had opus read a 10 year old proprietary file format. The idea that we will lose the ability to use file formats is not consistent with reality.

  • FLeXMurphy 1 day ago
    @dang @tomhow

    My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.

    Anyway, flagged.

    • bityard 1 day ago
      I was about to argue with you, but I now believe this article (and indeed the whole blog) is AI-generated. Here's my evidence:

      * All the posts are walls of text. In many of the posts, all of the paragraphs are just about exactly the same length.

      * None of the posts contain personal stories, experiences, or anecdotes.

      * All of the posts are hyper-focused around advocacy of "old tech" and decentralization. Which is great, but real blogs with this many posts have at least _some_ variation in subject matter and quality. This one's surprisingly uniform.

      * The blog is relatively new, posts began about two months ago and there are 17 posts, that's a little over 2 posts a week. Sure, there are bloggers who post that often, sometimes daily, but they are usually much shorter posts, tend to be shorter, or do it for their job.

      * The picture of the author on the About page is very obviously AI generated.

      * The author has a GitHub account that had virtually no activity until March of this year, and basically all of the commits were co-authored by Claude.

      To be clear, I'm sure a real human is behind the site/posts (as opposed to a fully autonomous agent), but I'd bet dollars to donuts that each article started out its life as a very short prompt.

      • breakwaterlabs 7 hours ago
        Of course it's AI generated, its a benign, beige opinion being drummed up as some incredible insight. You can use Awk on plaintext? Incredible! No one has ever in the last 15 years of journald debates mentioned this.

        And every such burst of inspiration is accompanied by bullet-points to flog the dead horse further.

        Lest you think it wasn't mind-numbing enough, now we're turning it into a narrative arc worthy of a screenplay: text, coming of age, has become markdown (which the author certainly loves). Completing the arc, we're going to acknowledge its limitations ("be unbiased / balanced"), and then come to a denouement where we tie it up with a bland bow.

      • solaire_oa 21 hours ago
        The submitter posts approximately 5 articles per day, has relatively low comment engagement per submission, and the account was created in 2023, for what that's worth. Low trust.

        I vibed TrustSpark for HN plugin some months ago, and I get true positive flags from this account all the time. Unfortunately my RSS feed doesn't see the flags, and I have to click the link to find out.

        • FLeXMurphy 8 hours ago
          You should send TrustSpark to dang/tomhow because it will probably work better than what they seem to have running.
    • sillysaurusx 1 day ago
      Where did you get that understanding from? I might have missed something. They penalize LLM comments. Do they also penalize submissions? Any evidence of that?
      • FLeXMurphy 1 day ago
        dang mentioned it in passing some weeks back (going off of memory). Also please publish new NoH/Kongor build.
    • sthatipamala 1 day ago
      I skimmed this whole article and didn't see any classic LLM tells about it. Why do you think it is LLM generated?
      • tux3 1 day ago
        The recently released models have improved their writing styles. I don't know about this article, but if that's any indication of their writing habits, the author's github pages repo is full of posts generated by Claude.
        • pixl97 1 day ago
          Here's the thing, HN'ers have detected 200 of the last 10 AI written articles posted. We really seem to suck at it.

          Worse the tools for detecting it have an insane false detection rate. ESL = FP. Good writer = FP. Garbage human slop writer = A'ok.

          • rerdavies 22 hours ago
            I am very much looking forward to the day when I can use em-dashes again. I genuinely like em-dashes. Such a great punctuation mark!

            To my ears, I find accusations of "it was written by AI" to be a bit MAGA-ish. It's a far too convenient way to say that I don't like what someone is saying without actually having to think for yourself (or heaven forbid, actually read it).

          • tux3 1 day ago
            Detection tools in general have a bad selectivity and sensitivity, but there's consensus that Pangram is generally accurate and reliable. I don't use these tools much, but that's the word among people I trust, and that's what I see when I try it.

            And, well, neither of us really knows the ground truth. Maybe the guy who has a repo filled with articles co-authored by Claude is actually not posting slop this time and HN users are wrong and Pangram flagging it is just a false positive. But I can't blame anyone for being tired of giving this the benefit of the doubt.

            I'm ESL. The mistakes I make are part of what makes me human. If HN'ers are having an allergic reaction at the first hint of slop, it's because we've been drowning in it for months.

            • pixl97 4 hours ago
              It's a problem of Bullshit Asymmetry. The problem is at the end of the day Bullshit wins in unregulated environments where all actors aren't playing by the same rules.

              Furthermore, Panagram has something like 98% correctness now, but that's the best it's ever going to get. "If Panagram is effective it cannot remain effective" as it will just be used as a GAN.

      • gwbas1c 1 day ago
        IMO, the article is well-written.

        The author's picture clearly has an LLM-generated background, (and is a little tacky, IMO.) That being said, I wouldn't dismiss an article merely because someone used an LLM while setting up their blog.

      • loup-vaillant 1 day ago
        The whole structure had a faint smell to it. It's hard to put into word, there was no obvious "it's not the quiet part. It's the load bearing bedrock", but the way the thing flowed made me strongly suspect LLM involvement.
        • comradesmith 1 day ago
          First we had AI psychosis, now we have AI schizophrenia
          • loup-vaillant 1 day ago
            Well, I could be wrong, and it wouldn't be the first time. Besides, some of those "tells" are actually good tricks, when used sparingly enough. And we humans are picking up on them, making us sound a bit more like a chat bot than we used to.

            In the end we can't even tell what's what, it's exhausting.

    • john_strinlai 1 day ago
      they don't get @ mentions, you're better off emailing hn@ycombinator.com
    • tim333 9 hours ago
      Rather than banning you could have something like a pangram score in the 'new' list, then humans can vote or not as usual?
    • grebc 1 day ago
      I think you’ve posted on an unrelated topic.
      • FLeXMurphy 1 day ago
        No, this is the right article.
        • grebc 1 day ago
          Doesn’t read like AI.

          Even has a simple an/a mistake that I doubt a chatbot is going to make.

    • thuruv 1 day ago
      perused the posts and incidentally all are touching high-octane topics, posted only recently but surely warrants discussion/disputes, clearly an evidence of karma farming!
    • beej71 1 day ago
      The is real is real.
      • FLeXMurphy 1 day ago
        It is littered with little tells like that.