Re: Unicode lambda Marc Nieper-Wißkirchen 21 Oct 2019 18:43 UTC

Am Mo., 21. Okt. 2019 um 19:39 Uhr schrieb John Cowan <xxxxxx@ccil.org>:
>
>
>
> On Mon, Oct 21, 2019 at 1:19 PM Shiro Kawai <xxxxxx@gmail.com> wrote:
>
>> So I suggest to keep encoding directive simple, parsable with a small finite automaton without lookahead.
>
>
> Agreed.   That makes it tantamount to a known-plaintext attack, and since there is no attempt at secrecy, that's easy.  See <http://recycledknowledge.blogspot.com/2005/07/hello-i-am-xml-encoding-sniffer.html> for an algorithm for sniffing XML encodings, where a declaration of the form <!?xml encoding="blahblah"?> (but with some additional bells and whistles possible) is not required if the encoding is UTF-8, any UTF-16 variant, or any UTF-32 variant.
>
>> Does #!encoding count?  Or should it be ignored?
>
>
> Ignored unless it is on a line by itself and as early as possible.  Except in EBCDIC files (Crom forbid it!), no non-ASCII characters should appear before it.

Is R7RS-small crystal-clear about that matter? Does

#;(bla #!fold-case blubb)

cause subsequent identifiers being case-folded or not? It's probably a
good idea if directives are skipped in comments but I am not so sure
whether it is what the R7RS document says.

>
> That said, UTF-8 is as near as not universal now: 95% of all web documents, though rather less on particular local systems.
>
>
>
> John Cowan          http://vrici.lojban.org/~cowan        xxxxxx@ccil.org
> If I have seen farther than others, it is because I was looking through a
> spyglass with my one good eye, with a parrot standing on my shoulder. --"Y"
>