Re: Unicode lambda Lassi Kortela 12 May 2019 12:23 UTC

> 'read' can occur strictly before interpreting any of S-expressions, and reading in incorrect encoding can
> cause an I/O error so you may not have a chance to interpret those forms.

> The source file encoding should be a property of the port (as is the
> case-folding property). It could be set with a "#!" directive (at the
> top of the file).

The 'read' procedure that looks for the encoding declaration should be a
special reader that's much simpler than the normal Scheme reader and
handles encoding errors gracefully (perhaps it should simply read raw
bytes from a binary port, treat 0..127 as ASCII characters, and ignore
all other characters or treat them as whitespace or symbol/string
constituent).

Skipping/parsing the Unix shebang line (#!) at the start of a script is
in many Schemes/Lisps a similar magic job that needs its own reader (or
a hack to the standard reader).

> BTW, the "magic encoding comment" is supported in a few languages:
>
> Python: https://www.python.org/dev/peps/pep-0263/
> Ruby: https://idiosyncratic-ruby.com/26-file-encoding-magic.html

The fact that it's widespread is nice, but the syntax is not really well
specified. Ruby uses this regexp to look for it:

ENCODING_SPEC_RE = %r"coding\s*[=:]\s*([[:alnum:]\-_]+)"

Python uses this regexp:

^[ \t\f]*#.*?coding[:=][ \t]*([-_.a-zA-Z0-9]+)

Since the Scheme/Lisp is traditionally to resist the temptation of quick
hacks and do things in a principled way, I would really like to avoid
this approach even though other languages and editors are doing it. If a
magic comment syntax is to way to go, I would at least like to have a
principled and well specified grammar that can also be used other kinds
of magic comments (I have been collecting samples from many languages in
the hopes of specifying such a grammar, but I don't know if anyone would
adopt it).

> Technically the encoding info should be a metadata of a file, not in the content of the file, so the "coding" comment
> is certainly a kluge.  What I thought is that it might be useful to codify the current practice.

Correct, but it's certainly good to have the info somewhere since the
whole world is not Unicode yet.