Re: Unicode lambda Lassi Kortela 12 May 2019 14:38 UTC

> I did something similar (but in pseudo-code, not code) in my old blog
> post "Hello!  I am an XML encoding sniffer" at
> <http://recycledknowledge.blogspot.com/2005/07/hello-i-am-xml-encoding-sniffer.html>.

That's great :) I just changed it so it skips all bytes except 1..126.
So now it can read UTF-16, both big- and little-endian, with or without
byte order mark.

To be sure, if there's more data in it than just the encoding
declaration, that data should be able to use bytes above 126. But for
parsing only the encoding declaration this skipping is appropriate.

> As long as you are using a custom reader, however, you might as well
> make the whole declare-file form a comment so that the regular reader
> can just start over.

It's nice to be able to do that, but maybe it shouldn't be a comment.
After all, it could contain metadata that is useful to read from Lisp
and much of the point of S-expression syntax is that it's extensible
like this. If we make it a comment, it's just a glorified version of the
original magic comments hack :-/

For implementations that don't support it, skipping the thing is a
one-line macro:

;; Scheme:
(define-syntax declare-file (syntax-rules () ((_ rest ...) #f)))

;; Common Lisp:
(defmacro declare-file (&body body) (declare (ignore body)) nil)

;; Emacs Lisp:
(defmacro declare-file (&rest _rest) nil)

;; Clojure:
(defmacro declare-file [& body] nil)