Re: SRFI 207: String-notated bytevectors Daphne Preston-Kendal 18 Aug 2020 09:48 UTC

On 17 Aug 2020, at 18:10, Marc Nieper-Wißkirchen <xxxxxx@nieper-wisskirchen.de> wrote:

> I think the main reason is that if we allowed the full Unicode range
> and specified UTF-8 encoding, sequences like \x80; would be ambiguous.
> Either the byte 80 is meant or the bytes corresponding to the UTF-8
> encoding of U+0080.

I don’t understand what you mean here. Do you mean that you think #u8"\x80;" is
ambiguous because it's not clear whether it's meant to be #u8(#x80) or #u8(#xC2
#x80)? I kind of understand this point, since it’s not too far removed from my
point about encoding confusions when non-ASCII characters are allowed to be
used. But I think the byte numbers being written out explicitly means there’s
no ambiguity here.

> In order to allow implementations to extend #u8"..." so that "..." can
> be any string allowed by the implementation (*), I want to suggest to
> rename the sequence "\xHH;" of this SRFI into something different like
> "\yHH;".

So what if I have a normal string and I use a \y escape in it? What happens
then? Do I get Unicode character U+00HH in my string? What’s the point of using
a different escape character in that case? That semantic still seems to break the
equivalence with string->utf8.

Sorry if I’m being obtuse here.

Daphne