From: Junio C Hamano <hidden> Date: 2016-06-15 22:55:46
John Keeping [off-list ref] writes:
quoted
That really feels wrong. Displaying is a separate issue and it is
the _right_ thing to punt the problem at the lower-level machinery
level.
But the display will require decoding the ref name to a Unicode string,
which depends on the encoding of the underlying ref name, so it feels
like it should be decoded where it's read (see [1]).
If you botch the decoding in a way you cannot recover the original
byte string, you cannot create a ref whose name is the original byte
string, no? Keeping the original byte string internally (this
includes where you use it to create new refs or update existing
refs), and attempting to convert it to Unicode when you choose to
show that string as a part of a message to the user (and falling
back to replacing some bytes to '?' if you cannot, but do so only in
the message), you won't have that problem.
From: John Keeping <hidden> Date: 2016-06-15 22:55:46
Although 2to3 will fix most issues in Python 2 code to make it run under
Python 3, it does not handle the new strict separation between byte
strings and unicode strings. There is one instance in
git_remote_helpers where we are caught by this, which is when reading
refs from "git for-each-ref".
Fix this by operating on the returned string as a byte string rather
than a unicode string. As this method is currently only used internally
by the class this does not affect code anywhere else.
Note that we cannot use byte strings in the source as the 'b' prefix is
not supported before Python 2.7 so in order to maintain compatibility
with the maximum range of Python versions we use an explicit call to
encode().
Signed-off-by: John Keeping <redacted>
---
On Tue, Jan 15, 2013 at 02:04:29PM -0800, Junio C Hamano wrote:
John Keeping [off-list ref] writes:
quoted
quoted
That really feels wrong. Displaying is a separate issue and it is
the _right_ thing to punt the problem at the lower-level machinery
level.
But the display will require decoding the ref name to a Unicode string,
which depends on the encoding of the underlying ref name, so it feels
like it should be decoded where it's read (see [1]).
If you botch the decoding in a way you cannot recover the original
byte string, you cannot create a ref whose name is the original byte
string, no? Keeping the original byte string internally (this
includes where you use it to create new refs or update existing
refs), and attempting to convert it to Unicode when you choose to
show that string as a part of a message to the user (and falling
back to replacing some bytes to '?' if you cannot, but do so only in
the message), you won't have that problem.
Actually, this method is currently only used internally so I don't think
my argument holds.
This is what keeping the refs as byte strings looks like.
git_remote_helpers/git/importer.py | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
@@ -18,13 +18,16 @@ class GitImporter(object):defget_refs(self,gitdir):"""Returns a dictionary with refs.++Notethatthekeysinthereturneddictionaryarebytestringsas+readfromgit."""args=["git","--git-dir="+gitdir,"for-each-ref","refs/heads"]-lines=check_output(args).strip().split('\n')+lines=check_output(args).strip().split('\n'.encode('utf-8'))refs={}forlineinlines:-value,name=line.split(' ')-name=name.strip('commit\t')+value,name=line.split(' '.encode('utf-8'))+name=name.strip('commit\t'.encode('utf-8'))refs[name]=valuereturnrefs
From: Pete Wyckoff <hidden> Date: 2016-06-15 22:55:47
john@keeping.me.uk wrote on Tue, 15 Jan 2013 22:40 +0000:
This is what keeping the refs as byte strings looks like.
As John knows, it is not possible to interpret text from a byte
string without talking about the character encoding.
Git is (largely) a C program and uses the character set defined
in the C standard, which is a subset of ASCII. But git does
"math" on strings, like this snippet that takes something from
argv[] and prepends "refs/heads/":
strcpy(refname, "refs/heads/");
strcpy(refname + strlen("refs/heads/"), ret->name);
The result doesn't talk about what character set it is using,
but because it combines a prefix from ASCII with its input,
git makes the assumption that the input is ASCII-compatible.
If you feed a UTF-16 string in argv, e.g.
$ echo master | iconv -f ascii -t utf16 | xargs git branch
xargs: Warning: a NUL character occurred in the input. It cannot be passed through in the argument list. Did you mean to use the --null option?
fatal: Not a valid object name: ''.
you get an error about NUL, and not the branch you hoped for.
Git assumes that the input character set contains roughly ASCII
in byte positions 0..127.
That's one small reason why the useful character encodings put
ASCII in the 0..127 range, including utf-8, big5 and shift-jis.
ASCII is indeed special due to its legacy, and both C and Python
recognize this.
@@ -18,13 +18,16 @@ class GitImporter(object): def get_refs(self, gitdir): """Returns a dictionary with refs.++ Note that the keys in the returned dictionary are byte strings as+ read from git. """ args = ["git", "--git-dir=" + gitdir, "for-each-ref", "refs/heads"]- lines = check_output(args).strip().split('\n')+ lines = check_output(args).strip().split('\n'.encode('utf-8')) refs = {} for line in lines:- value, name = line.split(' ')- name = name.strip('commit\t')+ value, name = line.split(' '.encode('utf-8'))+ name = name.strip('commit\t'.encode('utf-8')) refs[name] = value return refs
I'd suggest for this Python conundrum using byte-string literals, e.g.:
lines = check_output(args).strip().split(b'\n')
value, name = line.split(b' ')
name = name.strip(b'commit\t')
Essentially identical to what you have, but avoids naming "utf-8" as
the encoding. It instead relies on Python's interpretation of
ASCII characters in string context, which is exactly what C does.
-- Pete
From: John Keeping <hidden> Date: 2016-06-15 22:55:47
On Tue, Jan 15, 2013 at 07:03:16PM -0500, Pete Wyckoff wrote:
john@keeping.me.uk wrote on Tue, 15 Jan 2013 22:40 +0000:
quoted
This is what keeping the refs as byte strings looks like.
As John knows, it is not possible to interpret text from a byte
string without talking about the character encoding.
Git is (largely) a C program and uses the character set defined
in the C standard, which is a subset of ASCII. But git does
"math" on strings, like this snippet that takes something from
argv[] and prepends "refs/heads/":
strcpy(refname, "refs/heads/");
strcpy(refname + strlen("refs/heads/"), ret->name);
The result doesn't talk about what character set it is using,
but because it combines a prefix from ASCII with its input,
git makes the assumption that the input is ASCII-compatible.
If you feed a UTF-16 string in argv, e.g.
$ echo master | iconv -f ascii -t utf16 | xargs git branch
xargs: Warning: a NUL character occurred in the input. It cannot be passed through in the argument list. Did you mean to use the --null option?
fatal: Not a valid object name: ''.
you get an error about NUL, and not the branch you hoped for.
Git assumes that the input character set contains roughly ASCII
in byte positions 0..127.
That's one small reason why the useful character encodings put
ASCII in the 0..127 range, including utf-8, big5 and shift-jis.
ASCII is indeed special due to its legacy, and both C and Python
recognize this.
@@ -18,13 +18,16 @@ class GitImporter(object): def get_refs(self, gitdir): """Returns a dictionary with refs.++ Note that the keys in the returned dictionary are byte strings as+ read from git. """ args = ["git", "--git-dir=" + gitdir, "for-each-ref", "refs/heads"]- lines = check_output(args).strip().split('\n')+ lines = check_output(args).strip().split('\n'.encode('utf-8')) refs = {} for line in lines:- value, name = line.split(' ')- name = name.strip('commit\t')+ value, name = line.split(' '.encode('utf-8'))+ name = name.strip('commit\t'.encode('utf-8')) refs[name] = value return refs
I'd suggest for this Python conundrum using byte-string literals, e.g.:
lines = check_output(args).strip().split(b'\n')
value, name = line.split(b' ')
name = name.strip(b'commit\t')
Essentially identical to what you have, but avoids naming "utf-8" as
the encoding. It instead relies on Python's interpretation of
ASCII characters in string context, which is exactly what C does.
From: Pete Wyckoff <hidden> Date: 2016-06-15 22:55:48
john@keeping.me.uk wrote on Wed, 16 Jan 2013 09:45 +0000:
On Tue, Jan 15, 2013 at 07:03:16PM -0500, Pete Wyckoff wrote:
quoted
I'd suggest for this Python conundrum using byte-string literals, e.g.:
lines = check_output(args).strip().split(b'\n')
value, name = line.split(b' ')
name = name.strip(b'commit\t')
Essentially identical to what you have, but avoids naming "utf-8" as
the encoding. It instead relies on Python's interpretation of
ASCII characters in string context, which is exactly what C does.
Drat. The b'' syntax seems to work on 2.6.8, in spite of
the docs, but certainly isn't in 2.5.
I think you had hit on the best compromise with encoding,
but maybe ascii is a little less presumptuous than utf-8,
and more indicative of the encoding assumption.
-- Pete