Although 2to3 will fix most issues in Python 2 code to make it run under Python 3, it does not handle the new strict separation between byte strings and unicode strings. There is one instance in git_remote_helpers where we are caught by this, which is when reading refs from "git for-each-ref". Fix this by operating on the returned string as a byte string rather than a unicode string. As this method is currently only used internally by the class this does not affect code anywhere else. Note that we cannot use byte strings in the source as the 'b' prefix is not supported before Python 2.7 so in order to maintain compatibility with the maximum range of Python versions we use an explicit call to encode(). Signed-off-by: John Keeping <john@xxxxxxxxxxxxx> --- On Tue, Jan 15, 2013 at 02:04:29PM -0800, Junio C Hamano wrote: > John Keeping <john@xxxxxxxxxxxxx> writes: >>> That really feels wrong. Displaying is a separate issue and it is >>> the _right_ thing to punt the problem at the lower-level machinery >>> level. >> >> But the display will require decoding the ref name to a Unicode string, >> which depends on the encoding of the underlying ref name, so it feels >> like it should be decoded where it's read (see [1]). > > If you botch the decoding in a way you cannot recover the original > byte string, you cannot create a ref whose name is the original byte > string, no? Keeping the original byte string internally (this > includes where you use it to create new refs or update existing > refs), and attempting to convert it to Unicode when you choose to > show that string as a part of a message to the user (and falling > back to replacing some bytes to '?' if you cannot, but do so only in > the message), you won't have that problem. Actually, this method is currently only used internally so I don't think my argument holds. This is what keeping the refs as byte strings looks like. git_remote_helpers/git/importer.py | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/git_remote_helpers/git/importer.py b/git_remote_helpers/git/importer.py index e28cc8f..c54846c 100644 --- a/git_remote_helpers/git/importer.py +++ b/git_remote_helpers/git/importer.py @@ -18,13 +18,16 @@ class GitImporter(object): def get_refs(self, gitdir): """Returns a dictionary with refs. + + Note that the keys in the returned dictionary are byte strings as + read from git. """ args = ["git", "--git-dir=" + gitdir, "for-each-ref", "refs/heads"] - lines = check_output(args).strip().split('\n') + lines = check_output(args).strip().split('\n'.encode('utf-8')) refs = {} for line in lines: - value, name = line.split(' ') - name = name.strip('commit\t') + value, name = line.split(' '.encode('utf-8')) + name = name.strip('commit\t'.encode('utf-8')) refs[name] = value return refs -- 1.8.1 -- To unsubscribe from this list: send the line "unsubscribe git" in the body of a message to majordomo@xxxxxxxxxxxxxxx More majordomo info at http://vger.kernel.org/majordomo-info.html