Re: GIT character codecs

2 messages, 2 authors, 2016-06-15 · open the first message on its own page

Re: GIT character codecs

From: Marco Costalba <hidden>
Date: 2016-06-15 22:42:11

Junio C Hamano wrote:
Marco Costalba [off-list ref] writes:
quoted
I modified qgit to force the use of utf-8 codec instead of the local one, 
in my case ISO-8859-15.

I suspect that trying to have a globa single encoding is a wrong
approach overall.  There are a handful issues to think about.
I agree.
First the easiest one -- the commit log messages.  We encourage
use of UTF-8, 

[cut]

But we do not _enforce_ UTF-8.

What this means for qgit is that it is often sufficient for its
log browser to support one encoding at a time, provided if it
allows the user to switch which encoding to use depending on
what project is being viewed.


[cut]


But other projects may use different
encodings, and even a single project can have its i18n message
files in different encodings and charsets in different files.
Users probably want to be able to view all of them, even if they
only understand a couple of languages and not others.

What this means for qgit is that at least you should be able to
show a whole file in a single encoding, but if you show more
than two files at the same time, one in each window, these
windows may be showing its contents in different encodings and
charsets.  So you would need to give a way to your users to tell
you what encoding each file is in.  Using global locale as the
default and having a way to override that per file basis would
be sufficient.
If encoding is a per-blob _and_ per-log message property a real solution, although cumbersone,
could be that git stores encoding togheter with the blob and the commits.

The interfaces to read out encoded info could be something like 

git-rev-list --pretty  --encoding
git-ls-files -e

for commits and blobs respectively.

This avoids all user settings. More, the user could do not know the encoding, as
example when browsing a public repo. 


But I understand is too late for something like this. So what qgit can do is
to show a setting with a list of all supported encodings from the user to choose from.

But this is also a problem because, first we should have _two_ settings, one from log 
messages and one from blobs _and_ also in this case we miss the problem of different 
blobs with different encoding.

A practical workaround, *that do not catches all cases at all* could be qgit has only one global
setting (defaulting to local codec) and applies that setting blindly to everything, from commits
to blobs, the user changes that setting and refreshes the view until he finds the correct one.

But I'm a bit reclutant to add this 'codec list' stuff because, as said, is *not* the real
solution.

   Marco




		
__________________________________ 
Start your day with Yahoo! - Make it your home page! 
http://www.yahoo.com/r/hs

Re: GIT character codecs

From: Linus Torvalds <torvalds@osdl.org>
Date: 2016-06-15 22:42:11


On Sun, 13 Nov 2005, Marco Costalba wrote:
If encoding is a per-blob _and_ per-log message property a real solution, although cumbersone,
could be that git stores encoding togheter with the blob and the commits.
We'd be much better off with just saying "we encourage people to use 
utf-8, but if you don't, just set your locale to make things show up 
properly".

utf-8 is clearly the future, and if we make git internally aware of 
locales, that's just going to complicate things. And usually for no good 
reason, since users don't really care that much.

I really hate codepages. I'd much rather say:
 - git is 8-bit clean, so you can use any damn encoding you want
 - utf-8 is strongly recommended for all the same reasons it's recommended 
   for anything else.

Hmm?

		Linus
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help