I played with making a git alternative that was streamlined according to workflow. I haven't completely abandoned it, but my core idea was that you can use ast path to compress the data further.
Like you said git is really efficient, and even though I went in with some criticism because of the work(mess) that is chromium and the deep hacking I did there to fix stuff between versions without forcing a complete rebuild
I came to the conclusion that as a compression alternative, my approach wasn't worth it. Git did a better job!
This is one of my biggest pet peeves! I hate reading shell scripts which uses small options. It is okay to use short options when interactively using shell, but there is no excuse to not spend time and try converting a small options shell script to use long options if its going to be read by someone else, especially if you anyway have to write comments explaining it.
The biggest part i hate is that, some utilities don’t have long option counterpart!
No. You can configure the compression level but that's at far as it goes.
You can configure loose object and pack compression separately (core.compression and pack.compression) so loose object compression could be switched to an other algorithm as it's completely local[1] (likely a low-compression zstd), but pack files are part of the network exchange, so things are a bit more complicated there. Also an annoyance there is that git uses raw zlib there is no identifying header or anything.
There are old threads on the mailing list, as well as a gitlab issue.
Although you can also tune compression per pack when you create them explicitly, on $DAYJOB's repository in order to speed up gc some I have .keep packs which segregate images uncompressed in their own packs, because there's little point wasting cycles on zlib-compressing jpegs and pngs.
[1]: unless you're using the "dumb http" protocol, or performing (n)fs-mediated clones
It can get even smaller when loose objects get packed into packfiles because of delta compression: a whole bunch of identical files can be stored in minimal overhead over storing one copy.
Right, packfile size is really what matters in the long-term. A tiny change to a file in the short term just adds the whole thing again (compressed), whereas a tiny change in a packfile shouldn't add much at all. Unless you've got very limited disk space, the long-term is the only size that matters.
Like you said git is really efficient, and even though I went in with some criticism because of the work(mess) that is chromium and the deep hacking I did there to fix stuff between versions without forcing a complete rebuild
I came to the conclusion that as a compression alternative, my approach wasn't worth it. Git did a better job!
Anyway here's the attempt: https://github.com/pankon/gat
Thanks!
Small nitpick: Use long options and your code samples become self explanatory!
You can configure loose object and pack compression separately (core.compression and pack.compression) so loose object compression could be switched to an other algorithm as it's completely local[1] (likely a low-compression zstd), but pack files are part of the network exchange, so things are a bit more complicated there. Also an annoyance there is that git uses raw zlib there is no identifying header or anything.
There are old threads on the mailing list, as well as a gitlab issue.
Although you can also tune compression per pack when you create them explicitly, on $DAYJOB's repository in order to speed up gc some I have .keep packs which segregate images uncompressed in their own packs, because there's little point wasting cycles on zlib-compressing jpegs and pngs.
[1]: unless you're using the "dumb http" protocol, or performing (n)fs-mediated clones
Well to the extent that git is able to find a suitable delta, which is the difficult part.