Join GitHub today
GitHub is home to over 50 million developers working together to host and review code, manage projects, and build software together.
Sign upVersion Information for debugging #2275
Comments
|
Hi @nand28, Yes, it is intended that both old-to-new and new-to-old roundtrips should always work (as long as both endpoints are >=v0.8). The compression format is stable. The only exception to that is that we sometimes find bugs in the decompressor and fix them (you can search closed PRs with the
The first four bytes of the compressed frame are, in effect, a version identifier.
Yes, you can enable checksumming on any frame, though you may have to use a different API entry point than you are used to in order to do so. For example,
See this comment for more. |
|
Thanks Felix for the quick response. In case of unexpected bugs during decompression, we would like to detect the exact version of the zstd with which the data was compressed with. The magic, provides info only about the stable release and we will mostly be using zstd versions later to this. Is it possible to accommodate version in the compressed data? For checksum inclusion, is it necessary to reset the context everytime as that could be adding to delays. I mean, what are the sticky parameters that should be taken care while reusing context for singlestep compression? Does decompression validate the data with the checksum stored? |
There aren't any fields in a zstd frame that are really appropriate for that purpose. You could however append a skippable frame with whatever metadata you want.
If you want to use the same parameters every time there's no need to reset. Though it should also be noted that resetting the context is extremely cheap.
Yes. |
|
The skippable frame is designed for watermarking. The content of the skippable frame is yours, you can add anything you want into it, from the Note though that the content is not free, and will consume bandwidth. So in regular production mode, you will likely want to limit the bandwidth impact of watermarking, either reducing the payload, and/or sampling, or disable it altogether. |
|
Thats cool. These skippable frames have to be appended to compressed buffer or are there any apis for skippable frame handling? Upon checksum validation failure, what is the error returned? Is there any way to read this checksum? |
Skippable frames can be in front or after a zstd frame. From an exploitation perspective it may matter :
Very little. The main document is the format specification, There is no api to generate a skippable frame. There is however an (optional) api symbol able to detect a skippable frame. Note though that this function is part of the advanced API, which is not labelled stable,
The content size of the skippable frame is part of its header.
This is a form of multi-frames. Impact on decoding stage varies, from being completely transparent, when invoking
Read the last 4 bytes from the compressed frame. |
|
Thanks Yann for clarifying. Can the content size of skippable frame be zero at times? Will it have any impact? |
Yes it can
Well, it just occupies 8-bytes. That's all. |
|
Thanks for all the clarifications.
Can we get an API around this to return the content checksum given a compressed buf, if the checksum flag is present? Also, is this checksum calculation done as a separate pass of the input data? Or is it calculated in the same pass as the compression happens? In short, if we enable content checksum, will the compression time increase by the time taken by xxH64() to compute the checksum? |
To be fair, I don't see a good enough use case to justify an additional entry point.
Yes
Yes. |
Hi,
ZSTD guarantees successful decompression of data compressed with older versions of zstd. But is the same guaranteed vice versa? That is if we need to compress with newer version of zstd (at one node) and decompress with older version of zstd (at another node). If that can break at any point in future, we would like to identify such cases by knowing the version with which the data is compressed. The version of zstd which is currently used for compression can be obtained via ZSTD_version() api but similarly, can we obtain the zstd version of the compressed data? Like an optional checksum appended to the compressed data, can we have an optional version number appended to the compressed data and the same is verified for compatibility during decompression?
Also, is optional checksum allowed for single step compression using context?
Thanks.