snapshottransfer
snapshottransfer — Transfers a snapshot to external storage in either one of two formats.
Synopsis
snapshottransfer [--servers {server-url[:port]}[,...]]
[--destination {path-or-url}]
[--format {csv|parquet}]
[--ssl {ssl-config-file}]
[--new]
[--dry-run]
[destination-specific-qualifiers...]
[format-specific-qualifiers...]
[ssl-security-qualifiers...]
Description
SnapshotTransfer sends the latest snapshot to an external storage system, which can be local disk storage, Amazon S3
storage, or Google Cloud storage. The argument to the --destination qualifier determines the target of
the transfer. For example, a directory path will save the output to a local file system. A URL starting with "gs://" will
send the output to Google cloud storage, while a URL starting "s3://" will deliver the files to an S3 "bucket". You can
also specify the format of the output with the --format qualifier The default is comma-separated value
(CSV) files. The alternative is Parquet format.
By default. SnapshotTransfer copies the latest snapshot. If you specify --new, it first initiates
a new snapshot and then sends that to the target destination.
The global qualifiers for the snapshottransfer command are:
- --servers {server-url[:port]} [,...]
Specifies the source of the snapshot as the client port(s) of one or more nodes of a VoltDB database. The default is "http://localhost:21212".
- --destination {path-or-url}
Specifies where the files are to be sent. A file path indicates the files will be saved to local disk storage. A URL starting "s3://" identifies an Amazon S3 bucket. A URL starting "gs://" identifies a Google Cloud storage destination. The default is the user's current default working directory.
- --format {csv|parquet}
Specifies the format of the generated files. The options are either comma-separated value (CSV) or Parquet format. The default is CSV.
- --ssl {ssl-properties-file}
Specifies the use of TLS/SSL encryption when communicating with the database server. Only necessary if the cluster is configured to use TLS encryption for the external ports. See the section called “Using CLI Commands with TLS/SSL” in the Using VoltDB manual for more information.
- --new
Initiates a new snapshot and sends its contents to the destination. If the
--newqualifier is not specified, the most recent snapshot is used as input.- --dry-run
Verifies that the command arguments are valid and that the command can be executed. No data is actually transferred during a dry run.
Destination Specific Qualifiers
Depending on the destination type, you may need to provide destination-specific information to complete the transfer. The following tables describe the qualifiers specific to Amazon S3 and Google Cloud delivery.
- Amazon S3 Qualifiers
Qualifier Description --s3-region The region where the snapshot will be stored. This is a required value, there is no default. --s3-endpoint For S3-compatible services other than Amazon, the host name address of the provider. At runtime, this host name is combined with the --destinationto create the target URL {hostname}/{destination}. For Amazon, the host name is generated automatically and this qualifier is not needed.--s3-access-key-id The account access key ID. --s3-secret-access-key The account secret access key. --s3-credentials A properties file specifying the access-key-id and secret-access-key values. You can either provide the credentials on the command line using the --s3-access-key-idand--s3-seret-access-keyqualifiers, or you can use--s3-credentialsand a properties file to avoid including sensitive information on the command line itself.- Google Cloud Qualifiers
Qualifier Description --gcp-project An optional Google Cloud project ID. For informational purposes only. --gcp-credentials The path to a JSON key file containing the Google account credentials.
Format Specific Qualifiers
Depending on the output format chosen, you can customize the delivery using format-specific qualifiers. The following tables describe the qualifiers specific to CSV and Parquet output.
- CSV Qualifiers
Qualifier Description --csv-separator The field delimiter. The default is a comma (,). --csv-quotechar The quote character. The default is the double quotation mark ("). --csv-escape The escape character. The default is the backslash (\). --csv-header
--no-csv-headerWhether column names are included as the first line. Column names are included by default. --csv-charset The output encoding. The default is UTF-8. --csv-null-string How null values are represented. The default is an empty space. - Parquet Qualifiers
Qualifier Description --parquet-compression Which compression codec to use. Supported codecs include SNAPPY, ZSTD, and UNCOMPRESSED. The default is SNAPPY. --parquet-row-group-size-mb The row group size in megabytes. The default is 256.
Usage
Unlike the other snapshot utilities that operate on existing local files regardless of the database state, snapshottransfer requires access to a running VoltDB database to perform its tasks. On the other hand, the snapshottransfer command transfers the snapshot files for the entire database cluster with a single command, not just partial files from a single server.
When you invoke snapshottransfer, your local process acts as the transport for the files. It receives the data from the Volt server(s) and then sends it to the target destination. This means your terminal session will not return to the command prompt until the transfer is complete (or an error occurs). If the current process stops for any reason (for example, if you press CTRL/C), the transfer will be interrupted.
Both CSV and Parquet format are transmitted as multiple files The files are organized into a directory structure:
The top level is a directory named after the snapshot itself.
Within the top level directory, there are subdirectories for each table.
Within each table directory there can be multiple files that contain the table data.
At both the top level and in each table directory, there is a zero-byte file called _SUCCESS indicating that the transfer of the table contents or database as a whole has been completed successfully.
For example, saving CSV files of a database with two tables to the local directory
/var/archive/ might create the following file structure:
/var/
archive/
snapshot-20260824T143022/
ORDERS/
part-00000.csv
part-00001.csv
_SUCCESS
CUSTOMERS/
part-00000.csv
_SUCCESS
_SUCCESS
Examples
The following command transfers the most recent database snapshot to an Amazon S3 bucket in the
us-east-2 region in Parquet format. The credentials for accessing S3 are provided on the
command line in the --s3-access-key-id and --s3-secret-access-key qualifiers:
$ snapshottransfer --servers mydbserver1:21212 \ --format parquet \ --destination s3://mycompany/snapshots/ \ --s3-region us-east-2 \ --s3-access-key-id myaccesskey \ --s3-secret-access-key topsecret
The next example initiates a new snapshot and then transfers its contents to Google Cloud storage as CSV files. In
this case, The credentials for accessing S3 are provided in the properties file
~/credentials/googlecred.properties:
$ snapshottransfer \ --servers mydbserver1:21212 \ --new \ --format csv \ --destination gs://mycompany/snapshots/ \ --gcp-credentials ~/credentials/googlecred.properties
Documentation