Documentation Index

Fetch the complete documentation index at: https://docs.dataddo.com/llms.txt

Use this file to discover all available pages before exploring further.

Amazon S3

Prev Next

Amazon S3 (Simple Storage Service) is a scalable cloud storage service provided by Amazon Web Services (AWS). It offers a way to store and retrieve data, such as files, images, videos, and backups, in a highly durable and easily accessible manner, making it a foundational component for various cloud-based applications and services.

Authorize Connection to AWS S3

Authorize the connection so Dataddo can write files to AWS S3:

AWS S3

Field Description
A name for this authorizer in Dataddo A label so you can recognize the connection later.
Bucket Provide the identifier of S3 bucket you want to use for reading or writing the data.
Region Region of the S3 bucket.
Key Provide the value of AWS Key. Please check our documentation if unsure how to configure it.
Secret Provide the value of AWS Secret. Please check our documentation if unsure how to configure it.

AWS Assume Role

Field Description
A name for this authorizer in Dataddo A label so you can recognize the connection later.
External ID System-generated and stable for your account. Add it to your IAM role's trust policy as the sts:ExternalId condition.
Role ARN The ARN of the IAM role in your AWS account that Dataddo will assume.

Dataddo validates the connection when you save it.

Create an AWS S3 Destination

Go to Destinations, click Create Destination, and select AWS S3. Give the destination a name, choose the authorizer, then set:

Field Description
Bucket Required when using an AWS Assume Role authorizer. Leave empty with a classic S3 authorizer, which already carries the bucket. Shown when oAuthId is ``.
Region Region of the S3 bucket. Required when using an AWS Assume Role authorizer. Shown when oAuthId is ``.
Path Enter or select the directory where the data file will be created. Shown when oAuthId is ``.

Click Save.

Supported File Formats

Each flow run writes the data as a file in the format you pick.

Format Description
csv AWS S3 writes the data as CSV files.
json AWS S3 writes the data as JSON files.
jsonl AWS S3 writes the data as JSONL files.
parquet AWS S3 writes the data as PARQUET files.

For CSV you can set the delimiter, header row, and date formatting. For Parquet, JSON, and JSONL you can set the timestamp unit.

File Naming

You can also build a custom filename with these placeholders (see Dynamic File Naming Patterns for the full list):

  • {{objectLabel}} and {{objectId}}: the flow name and id.
  • {{today}} and {{yesterday}}: the run date.
  • {{dateRangeStart}} and {{dateRangeEnd}}: the bounds of the flow's date range.
  • Date-range expressions such as {{1d1}} (yesterday) or {{90d1}} (the last 90 days through yesterday).
  • Add a date format after a |, for example {{today|Ymd}} gives 20201231 and {{1d1|Y-m-d}} gives 2020-12-31.

File Partitioning

File partitioning splits a large dataset into smaller files based on a criterion such as date, which improves how a data lake organizes and queries the data (see Data Lake Ingestion). In Dataddo you partition by putting date variables from File Naming into the file name, so each flow run writes its own dated file.

For example, events_{{1d1|Y-m-d}}.parquet writes one Parquet file per day, so a lake engine such as AWS S3 can read the set of files as a date-partitioned dataset. Pick a file format like Parquet or CSV that your lake reads, and schedule the flow to match the partition period (for example daily for a daily date token).

Write Modes

Each flow run writes a file at the resolved name. The default is truncate_insert: insert keeps writing new files, and truncate_insert replaces the file at the same name.

How Data Is Delivered

Every run produces one file at the resolved name. A date-stamped name accumulates a new snapshot file per run, which is the pattern described in Data Lake Ingestion.

How to Create a Flow to AWS S3

  1. Go to Flows and click Create Flow.
  2. Add one or more sources.
  3. Add AWS S3 as the destination and pick the authorizer.
  4. Choose the file format and the file name.
  5. Set the schedule and click Save.

Troubleshooting

Cannot connect to AWS S3

Dataddo cannot reach AWS S3. Check the credentials, path, and permissions in the authorizer and destination, and make sure AWS S3 is reachable by Dataddo.

File is not created

The account used by Dataddo lacks write permission on the target path. Grant write access to the folder or bucket and restart the flow.

Related Articles