Amazon Redshift is a fully managed data warehousing service provided by Amazon Web Services (AWS). It's designed to handle large-scale data analytics and complex querying, offering high-performance columnar storage and parallel processing capabilities to efficiently process and analyze large datasets for business insights.
Authorize Connection to Amazon Redshift
Prerequisites
- A running Amazon Redshift instance reachable by Dataddo. If you use a firewall, allow the Dataddo IPs.
- A database user with rights to create tables and to insert, update, and delete rows in the target schema.
Create the authorizer
In Authorizers, click Authorize New Service and select Amazon Redshift, then fill in the fields:
| Field | Description |
|---|---|
| A name for this authorizer in Dataddo | A label so you can recognize the connection later. |
| Server IP or Hostname | Public IP or hostname of your Redshift cluster. |
| Database | Name of the database you will use for writing or reading the data. |
| Username | Username for authentication. |
| Password | Password for authentication. |
| Port | Port to connect to Redshift. The default value is 5439. |
| Use SSH tunnel | In case you need to connect via SSH tunnel, create it first on the Security setting page and use pick it here afterwards. |
Click Save. Dataddo validates the connection.
Create an Amazon Redshift Destination
Go to Destinations, click Create Destination, and select Amazon Redshift. Give the destination a name, choose the authorizer you created, and click Save.
Write Modes
The write mode decides how each flow run changes the target table. Amazon Redshift offers the standard set below, and the default is insert. Which modes are available, and whether you can switch modes on an existing flow, can depend on the table's current state.
| Mode | Meaning | Composite Key |
|---|---|---|
insert |
Appends new rows on every run. | No |
insert_ignore |
Appends rows, skipping rows that already exist. | Yes |
truncate_insert |
Empties the table, then writes the current data. | No |
upsert |
Inserts new rows and updates existing rows matched by the Composite Key. | Yes |
update |
Updates existing rows matched by the Composite Key; fails rows with no match. | Yes |
update_ignore |
Updates matched rows, skipping rows with no match. | Yes |
delete |
Deletes rows matched by the Composite Key. | Yes |
Composite Key
The Composite Key is the set of columns whose combined values identify each row. Every mode except insert and truncate_insert needs one. You can select up to 4 columns.
Table Naming
Dataddo creates the table on the first write if it does not exist. When you name the table:
- Table name: table name should contains only letters, numbers and/or underscores/dashes.
- You can put date-range patterns such as
{{1d1}}in the table name.
How to Create a Flow to Amazon Redshift
- Go to Flows and click Create Flow.
- Add one or more sources.
- Add Amazon Redshift as the destination and pick the authorizer.
- Choose a write mode, and a Composite Key for modes that need one.
- Set the schema and the table name.
- Set the schedule and click Save.
Troubleshooting
Invalid schema or table name
The schema or table name does not match Amazon Redshift's naming rules. Fix the name in the flow to match the rules in Table Naming, then restart the flow.
Insufficient privileges
The database user cannot create tables or write rows in the target schema. Grant the user create, insert, update, and delete rights, then restart the flow.
Cannot connect to Amazon Redshift
Dataddo cannot reach the database. Check the host, port, and credentials in the authorizer, and make sure the server is reachable by Dataddo. If you use a firewall, allow the Dataddo IPs.