Amazon S3 - Frequently Asked Questions

Prev Next

If you cannot find the answers you need, please contact us at customersupport@hginsights.com, and we will be happy to assist you


You can find the main documentation on the S3 integration here:

Q: How does the S3 integration work with RGIP?

Amazon S3 is a storage service that RGIP uses to transfer data from data warehouses like Redshift or from systems that do not have a direct integration with us. You will transfer data to an S3 bucket, from which RGIP Scoring will pull this data for scoring or Sales Copilot purposes.

Q: Should I host the S3 bucket or does RGIP have one?

You will need to host the S3 bucket on your side according to our security policy and grant RGIP access to your S3 bucket using an IAM role (see How to create an S3 bucket and give RGIP access).

Q: What are the requirements to set up an S3 integration with RGIP?

You will need an AWS account and assistance from someone, often a data engineer, who can stream data from your source to the S3 bucket.

Q: What type of data should I send RGIP?

RGIP can only ingest events from your S3 bucket. We plan to add support for contact and account attributes in the future.

What events specifically? You can refer to this article to know:What type of events can be used in a behavioral segmentation

Q: Should the data be transformed before sending it to S3?

Yes, we recommend following the instructions in How to format data and files in the S3 bucket to format the data from your system or data warehouse to the S3 bucket.

Q: How should the data be transferred?

We require 9 months of historical data to train the predictive models, along with fresh data every 4 to 12 hours.

The data should be uploaded as JSON or CSV files in the S3 bucket:

  • Each file can contain only the most recent data, in which case you should load each file separately.

  • Alternatively, a file can contain all data, including any recent data, allowing you to replace the existing file with each upload.

The goal is to maintain both fresh and historical data in the bucket at all times, not just the most recent data.

Q: How fast should the data be loaded?

Although it varies from customer to customer (depending on your setup), RGIP Scoring updates every few hours, so data should be uploaded within that timeframe. We highly recommend compressing files to speed up transfers.

Q: What is the volume of data to be loaded?

If you plan to send data in the range of multiple billions of records per month, we may need to trim that number to only the most relevant data to expedite the overall process of pulling, scoring, and pushing back scores to you.

Q: Is a historical upload of data required?

Yes, RGIP Scoring requires 9 months of historical data to train the predictive models.

Q: For which time period should the data be loaded (specific month, week, year, etc.)?

If sending events:

  • 1 record = 1 event with a timestamp and user ID/email.

  • The data should be sent as events with timestamps, with the oldest timestamp being 9 months back.

Q: Should the data be transferred on a periodic schedule?

Yes, we recommend transferring data at least once a day to provide RGIP with fresh data. Historical data can be loaded only once with a fixed timeframe of 9 months.

For specific use cases, data should be loaded only once, such as a historical data dump for a fixed timeframe. In other scenarios, we need to enable a continuous live sync to fetch new data as soon as it arrives.

Q: Can RGIP pull data directly from our Snowflake?

Yes, please refer to our Snowflake integration article.

Q: Can RGIP pull data directly from our BigQuery / GCS?

Yes, RGIP can pull data from BigQuery. Please see our BigQuery integration article.

RGIP does not integrate directly with GCS, but you can stream data from GCS to S3 using gsutil.

Q: Can HG Insights’ RGI Platform push data to an S3 bucket?

This use case is not supported at this moment.

Q: Can I upload several files per day instead of one?

Yes! We allow folder paths as the extract location, so as long as you have one folder per day of the month, our script will ingest any files found at that folder path.