This guide provides a 3-week approach to adding a new behavioral data source to an existing Engagement (behavioral) model, from the first connection through deployment. The point of the sequence is that your live scoring keeps running the whole time: you connect the source, map its events, and rebuild the model on a duplicate, then deploy once the new model has been compared to the live one.
Use this guide when you are adding a source alongside the ones already connected. Two neighbouring guides cover the other cases: Behavioral Source Migration when a new source replaces an existing one, and Behavioral model refresh when you are re-tuning a model without adding a source. If all you need is one extra event from a source that is already connected, see How to add a new event to your Engagement model? instead.
💡 Adding a source is additive. Your existing event mapping stays on, your live model keeps scoring, and batch processing does not need to be paused. Nothing changes for your Sales team until you deploy the new model in Week 3.
Prerequisites - Week 1
Confirm which behavioral integrations are already connected, and which models the new source should feed: lead-level, account-level, or both
Confirm the new source is supported and how it will send data to RGIP: see Integrations overview and What type of events can be used in a behavioral model?
Decide the source of truth for every event the new source brings. This is the step that is specific to adding a source, and the one that goes wrong most often
A single action (a form submission, a webinar registration, a campaign response) is often captured by several connected systems at once. Map it from more than one and the model counts it twice, so the score stops meaning what the weights say
Pick one source per event. A sensible default order is Salesforce campaigns, then your marketing automation platform, then product analytics, but the decision is yours: plenty of teams treat their marketing automation platform as the system of record for campaign response
Write the decision down per event before you map anything, and note which existing mappings you will need to drop as a result
Confirm which audience and conversion mappings should be used to measure the model, so the before and after are measured the same way
Record the current performance of the live model as your baseline, and spot check the deep dive reports for the patterns you hope the new source will fix (which false negatives do you expect it to catch?)
Internally align on the hand raiser events the model must always qualify (demo request, contact us, pricing page), including any that will now arrive from the new source
Establish the success criteria for the project before you start building
Connect the new data source - Week 1
Configure the integration for the new source
Check the Data Discovery page and confirm that data is arriving
Compare the event types and the volumes per event, per person, per month against what the source shows on its own side
Check whether the events you expected to be new are in fact copies of events you already receive from another system, and revisit the source of truth decisions if so
Check how much history the new source sends. A brand new connection usually starts from the day it was connected, and that matters for training
Training datasets are cohort snapshots taken at base dates 3 to 5 months in the past, each with a 90-day lookback. Events that only started arriving this week are absent from those windows, so the model has no statistics on them and cannot weight them
Where the source supports it, backfill the history: Segment replay, or a one-off historical load from a warehouse or S3 bucket. RGIP stores 9 months of behavioral data, so there is no value in sending more than that
If a backfill is not possible, plan for two passes: qualify the new events with business rules now (below), then re-run a behavioral model refresh in 3 to 4 months, once the events have enough history for the statistics to be meaningful
Map the new events - Week 2
Create the event mapping for the new connector. Each integration has its own mapping and its own shape: Segment, S3, Amplitude, Mixpanel, HubSpot, Marketo, Salesforce Campaigns
Reuse the meta event names already in use. If the new source brings a form fill and your existing mapping already has one, the same action should land under the same meta event name, with the same activity type. Different names for the same action split the statistics in two and make the weights harder to read
Leave out the events you decided against in Week 1, and turn off any existing mapping rule that the new source now owns. One source per event
Check the event mapping output report
Confirm the new events appear with the volumes you saw in Data Discovery, and that no event you already had has doubled
Check the unmapped events report for anything from the new source worth catching
Run through the event mapping sanity checklist before publishing
Pair review the mapping and the output report with a colleague before going further. Everything downstream inherits mapping mistakes, and they are much cheaper to catch here
Rebuild the dataset - Week 2
Duplicate the live behavioral model(s) and work on the duplicate. Never load a dataset on the live model
Reload the training and validation datasets on the duplicate, with the audience and conversion confirmed in Week 1
Training: three base dates, one per month, at today minus 3, 4 and 5 months
Validation: one base date at today minus 2 months
Base dates cannot go much beyond 6 months back, because the oldest one still needs its full 90-day lookback inside the 9 months of behavioral data that RGIP stores
Confirm the new events actually made it into the model: how to make newly mapped events appear in the Event Weights tab. An event that is missing here was either mapped after the dataset was loaded, or has no occurrences inside the lookback windows
Handle the high-frequency events the new source brings with aggregations rather than per-occurrence scoring
Create the aggregation, then run the frequency analysis to find the occurrence counts worth scoring, then create the aggregated events at those thresholds
Event names in the aggregation must match the event mapping output names exactly. If they do not, the aggregation silently returns nothing, with no error
Aggregated events only appear in the model after you reload the dataset again, so create them before the weighting pass rather than after
Assign the weights - Week 2
Start from the statistically suggested weights, then apply your business rules on top. The suggested importance is lift based and lands in a narrow 1 to 9 range; the scale you set is unbounded and only the ratios matter, so widen the spread.
Qualify the hand raisers by business rule. Demo requests and contact sales events get the top weights, in the 30 to 50 range, so that a single one of them puts the person in the very high segment on its own. This is the rule that makes the model behave the way Sales expects, and it is the one place where the business decision outranks the statistics
Set a weight of 0 for the events that would otherwise qualify people cheaply, which is where a new source most often causes trouble, because a fresh product analytics or web feed arrives with enormous volume
High frequency, low intent events: email opens, generic blog and web page views. Score them through an aggregated event ("at least 10 in the last 30 days") and leave the single occurrence at 0, so nobody qualifies for doing one small thing many times
Sales activity: responses to your own outreach. Scoring people who are already in conversation with Sales defeats the purpose of a model whose job is to surface organic intent
Keep negative weights for true anti-signals only: unsubscribe, email bounce, spam complaint, opt-out. Everything that is merely not a positive signal belongs at 0, not below it. A positive lift on a bounce is an artifact of active people receiving more email, not a buying signal
Leave the new source's thinly attested events neutral. An event with a handful of occurrences in the training data has a lift that is noise, whichever direction it points. Park it at a low weight until it has history
Leave the suggested weights in place for everything else
Set the lifespans. Weights decay linearly over the lifespan, up to 90 days. Do not go below about 5 days, or scores flip-flop as events expire; high frequency events can take a shorter lifespan so they do not accumulate noise
Adjust the thresholds - Week 2
Re-derive the segment thresholds. Do not carry the live model's over. New events and a wider weight spread change the raw score distribution, so the old very high / high / medium / low cutoffs land somewhere different
Check the hand raiser rule holds: score a person with a single demo request and confirm they land in very high
Check the volume in each segment against what your Sales team can work. A new behavioral source usually raises scores across the board, and a model that suddenly triples the very high population breaks the SLA even when the ranking is better
Review the conversion rate of each segment on the validation dataset. Segments should separate cleanly, with medium sitting close to the population baseline
Compare to the live model - Week 2&3
Compare the new model against the live one on the same validation dataset, using the audience and conversion you fixed in Week 1
Pay close attention to the Sample page, and review it with a pair. Read the actual people: does the new source's data explain why each one moved, and does the explanation hold up?
Look specifically at the movers. The people who changed segment because of the new source are the whole point of the project, and they are where a double counted event or a too-generous weight shows up first
Check the Signals a rep will see on a scored person, so the new events read sensibly in the CRM
Share a handful of examples with Sales and collect their feedback
Iterate on the weights and thresholds, then review a fresh sample. Expect to go round this loop more than once
Go Live - Week 3
Publish the new event mapping if it is not published yet
Deploy the new behavioral model(s)
Consider deploying on a Friday afternoon, so any surprise lands outside the working week
Monitor performance for at least one week after deployment, and compare the segment volumes against the baseline you recorded in Week 1
Keep an eye on the new source's volumes in Data Discovery for the first month. A feed that stops sending is not obvious from the scores: the model simply stops seeing those events and the people who relied on them quietly drop down the ranking
Gather feedback from the Sales team on the leads they are now receiving
FAQ
Do I need to turn off my existing behavioral integration?
No. Adding a source is additive: the existing integration and its event mapping keep running. The only rules you turn off are the ones for events that the new source now owns, so that a single action is not counted twice. If the new source is meant to replace an existing one entirely, follow Behavioral Source Migration instead.
Will scoring stop while I add the new source?
No. You build the new model on a duplicate of the live one, so the live model keeps scoring throughout and nothing changes for Sales until you deploy in Week 3. If you also need historical events backfilled from the new source, get in touch with the RGIP team to plan it.
The same event arrives from two systems. Which one should I map?
Only one of them, and which one is your call. A reasonable default is Salesforce campaigns first, then your marketing automation platform, then product analytics, but if your marketing automation platform is your system of record for campaign response, map that. What matters is that exactly one source is mapped per event: missing one system's copy costs nothing, because you still have the event once, while mapping both doubles its weight and distorts the ranking.
The new source has no history. Can I still use its events?
Yes, but through business rules rather than statistics. Training runs on base dates 3 to 5 months in the past, so events that started arriving this week are not in the training data and have no measured lift. Weight the obvious hand raisers by hand, keep the rest neutral, and plan a behavioral model refresh in 3 to 4 months once the events have accumulated history. Where the source supports a historical backfill, doing that first is the better option.
My new source sends a huge volume of page views. How do I keep it from swamping the score?
Set the single occurrence event to a weight of 0 and score the behavior through an aggregated event instead, at a threshold such as "at least 10 in the last 30 days". The person then scores for the pattern once it is real, rather than accumulating points for repeating one small action. Generic page views and email clicks are aggregation candidates on almost every build.