Skip to main content
Guides›Sandbox Security5 min

What Is Salesforce Sandbox Seeding?

Sandbox seeding is the process of populating a Salesforce sandbox with records so that it contains realistic data to build and test against. Developer and Developer Pro sandboxes are created empty of records; Partial Copy sandboxes carry a sample; only a Full sandbox is a complete copy. Seeding fills the gap — either by generating synthetic records or by copying a subset of production. Masking, by contrast, changes the values of records that are already there so they no longer identify real people. Most regulated orgs need both, and the order matters.

Why sandboxes need seeding at all

Salesforce sandbox types differ in what they copy. A Developer sandbox copies metadata only, no records, with 200 MB of data storage (400 MB with the storage upgrade). A Developer Pro sandbox is the same shape with 1 GB (2 GB upgraded). A Partial Copy sandbox copies metadata and sample data up to 5 GB, and a sandbox template is required to choose which records come across. A Full sandbox copies metadata and all data, with the same storage limit as your production org.

That leaves every developer on a Developer sandbox with an org that has objects, fields, flows and Apex — and nothing in them. A test that needs an Account with ten Contacts, three open Opportunities and a Case history cannot run against an empty schema. Seeding is how that data gets there.

There are two ways to seed:

  • Synthetic generation — a tool creates records from scratch to a shape you describe: this many Accounts, this many Contacts per Account, values drawn from libraries or generated to match field types. Nothing in the sandbox came from a real person.
  • Production subset copy — a tool selects a slice of production records (by object, filter, or relationship depth) and copies them into the sandbox, with relationships intact. Everything in the sandbox came from a real person, unless it is masked on the way in.

The choice between them is the whole compliance question, and it is covered below.

Seeding and masking are different jobs

It is common to hear "seeding" and "masking" used as if they were alternatives. They are not. They answer different questions.

SeedingMasking
Question it answersIs there data to test against?Does the data identify real people?
Acts onEmpty or thin sandboxesSandboxes that already hold production records
ProducesNew records (synthetic or copied)The same records with sensitive values replaced
Compliance roleNone on its own; a copied subset is still personal dataThe control that takes personal data out of the sandbox
Typical failureTest data that does not look like production, so bugs hideA masking rule that stops matching a field, so real data leaks

A Full sandbox does not need seeding — it has everything — but it needs masking before anyone opens it, because everything in it is real. A Developer sandbox needs seeding, and if the seed is a copy of production it needs masking too. Only a synthetic seed avoids masking entirely, and synthetic data has its own cost: it looks like production only to the extent someone described production accurately.

Salesforce's own tooling reflects the pairing. Data Mask & Seed, the Core application that replaced the Data Mask managed package in the Summer '26 release, puts both functions in one app: masking policies for records already in a sandbox, and seeding to populate one. Both require the Salesforce Data Mask add-on licence.

The order to run them after a refresh

A sandbox refresh replaces the sandbox's contents with a fresh copy from production — and with it, every live email address, remote-site setting and scheduled job production had. The sequence after a refresh determines whether the sandbox is safe to open.

1. Mask first, on the Full or Partial Copy sandbox, before any user gets access. Every record in it is real until it is masked. Do not seed on top of unmasked data; it only makes the unmasked data harder to find.

2. Clean up what the refresh copied besides records. Email deliverability settings, remote-site settings that point at production endpoints, scheduled jobs, integration credentials. Masking the Email field does not stop a workflow from sending. This is the step most refresh checklists miss, and it is the cause of the sandbox-emails-real-customers incident.

3. Then seed, if the sandbox is thin. For Developer sandboxes, seed synthetic records or a masked subset. If seeding copies from production, the masking step applies to the copy — either the tool masks on the way in, or you mask the sandbox again after seeding.

4. Validate that named fields are masked. The deliverable is not "the jobs ran". It is "these fields on these objects are demonstrably masked" — checked, because a masking rule that silently stops matching a renamed field throws no error.

DataMasker treats steps 1 and 2 as one job: it masks in place, preserving record IDs and parent–child relationships, and auto-updates remote-site settings and custom labels and deletes files, case comments and tasks that carry production content. It is a masking tool; it does not generate synthetic records, which is why it sits alongside a seeding tool rather than replacing one.

When synthetic data is the wrong choice

Synthetic seeding is the cleanest answer to the compliance question — nothing came from a real person — and it is often the wrong answer to the testing question.

Production data has a shape that is hard to describe: the long tail of Accounts with one Contact and the handful with four hundred; the Opportunities stuck in a stage for three years; the picklist values nobody uses but that break a report when they appear; the custom field that holds a date in a text type because of a decision made in 2019. Bugs live in that shape. A synthetic seed reproduces it only as well as the person who configured the generator understood it, and generated data tends to be tidier than reality.

That is why UAT and integration testing usually run on masked copies of real data rather than on synthetic seeds: the shape is preserved because it is the real shape, with the identifying values swapped for plausible substitutes of the same format. Synthetic seeds are the right tool for unit tests, for demos, for training orgs, and for developers who need "an Account with some Contacts" rather than "our Accounts".

Choose by what the test needs. If it needs production's shape, mask a copy. If it needs any data at all, seed.

What to check in any seeding tool

Whether the tool is Salesforce's own or a third party's, the questions are the same.

  • Does it preserve relationships? Seeded Contacts that do not belong to Accounts, or Opportunities without OpportunityLineItems, produce a sandbox that passes a record count and fails every real test.
  • If it copies from production, where does masking happen? On the way in, or not at all? "We only copy a subset" is not a privacy control — a subset of personal data is personal data.
  • Can it be run from a pipeline? Sandboxes created by CI/CD need seeding and masking as steps with logged results, not as reminders.
  • What does it do with the rest of the refresh? Emails, endpoints, scheduled jobs. If the answer is "that's a separate checklist", plan for the checklist to be skipped once.
  • What are its documented limits? Salesforce's page for Data Mask & Seed, for example, documents that above 20 million records a masking job no longer preserves data distributions or masks identical fields consistently across objects, and that a sandbox can run at most 12 masking jobs in 24 hours. Every tool has a page like that; read it before the licence, not after.

Key Takeaways

Sandbox seeding populates a sandbox with records — synthetic ones or a copied subset of production — because Developer and Developer Pro sandboxes are created with no records at all.

Seeding and masking answer different questions: is there data to test against, and does the data identify real people. A copied subset is still personal data until it is masked.

After a refresh, the order is: mask the real data, clean up what the refresh copied besides records, then seed thin sandboxes, then validate named fields — in that sequence.

Synthetic data avoids the compliance problem and often fails the testing problem: bugs live in production's shape, which a generator reproduces only as well as it was described.

Salesforce's Data Mask & Seed combines both functions in one licensed Core app; DataMasker is a masking tool that also handles post-refresh cleanup and runs from a pipeline, and pairs with a seeding tool rather than replacing it.

Frequently Asked Questions

See how this works in your Salesforce org

30-minute demo tailored to your specific use case and data model.