What Is Salesforce Sandbox Seeding?
Sandbox seeding is the process of populating a Salesforce sandbox with records so that it contains realistic data to build and test against. Developer and Developer Pro sandboxes are created empty of records; Partial Copy sandboxes carry a sample; only a Full sandbox is a complete copy. Seeding fills the gap — either by generating synthetic records or by copying a subset of production. Masking, by contrast, changes the values of records that are already there so they no longer identify real people. Most regulated orgs need both, and the order matters.
Why sandboxes need seeding at all
Salesforce sandbox types differ in what they copy. A Developer sandbox copies metadata only, no records, with 200 MB of data storage (400 MB with the storage upgrade). A Developer Pro sandbox is the same shape with 1 GB (2 GB upgraded). A Partial Copy sandbox copies metadata and sample data up to 5 GB, and a sandbox template is required to choose which records come across. A Full sandbox copies metadata and all data, with the same storage limit as your production org.
That leaves every developer on a Developer sandbox with an org that has objects, fields, flows and Apex — and nothing in them. A test that needs an Account with ten Contacts, three open Opportunities and a Case history cannot run against an empty schema. Seeding is how that data gets there.
There are two ways to seed:
- Synthetic generation — a tool creates records from scratch to a shape you describe: this many Accounts, this many Contacts per Account, values drawn from libraries or generated to match field types. Nothing in the sandbox came from a real person.
- Production subset copy — a tool selects a slice of production records (by object, filter, or relationship depth) and copies them into the sandbox, with relationships intact. Everything in the sandbox came from a real person, unless it is masked on the way in.
The choice between them is the whole compliance question, and it is covered below.
Seeding and masking are different jobs
It is common to hear "seeding" and "masking" used as if they were alternatives. They are not. They answer different questions.
| Seeding | Masking | |
| Question it answers | Is there data to test against? | Does the data identify real people? |
| Acts on | Empty or thin sandboxes | Sandboxes that already hold production records |
| Produces | New records (synthetic or copied) | The same records with sensitive values replaced |
| Compliance role | None on its own; a copied subset is still personal data | The control that takes personal data out of the sandbox |
| Typical failure | Test data that does not look like production, so bugs hide | A masking rule that stops matching a field, so real data leaks |
A Full sandbox does not need seeding — it has everything — but it needs masking before anyone opens it, because everything in it is real. A Developer sandbox needs seeding, and if the seed is a copy of production it needs masking too. Only a synthetic seed avoids masking entirely, and synthetic data has its own cost: it looks like production only to the extent someone described production accurately.
Salesforce's own tooling reflects the pairing. Data Mask & Seed, the Core application that replaced the Data Mask managed package in the Summer '26 release, puts both functions in one app: masking policies for records already in a sandbox, and seeding to populate one. Both require the Salesforce Data Mask add-on licence.
The order to run them after a refresh
A sandbox refresh replaces the sandbox's contents with a fresh copy from production — and with it, every live email address, remote-site setting and scheduled job production had. The sequence after a refresh determines whether the sandbox is safe to open.
1. Mask first, on the Full or Partial Copy sandbox, before any user gets access. Every record in it is real until it is masked. Do not seed on top of unmasked data; it only makes the unmasked data harder to find.
2. Clean up what the refresh copied besides records. Email deliverability settings, remote-site settings that point at production endpoints, scheduled jobs, integration credentials. Masking the Email field does not stop a workflow from sending. This is the step most refresh checklists miss, and it is the cause of the sandbox-emails-real-customers incident.
3. Then seed, if the sandbox is thin. For Developer sandboxes, seed synthetic records or a masked subset. If seeding copies from production, the masking step applies to the copy — either the tool masks on the way in, or you mask the sandbox again after seeding.
4. Validate that named fields are masked. The deliverable is not "the jobs ran". It is "these fields on these objects are demonstrably masked" — checked, because a masking rule that silently stops matching a renamed field throws no error.
DataMasker treats steps 1 and 2 as one job: it masks in place, preserving record IDs and parent–child relationships, and auto-updates remote-site settings and custom labels and deletes files, case comments and tasks that carry production content. It is a masking tool; it does not generate synthetic records, which is why it sits alongside a seeding tool rather than replacing one.
When synthetic data is the wrong choice
Synthetic seeding is the cleanest answer to the compliance question — nothing came from a real person — and it is often the wrong answer to the testing question.
Production data has a shape that is hard to describe: the long tail of Accounts with one Contact and the handful with four hundred; the Opportunities stuck in a stage for three years; the picklist values nobody uses but that break a report when they appear; the custom field that holds a date in a text type because of a decision made in 2019. Bugs live in that shape. A synthetic seed reproduces it only as well as the person who configured the generator understood it, and generated data tends to be tidier than reality.
That is why UAT and integration testing usually run on masked copies of real data rather than on synthetic seeds: the shape is preserved because it is the real shape, with the identifying values swapped for plausible substitutes of the same format. Synthetic seeds are the right tool for unit tests, for demos, for training orgs, and for developers who need "an Account with some Contacts" rather than "our Accounts".
Choose by what the test needs. If it needs production's shape, mask a copy. If it needs any data at all, seed.
What to check in any seeding tool
Whether the tool is Salesforce's own or a third party's, the questions are the same.
- Does it preserve relationships? Seeded Contacts that do not belong to Accounts, or Opportunities without OpportunityLineItems, produce a sandbox that passes a record count and fails every real test.
- If it copies from production, where does masking happen? On the way in, or not at all? "We only copy a subset" is not a privacy control — a subset of personal data is personal data.
- Can it be run from a pipeline? Sandboxes created by CI/CD need seeding and masking as steps with logged results, not as reminders.
- What does it do with the rest of the refresh? Emails, endpoints, scheduled jobs. If the answer is "that's a separate checklist", plan for the checklist to be skipped once.
- What are its documented limits? Salesforce's page for Data Mask & Seed, for example, documents that above 20 million records a masking job no longer preserves data distributions or masks identical fields consistently across objects, and that a sandbox can run at most 12 masking jobs in 24 hours. Every tool has a page like that; read it before the licence, not after.
Key Takeaways
Sandbox seeding populates a sandbox with records — synthetic ones or a copied subset of production — because Developer and Developer Pro sandboxes are created with no records at all.
Seeding and masking answer different questions: is there data to test against, and does the data identify real people. A copied subset is still personal data until it is masked.
After a refresh, the order is: mask the real data, clean up what the refresh copied besides records, then seed thin sandboxes, then validate named fields — in that sequence.
Synthetic data avoids the compliance problem and often fails the testing problem: bugs live in production's shape, which a generator reproduces only as well as it was described.
Salesforce's Data Mask & Seed combines both functions in one licensed Core app; DataMasker is a masking tool that also handles post-refresh cleanup and runs from a pipeline, and pairs with a seeding tool rather than replacing it.
Frequently Asked Questions
Seeding puts records into a sandbox that has few or none — either generated synthetic records or a copied subset of production. Masking changes the values of records already in a sandbox so they no longer identify real people. Seeding solves the empty-sandbox problem; masking solves the sensitive-data problem. A copied seed needs masking; a synthetic seed does not.
Yes. Data Mask & Seed, a Core application introduced in the Summer '26 release, includes seeding alongside masking policies. It is available in Professional, Enterprise, Unlimited and Developer Editions with the Salesforce Data Mask add-on licence, and replaced the Data Mask managed package, whose support ended on 9 July 2026.
Only if it is synthetic. Records copied from production into a sandbox are personal data wherever they sit, and every GDPR obligation — lawful basis, minimisation, access rights, erasure — follows them. A copied seed must be masked or anonymised before the sandbox is opened to users.
Mask. A Full sandbox is a complete copy of production and needs no seeding; it needs every sensitive field masked before anyone gets access, and the refresh's copied email settings, remote sites and scheduled jobs cleaned up.
Mask the production copy first, before granting access. Clean up what the refresh copied besides data — email deliverability, remote-site settings, scheduled jobs. Then seed thin sandboxes, masking any copied subset. Finally validate that specific named fields are masked, because a rule that stops matching a field does not raise an error.
No. DataMasker is a masking tool: it masks real records in place, preserving IDs and relationships, and handles post-refresh cleanup as part of the same job. For empty Developer sandboxes it pairs with a seeding tool, and the seeded data can be masked by DataMasker if it was copied from production.
See how this works in your Salesforce org
30-minute demo tailored to your specific use case and data model.