An apex test data factory is a class that builds records in memory for unit tests, and it works fine until you need volume, cross-object relationships, or repeatable data outside the test context. Bulk API v2 solves a different problem: getting thousands or millions of schema-compliant records into an org for performance testing, UAT, or sandbox seeding. Most teams pick one and try to force it to do the other's job, which is where the real time gets lost.

What an Apex Test Data Factory Is Built to Do

A factory class is a wrapper around SObject construction. You call TestDataFactory.createAccount() or something similar, it returns a populated Account, and you insert it inside a test method. The value is consistency: every developer on the team builds an Account the same way, with the same required fields set, instead of copy-pasting insert statements across forty test classes.

Done well, a factory class also handles relationships. It creates a Contact, links it to an Account, maybe attaches an Opportunity, and returns the whole graph so the test method does not have to think about referential order. That is genuinely useful for isolating logic in a trigger or a Flow, because you control exactly what data exists and nothing else.

The catch is scope. Factory classes are written for the unit test context: @isTest annotations, System.runAs blocks, and a maximum of 10,000 DML rows per test method under Salesforce's governor limits. They are not designed to populate a full sandbox or simulate a production-scale data volume test. Nobody wrote them for that job, and that is fine, as long as nobody tries to stretch them into it.

Where Factory Classes Hit a Wall

The first wall is volume. If your data skew testing needs 500,000 Account records with a realistic ownership distribution across fifteen sales teams, an Apex test factory running inside a single transaction will not get you there. You will hit the 10,000 DML row governor limit, and even if you split the load across multiple test methods, you are still bound by test execution time limits and CPU time per transaction.

The second wall is schema drift. Factory classes hardcode field names and values at the time they are written. Add a required custom field to Opportunity six months later, and every factory method that creates an Opportunity needs a manual update, or your entire test suite starts failing with REQUIRED_FIELD_MISSING errors. Teams that own dozens of factory classes spend real hours every release just patching them to match schema changes nobody told QA about.

The third wall is realism. A factory class usually assigns the same three or four values to a picklist field because that is what the developer typed in 2021. It will not reflect actual data distribution, will not exercise edge cases in validation rules tied to less common values, and will not surface the kind of record-type or page-layout-specific field requirements that only show up when you generate data across the full schema rather than a hand-picked subset.

What Changes When You Move to Bulk API v2

Bulk API v2 is not a testing framework, it is a data-loading protocol. Salesforce chunks your records automatically, manages job state, and reports success or failure per batch. That makes it the right tool for moving large volumes of data into or out of an org without babysitting batch sizes yourself the way you had to with Bulk API 1.0.

The advantage for test data specifically is that Bulk API v2 loads run outside the Apex test context. There is no 10,000 row DML ceiling, no per-transaction CPU limit tied to a test method, and no requirement to wrap everything in @isTest. You can generate a million-record Account hierarchy, load it into a full sandbox overnight, and have it ready for a load test the next morning.

SproutEzee reads the org's actual schema before generating anything, so field lengths, picklist values, required fields, and record-type-specific layouts all get respected automatically. That solves the schema drift problem that plagues hand-written factory classes: when a field changes, the generated data changes with it, because the generation logic reads the schema fresh every run instead of relying on a class someone wrote two years ago.

Choosing Between Factory Classes and Bulk Loads

The decision usually comes down to what you are actually testing. Unit-level logic checks belong in Apex test methods with factory-built data. Anything involving volume, integration, or a realistic full-org data set belongs in a Bulk API v2 load.

ScenarioBetter Fit
Testing a single trigger's field validationApex test data factory
Seeding a full sandbox for UATBulk API v2
Verifying a Flow's decision logic on one recordApex test data factory
Load testing report performance at 1M+ recordsBulk API v2
Regression testing after a schema changeBulk API v2 with schema-aware generation

Notice the pattern: factory classes win when the test is narrow and the assertion is about behavior. Bulk loads win when the test is about scale, integration, or how the org behaves once real volume and realistic variety enter the picture. Teams that only have factory classes tend to discover data volume bugs in production, because nobody ever tested with more than a few hundred records.

Keeping Both in Sync With Schema Changes

The failure mode that catches most teams is not choosing the wrong tool, it is letting the two drift apart. A factory class creates an Opportunity with five fields set. A Bulk API v2 load, generated from a schema-reading tool, populates thirty fields because it reads every required and commonly-used field on the object. Run both against the same validation rule, and you can get different pass/fail results depending on which path created the record.

The fix is not to abandon factory classes, since they still earn their keep for fast, isolated unit tests. The fix is to stop treating them as your only source of truth for what a valid record looks like. Schema-aware generation, whether through SproutEzee or a similar tool, should be the reference for what a complete, valid record actually requires, and factory classes should be updated against that reference rather than against whatever the last developer remembered.

In practice this means running periodic audits: pull the schema, compare required fields against what your top ten factory methods actually set, and flag gaps. It is not glamorous work, but it is a lot cheaper than debugging a production incident that traces back to a test suite that was quietly out of date for eight months.

Frequently Asked Questions

Can Bulk API v2 replace Apex test data factories entirely?

No, they serve different purposes. Apex test data factories run inside unit tests and are built for fast, isolated assertions about code behavior. Bulk API v2 is a data-loading protocol meant for volume and does not run inside the Apex test context, so it cannot be asserted against directly in a test method the way factory-built data can.

Why do Apex test data factory classes fail after a schema change?

Factory classes hardcode field names and values at the time they are written, so they do not automatically pick up new required fields or validation rule changes. When a schema change adds a required field, every factory method creating that object needs a manual update or it will throw a required field missing error. Schema-aware generation tools avoid this because they read the current schema on every run.

How many records can an Apex test method insert before hitting a governor limit?

A single test method transaction is capped at 10,000 DML rows under Salesforce governor limits. This makes Apex test factories impractical for volume testing scenarios that require hundreds of thousands of records. Bulk API v2 has no such per-transaction cap because it runs outside the test execution context.

Does SproutEzee generate data that respects record types and page layouts?

Yes. SproutEzee reads the org's schema directly, including record-type-specific required fields and picklist value sets, before generating data through Bulk API v2. This means generated records reflect the same field requirements a user would see on the actual page layout, rather than a generic subset a developer chose manually.

Should test data factory classes be retired once a team adopts Bulk API v2 loads?

Not entirely. Factory classes still have a clear role in fast unit tests that check specific logic paths without needing volume. The better approach is to use schema-aware generation as the reference for what a complete valid record requires, and periodically audit factory classes against that reference so the two do not drift apart over time.