An empty provider balance does not need to stop product engineering. Much of a delivery system is ordinary software: state transitions, authentication boundaries, persistence, error handling, user interfaces, and accounting. Test those parts without calling a live model.
Use deterministic provider fixtures
A fake provider can return valid output, malformed output, a transient error, or a permanent failure on demand. Assert the exact attempt count and final state. Keep the fixture explicit so nobody mistakes the result for a demonstration of real model quality.
Local HTTP servers add another layer: they can exercise request serialization, response parsing, timeouts, and heartbeat behavior. A desktop start/pause test can use a fake control plane that reports no eligible tasks. That proves process lifecycle behavior without pretending to complete a real repository change.
Exercise persistence with real local components
Use a disposable database to test enrollment exchange, one-time proofs, replay rejection, and migrations. Use a temporary OS key-store entry with a fake credential to test storage and rotation. Never point destructive test cleanup at a user's active state directory.
Installer tests should also preserve fake connection settings through upgrades and uninstall/reinstall. A successful package build is not proof that an installed launcher can find its runtime or native credential library. Test the installed artifact, not only the source checkout.
Test the interface people actually use
Reopen the app after saving settings. Is a blank password field explained? Can the user tell whether a connection is restored? What happens when login returns to the wrong page? These are ordinary usability defects that do not require paid inference to find.
For ForgeLoop's desktop preview, prerequisite checks, a saved-connection heartbeat, and local key-store verification do not invoke a model. Start runner is different: it can claim eligible work and is not a free test mode.
Keep the final boundary honest
Offline tests cannot establish that a model follows the instructions, that a provider accepts the real key, or that the account has credit. Record those as separate live validation steps. A useful release checklist says both what passed and what remains unverified. That distinction lets engineering continue without turning a simulated result into a production claim.