To test Shopify Functions locally, run three layers. Vitest unit tests call your exported JavaScript function with plain input JSON. Fixture tests compile the function to WebAssembly and run it through Shopify’s function-runner, checking input and output against the API schema. Then shopify app function replay re-runs a real checkout’s captured input against your local build.
Each layer catches a different class of bug. Unit tests on their own miss the ones that come from the contract between your code and Shopify. Everything below was checked against Function API version 2026-10 (stable since 1 October 2026), Shopify CLI 4.8 and @shopify/shopify-function-test-helpers 1.1.0, as of October 2026.
Why isn’t a Vitest unit test enough?
Because your unit test runs your JavaScript in Node, and production does not. Shopify compiles a JavaScript function to WebAssembly and runs it on a JavaScript engine inside that module, under a hard budget of 11 million instructions for carts of up to 200 lines. The test and debug guide says it directly: “Unit tests don’t run in WebAssembly, which in rare cases might cause different results than what’s found in production.”
There is also a duller failure. Your input query asks for one shape, your hand-written test input has another, and nothing checks that they agree. The function passes every test and then reads undefined in checkout because the field your test invented was never in the query.
How do you unit test a Shopify Function with Vitest?
Import the exported function and pass it an object shaped like your input query’s result. Shopify recommends Vitest for JavaScript and TypeScript functions, and for Rust the shopify_function crate provides run_function_with_input for cargo test.
Here is a small cart validation that stops checkout when any line has more than five units. The input query, in src/cart_validations_generate_run.graphql:
query CartValidationsGenerateRunInput {
buyerJourney {
step
}
cart {
lines {
quantity
}
}
}
The function, in src/cart_validations_generate_run.js:
// @ts-check
const MAX_PER_LINE = 5;
/**
* @param {import("../generated/api").CartValidationsGenerateRunInput} input
* @returns {import("../generated/api").CartValidationsGenerateRunResult}
*/
export function cartValidationsGenerateRun(input) {
// Only block at checkout; let people edit the cart freely.
if (input.buyerJourney.step === "CART_INTERACTION") {
return { operations: [] };
}
const tooMany = input.cart.lines.some((line) => line.quantity > MAX_PER_LINE);
if (!tooMany) {
return { operations: [] };
}
return {
operations: [
{
validationAdd: {
errors: [
{
message: `You can buy up to ${MAX_PER_LINE} of each item.`,
target: "$.cart",
},
],
},
},
],
};
}
And the unit test next to it:
import { describe, it, expect } from "vitest";
import { cartValidationsGenerateRun } from "./cart_validations_generate_run";
const cart = (step, ...quantities) => ({
buyerJourney: { step },
cart: { lines: quantities.map((quantity) => ({ quantity })) },
});
describe("cartValidationsGenerateRun", () => {
it("ignores the cart page", () => {
expect(cartValidationsGenerateRun(cart("CART_INTERACTION", 9))).toEqual({
operations: [],
});
});
it("passes at the limit", () => {
expect(cartValidationsGenerateRun(cart("CHECKOUT_COMPLETION", 5, 1))).toEqual({
operations: [],
});
});
it("blocks one over the limit", () => {
const result = cartValidationsGenerateRun(cart("CHECKOUT_COMPLETION", 1, 6));
expect(result.operations[0].validationAdd.errors[0].target).toBe("$.cart");
});
it("treats a null step as checkout", () => {
expect(cartValidationsGenerateRun(cart(null, 6)).operations).toHaveLength(1);
});
});
The last case earns its place. In the 2026-10 schema BuyerJourney.step is nullable, so a test that only ever passes one of the three enum values is testing a narrower world than checkout sends. The rule behind the function, and the checkout surfaces it does and does not run on, are in our cart validation function guide.
These tests run in milliseconds and you can step through them in a debugger, which is the reason to keep them once the fixture tests exist.
What does a fixture test check that a unit test cannot?
Three things: that the input query is valid against the function’s schema, that the fixture’s input matches what that query would return, and that the expected output is a valid result for the target. Then it builds the real Wasm module and compares its output with the fixture.
The generated function templates ship this already, as tests/default.test.js plus tests/fixtures/. A fixture is one JSON file:
{
"payload": {
"export": "cart_validations_generate_run",
"target": "cart.validations.generate.run",
"input": {
"buyerJourney": { "step": "CHECKOUT_COMPLETION" },
"cart": { "lines": [{ "quantity": 1 }, { "quantity": 6 }] }
},
"output": {
"operations": [
{
"validationAdd": {
"errors": [
{
"message": "You can buy up to 5 of each item.",
"target": "$.cart"
}
]
}
}
]
}
}
}
The test file leans on the @shopify/shopify-function-test-helpers package. buildFunction and getFunctionInfo shell out to shopify app function build and shopify app function info --json; the second returns the schema path, the Wasm path, the function-runner binary and each target’s input query. validateTestAssets returns three error arrays, and runFunction pipes the fixture input to function-runner and returns the output with an instructionCount, memoryUsageKiB and moduleSizeKiB.
The errors are specific. Add a field to the fixture that the query never selects and you get Extra field `cost` found in fixture data not in query. Drop the target from an expected error and you get Field "target" of required type "String!" was not provided. Both of those are bugs a Node-only unit test would have passed without comment.
If getFunctionInfo fails with The "shopify app function info" command is not available in your CLI version, upgrade the CLI. The helpers package also wants Node 20 or later.
How do you test Shopify Functions locally with real checkout input?
Capture it from real runs. When shopify app dev is running and you trigger the function on your dev store, the CLI streams each execution to your terminal and writes the input and output to .shopify/logs. Shopify’s guide says those log files can be dragged straight into tests/fixtures/ and used as fixtures.
To re-run one of those executions against your current code:
cd extensions/quantity-limit
shopify app function replay --log 9f1f0e --watch
The six-character identifier comes from the app dev output, and omitting --log lets you pick from a list of recent runs. --watch re-runs the function whenever the source changes, which makes replay the fastest loop for fixing a bug someone found by actually checking out.
For a one-off input you wrote yourself, shopify app function run --input input.json --export cart_validations_generate_run executes the built module once. Add --json for machine-readable output.
Read a captured log before you commit it as a fixture. It contains whatever your input query asked for, and if that includes a buyer’s email or address, so does your repo.
How do you catch instruction-count regressions?
Record a baseline and fail the build when a fixture gets more expensive. Instruction counts are deterministic, so the same module and the same input produce the same count every time. Shopify’s instruction count guide extends the default test to store counts in tests/instruction-count-baseline.json and assert each run stays within 5% of it. Create or refresh the baseline with:
UPDATE_INSTRUCTION_BASELINE=1 npm test
A function that crosses the limit in production is stopped with InstructionCountLimitExceededError, and above 200 cart lines the limit scales with the line count, per the Function API resource limits. The baseline test does not prove you are under that limit on a huge cart. It tells you when a refactor has made every run more expensive, before a merchant with a large cart finds out. When a number jumps, shopify app function run --input input.json --profile writes a profile you can open in Speedscope. The input and output caps (128 kB and 20 kB up to 200 lines) are in our discount function migration post, along with the query size rules.
What should run in CI?
Both test layers, on every pull request that touches an extension. Unit tests need nothing but Node. The fixture tests need Shopify CLI on the runner, because the helpers call the shopify binary to build the module and locate function-runner. Commit the fixtures and the baseline file; regenerate the baseline deliberately, in its own commit, so a reviewer can see the cost change.
Replay stays local. It depends on logs from your own app dev session, and its job is turning a real bug report into a fixture.
What we would do
Start with the fixture tests, not the unit tests. Turn on app dev, check out three or four real carts on the dev store (a single-item cart, a big B2B cart, a guest with no customer), drop those logs into tests/fixtures/, and record the baseline. That gives you schema-checked coverage of the inputs checkout really sends. Add Vitest tests afterwards for the edge cases you cannot easily produce on a store, like a null journey step or a quantity exactly at the limit.
Skip the instruction baseline for a function that does two comparisons and never loops over lines. It will sit at a tiny fraction of the budget and the baseline file becomes churn. Anything that iterates over cart lines, parses a metafield, or runs in JavaScript on carts that can reach hundreds of lines should have one from the first commit.
We write and maintain Functions as part of our Shopify development work, including the test harness and CI that keeps a checkout rule from quietly breaking on the next API version.