Choosing between seasonal staff and AI is not one decision. It is three, because your peak inbox holds three kinds of message, and only one of them is a hiring problem.
The hiring market got thinner, and it has not recovered
For the 2025 holidays, US retailers planned between 265,000 and 365,000 seasonal hires. The year before, they hired 442,000. That is the lowest level in at least 15 years, according to NRF figures reported by CNBC and CBS News. Challenger, Gray & Christmas counts it a different way, tracking Q4 job gains rather than planned hires, and lands in the same place: the smallest seasonal gain in 16 years.

Two caveats before you use that number.
It is last season's figure. NRF and Challenger publish their outlook in the autumn, and at the time of writing the current season's forecast is not out. So this is the most recent reading available, not a forecast for the peak you are planning now. Treat it as the trend line rather than the weather.
And it does not prove AI replaced those jobs. Nobody has established that. What it does tell you is narrower and more useful: the pool of seasonal help you were planning to hire from has been shrinking, and more stores are reaching into it. Expect the practical version of that. Fewer candidates, later in the season, and more competition for the good ones.
So the honest answer to what saves more is this: whichever lane matches what your inbox is actually full of. Two of your three question piles are not hiring problems, and paying people to work them is the expensive way round. Counting them is the part only you can do. Here is how.
What does one more person actually buy you?
A seasonal hire costs roughly $7,000 to $11,000 for the season, fully loaded. That range comes from Ringly's peak staffing playbook, a vendor estimate rather than a payroll study, so treat it as a starting bracket and replace it with your own wage plus tax plus training cost as soon as you can.
Before you post the job, walk last November's inbox and ask one question of every kind of message that showed up:
Does a second person make this answer better, faster, or neither?
Three answers come back.
"Where is my order?" Faster, not better. A person opens the same screen the AI opens. They read the same tracking number back to the shopper, in 40 to 90 seconds instead of one, and two staff often word it differently. You bought throughput. You did not buy a better answer.
"Which of these two should I get?" Better, but only if they know the catalog. A seasonal hire two weeks into the job does not know the catalog. Neither does an AI that was never given the comparison. This is the answer nobody has, and hiring does not create it.
"This arrived broken and I am furious." Better, every time. Somebody has to read the tone, weigh what this customer is worth, and decide to bend a rule. That is not a lookup. It is a judgment call, and it is the only kind of message where an extra person reliably returns more than you paid.

Read that honestly in both directions. If judgment calls are a big share of your peak, hiring is the right spend and you should start earlier than you planned, because good seasonal candidates go early in a thin year. If lookups dominate, you are about to spend $7,000 or more on speed alone.
The break-even nobody runs
Turn the hire into a per-message number, because that is the only way the two lanes compare.
Take the middle of the bracket, $9,000, and divide it by the messages that person will actually answer. A seasonal rep working the peak weeks at 40 to 90 seconds per lookup, with the rest of a shift going to the harder piles, realistically closes somewhere around 1,500 to 2,500 messages across the season.
That puts a human answer at roughly $3.60 to $6.00 per message, and the arithmetic is unkind to the lookup pile in particular, because a lookup is the cheapest thing a person can do and you are paying the same rate for it.
Now price the same volume the other way. Per-conversation AI pricing across the category tends to land in the low single digits of dollars, and monthly plans with a floor work out lower still once volume climbs, which is exactly the direction peak goes. The comparison is not close on lookups. It is not the right comparison at all on judgment calls, where a $6 answer that keeps a customer is cheap.

Run your own version with your own wage. If the number comes out under $2, your local labour market beats the bracket above and you should weight this article accordingly.
A worked example
A composite of stores we have looked at, not a single customer. The figures are illustrative.
A home goods store did the sort on 214 messages from its last peak week. The split came back 61% lookups, 22% advice, 17% judgment. It had budgeted two seasonal hires at about $18,000.
The 17% judgment pile worked out to roughly 36 messages across the sample week. Two people were not needed for that. One was, comfortably.
What the sort actually exposed was the 22% advice pile. Those were product-fit questions with no written answer anywhere, so neither a new hire nor an AI could have handled them. The store had been planning to solve them by adding a person who would have needed the same missing information.
They hired one instead of two, moved part of the saved budget to writing comparison and sizing content in October, and automated the lookups. The honest caveat: this is one season and one store, and we cannot show you a controlled version where they did it the other way.
And throughput is not worthless. If you have no automation at all and volume triples, a second person genuinely helps. It is just the most expensive way to buy it, and it stops the day you stop paying.
Seasonal hire vs AI, compared on what actually differs
Four dimensions separate the two lanes. Cost is the one everybody stares at, and it is the least decisive of the four.
| Seasonal hire | AI | |
|---|---|---|
| Cost | $7,000 to $11,000 for the season | Per conversation, or a monthly plan with a floor |
| Time until useful | 2 to 3 weeks of ramp | Days, if the content already exists |
| At 3x to 4x volume | Escalations rise from 18-24% to 26-38% | Throughput holds |
| Quality at that volume | Handle time up 12% to 22%, answers drift apart | Flat, at whatever level the content sets |
| What slips under load | Consistency between staff | Anything nobody wrote down |
Sources: seasonal hire cost from Ringly's peak staffing playbook, a vendor estimate. Escalation and handle time attributed to Gartner and Zendesk, reported second-hand by Stealth Agents. We have not seen the underlying Gartner and Zendesk material, so read those two figures as directional.
The table does not name a winner, because a winner only exists once you know which pile the volume lands in. A store whose peak is 40% judgment calls reads this table in the opposite direction to a store whose peak is 40% product questions.
Three more numbers sit outside the table, and they bend the result harder than anything in it.
Three numbers the cost comparison leaves out
More than half the demand lands outside a standard working day. During BFCM 2025, 53% of conversations arrived outside 9am to 6pm UTC (Chatty conversation data, Nov 24 to Dec 1, 2025).
That is a global base across many time zones, so read it as shape rather than as a clock reading for your own store. The shape holds either way. Over half of the volume falls outside one standard business day, and hiring more daytime coverage does not reach it. That is a coverage problem rather than a cost problem, so a bigger budget does not touch it. Closing it means some form of 24/7 customer support, which is a different purchase from more headcount.
The second number is about timing. A new support rep needs 2 to 3 weeks before they are useful (Ringly, vendor estimate). Peak itself lasts about four weeks. By the time the person you hired is good at the job, the job is nearly over.

The third is what happens after. US retail added 494,000 jobs between October and December 2023, then cut 464,000 in January and February 2024, per BLS figures compiled by Netchex. Those are the most recent full build-and-teardown cycles published in that compilation, and the pattern is structural rather than particular to that year. Every seasonal build gets dismantled six weeks later, usually right as returns arrive.

Now argue against these numbers
First, the soft edges in the case against hiring
The three numbers above are the case against hiring, and each one has a soft edge.
The 53% is measured across stores in many time zones. A store selling into one country, whose shoppers mostly sleep when its staff sleep, will find a smaller share of its own volume lands after hours. Check yours before you lean on the number.
The ramp figure assumes you are hiring strangers. Rehire two people who worked your last peak and the two to three weeks mostly disappears, which is the strongest argument for hiring in this article.
And the January teardown only counts as waste if you wanted to keep them. Seasonal work ending in January is the arrangement, not a defect in it.
Then the case against AI, which comes off worse
Start with resolution rate. The median across 195 deployments and 38 vendors is 70%, with half the field between 56% and 80%, according to My AskAI's benchmark. That is below the 80% to 90% quoted in most sales decks, including the ones written by companies like ours.
Two things about that 70% matter more than the number. These are published rates, so vendors put forward their wins rather than their average customer, which puts the true field average somewhere below the median. And the full range runs from 15% to 98.3%.
A spread that wide is the real finding. It means "what percentage can AI resolve" has no answer without knowing which questions you are asking it to resolve, which is the argument this whole article is making.
The metric itself is also unstable. Lorikeet found that figures published under the word "Resolution" carry a median of 72.5%, while figures published under "Automation" carry 61%. The gap comes from the denominator: resolution counts the conversations the AI handled, automation counts every conversation that arrived. Before you compare two vendors, ask each one which definition produced their figure. Ask us the same thing.
And elasticity is not capability. Handling four times the volume is not the same as handling it well. A lane that never falls over can still be quietly failing every question it was never taught.
Klarna is the cautionary case here, with a caveat. Klarna said AI was doing the work of 700 agents, then began hiring human agents back in May 2025, a reversal also covered by eMarketer. Klarna is fintech with permanent staff, not seasonal e-commerce, so it is not a template for your store. Its value is narrower and more useful. It shows what happens when the decision gets made at the level of "support" instead of at the level of question types.
There is also a risk neither benchmark captures. A wrong AI answer arrives instantly, in your brand voice, at scale, and looks exactly like a right one. A wrong human answer usually arrives one at a time. That asymmetry is the reason the handoff rules matter more than the resolution rate.
What neither option fixes
Neither hiring nor switching on AI writes your sizing guide, your returns policy, or your shipping cutoffs. Both lanes deliver answers. Neither creates them.
The peak week makes that visible. Across a base of Shopify stores, shopper satisfaction fell from 77.2% in the week before to 66.3% during BFCM week. In the same stretch, daily conversation volume peaked 44% above baseline (Chatty conversation data, Nov 24 to Dec 1, 2025).
That is a market-wide pattern measured across more than 2,000 stores, not a verdict on any single store or tool. What caused it is not established, and this article does not guess. Volume, shipping strain, discount-driven expectations and answer readiness all move in the same week, and one season of data cannot separate them.
The honest reading is the pairing itself. Shoppers rated the answers worst in the week they asked the most. That points at whether the answers were ready beforehand, not at who typed them.
So before you price either lane, find out what a shopper gets today when they ask your store a hard question at 11pm. Open your own store in a private window and ask it three things: a shipping cutoff, a returns edge case, and a product comparison. Whatever comes back is your real baseline, and it is more informative than either number above.
Hire seasonal staff if, scale with AI if, do both when
None of these conditions is a formula. Each comes from the sort you did earlier, so find yourself in one of the three lists.
Hire seasonal staff if:
- Judgment calls are the biggest of your three piles.
- You sell high-value, custom or made-to-order products, where exceptions are the norm.
- You already have someone who can train, and a customer service training plan to train against.
- Your peak is a lift you have absorbed before without the inbox backing up.
- You can rehire people who already worked a peak for you, which removes most of the ramp.
Scale with AI if:
- Lookups and advice carry most of your peak inbox.
- A large share of your demand lands outside your working hours.
- Your product and policy content already exists, or you have four to six weeks to write it.
- You have somebody who will read the failed conversations weekly and fix what caused them.
Do both when: you look at those two lists and see yourself in both, which describes most stores. The ratio between them is set by what your shoppers ask, not by your budget.
There is one case where neither list applies. If the content does not exist and there is no time left to write it, neither lane saves the season. The honest move then is to cut what you promise shoppers, publish your cutoffs and policies plainly, and stop selling the experience you cannot staff.
Which turns the question into how the two fit together.
The staffing model that survives peak
Stop choosing between people and AI, and start deciding what each one is for. The model that holds through peak runs in four layers, and people are the last layer rather than the removed one.
Layer one is content you publish, not answers you give. Shipping cutoffs, the sizing guide, the returns window, all findable without asking. In practice that means building a knowledge base before the season, not during it. Every question this layer prevents costs nothing for the rest of the season.
Layer two is AI on chat, covering the lookups. The same answer every time, including outside business hours, where more than half the volume lands. This is the layer that makes the after-hours number survivable, because the alternative is a form that promises a reply on Monday.
Layer three is your people, working the judgment calls. Hire for exceptions, complaints and the calls that need a human to make them. The route into this layer has to exist, and it has to be short. A shopper who has already decided they need a person will not wait in a queue to get one. Your top-10 question list and your handoff triggers belong in this layer too, and the BFCM CX guide covers how to write both.
Layer four is January. Returns arrive after the season, when your seasonal staff have already gone. Almost nobody plans this layer, even though after-sales questions already ran to about one in eight conversations during the peak days themselves.

Hire for judgment calls. Automate the lookups. Write the advice down, because that is where most stores break.
Product note, so you can skip it. Whatever tool you use for layer two, check that the route to layer three is not gated behind a paid tier, because that is a common place for this model to break in practice. Chatty includes the live chat inbox a handover lands in on every plan, including the free one, which caps the team at one seat.

Your inbox already made most of this decision
The hire-or-automate question has no answer at the level it usually gets asked. A store does not have support volume. It has three kinds of message, and two of them are not staffing problems at all.
Sort last November's messages into lookups, advice and judgment calls. Then the budget mostly writes itself. Automate the lookups, write the advice down, hire for the judgment calls, and plan for the returns that land in January after everyone has gone home.
Do the sorting before you post a job or start a trial. Both spends work far better once you know which pile they are for.
If you want the numbers behind this article in full, read what 580,000 shopper conversations showed about BFCM 2025. It breaks down the same peak season by question type, timing and channel. And if you want the checklist version of the four-layer model above, the BFCM readiness playbook walks through what to have in place before the volume arrives.
Frequently asked questions
Roughly $7,000 to $11,000 for the season fully loaded, according to Ringly's peak staffing playbook, which is a vendor estimate rather than a payroll study. Divided across the 1,500 to 2,500 messages one seasonal rep realistically closes, that works out to about $3.60 to $6.00 per answered message. Replace the bracket with your own wage plus tax plus training cost as soon as you can, because local labour markets vary widely.
On lookups, yes, and not narrowly. Per-conversation AI pricing across the category tends to land in the low single digits of dollars, against $3.60 to $6.00 for a human answer to the same tracking question. On judgment calls the comparison stops being useful, because a $6 answer that keeps an angry customer is cheap. The cost question only resolves once you know which pile your volume lands in.
No, and the benchmarks say so plainly. The median resolution rate across 195 deployments and 38 vendors is 70%, with a range running from 15% to 98.3%. That spread is the real finding: what AI can resolve depends entirely on which questions you ask it and whether the answers were written down beforehand. Judgment calls, meaning complaints, exceptions and refund negotiations, still need a person.
About 2 to 3 weeks before they are useful, per Ringly's vendor estimate. Peak itself lasts roughly four weeks, so by the time a new hire is good at the job, the job is nearly over. The exception is rehiring people who worked your last peak, which removes most of the ramp and is the strongest argument for hiring in this article.
Sort last November's messages into three piles first. If judgment calls dominate, hire, and start early because good candidates go fast in a thin year. If lookups and advice carry most of the volume, hiring buys you throughput at the most expensive possible rate. Most stores see themselves in both lists and need both, with the ratio set by what shoppers actually ask rather than by budget.
Returns arrive, usually right as seasonal staff leave. US retail added 494,000 jobs between October and December 2023, then cut 464,000 in January and February 2024. After-sales conversations already ran to about one in eight during the peak days themselves, so the January layer deserves planning that almost nobody gives it.



