Key facts
- On 2026-09-23, gpt-5.4-mini took the SEO-agency-selection question three separate times, browsing switched off for the entire test and no setting adjusted from one attempt to the next.
- All three runs produced the same evaluation framework in substance: relevant industry experience, a clear first-90-days strategy, transparent reporting tied to business KPIs rather than vanity metrics, technical and content competence, safe link-building practices, named account ownership, and realistic pricing.
- Our capture of each response was limited to 4,000 characters, while the model's actual output ran 4,046, 4,105, and 4,019 characters respectively. All three responses were therefore cut off before the natural end of the model's answer.
- Run 1's capture ended mid-sentence just as it introduced a section titled “Agencies I'd point to,” before naming a single one. Run 2's capture ended mid-word inside its pricing section, before it ever reached an agencies section we could see.
- Only run 3's capture reached far enough to show named agencies: WebFX, HigherVisibility, Victorious, Straight North, and Thrive Internet Marketing Agency for small business and local SEO, and Seer Interactive and Directive for “more strategic or premium” SEO, with Directive's description itself cut off mid-word.
What we asked and how
gpt-5.4-mini took three clean runs at “How should a small business choose an SEO agency? Give specific criteria and name any agencies you would point to,” all on 2026-09-23, en-CA locale, no browsing, nothing carried over between them.
This prompt asks for two things at once: a set of evaluation criteria, and specific agency names. The two halves of the answer behaved very differently across the three runs, and a data limit in how we captured the responses affected what we can honestly report about the second half.
The evaluation framework, run by run
All three runs produced what amounts to the same checklist, organized differently but covering the same ground each time: has the agency worked with businesses like mine, do they explain a concrete first 30, 60, or 90-day plan rather than vague promises, do they report on real business outcomes like leads and revenue rather than only rankings and traffic, do they use white-hat, explainable methods, who specifically will work on the account and how many clients that person handles, and does the agency's pricing match the scope of what a small business actually needs.
All three runs also independently listed the same red flags: guaranteed number-one rankings, vague or secret methods, no case studies or references, and long contracts with no exit terms.
A limit in our own data: every response was cut off
We need to disclose something about how this particular prompt's data was recorded. Our capture recorded each response's true length, 4,046, 4,105, and 4,019 characters, but only stored the first 4,000 characters of the actual text for all three runs. That means every one of these three responses is missing its final section, in the model's own words, not something we edited out.
This matters specifically for the “name any agencies you would point to” half of the prompt. Run 1's stored text ends mid-sentence, right as it introduces a section called “Agencies I'd point to,” with the sentence “These are examples of agencies often recognized in the SEO space. Whether they're right for you depends” and nothing after it. Run 2's stored text ends mid-word in its pricing discussion, before reaching any agencies section we can see. We are reporting this limit plainly rather than guessing at what either run said next.
The agencies that surfaced, and how partial that list is
Run 3's stored text is the only one that reached a named-agencies section before the 4,000-character cutoff. It split its list into two tiers: for small business and local SEO, it named WebFX, HigherVisibility, Victorious, Straight North, and Thrive Internet Marketing Agency, each with a short qualifier, for example calling Thrive a “common SMB option, though you should review team quality closely by account.” For more strategic or premium SEO, it named Seer Interactive and then Directive, with Directive's description cut off mid-word by the same 4,000-character limit.
We are not treating this as the model's complete or definitive agency list, only as the one partial list our capture happened to reach before running out of stored characters. Whether runs 1 and 2 would have named the same seven agencies, an overlapping set, or a completely different set, in the same pattern we saw with the dental-marketing and Vancouver Google Ads prompts, is not something we can answer from this dataset.
What this does and doesn't tell you
This shows that the criteria half of this question is something the model answers with a stable, template-like framework across runs, similar to how the Google Ads pricing prompt in this same panel produced consistent numbers. It does not show a reliable, complete answer to the naming half of the question, because our own capture limit cut every run short before we could confirm the full picture.
A small business reading this should treat the evaluation framework as a reasonably stable checklist to bring to any agency conversation, and should treat the seven named agencies as one partial, unverified list from one AI model on one date, not a vetted shortlist.
Related questions
Yes, in substance. All three runs on 2026-09-23 produced the same core framework: relevant experience, a concrete first-90-days plan, outcome-based reporting, technical and content competence, safe link building, named account ownership, and realistic pricing, each worded a little differently.
Only run 3's captured text reached a named-agencies section before being cut off: WebFX, HigherVisibility, Victorious, Straight North, and Thrive Internet Marketing Agency for small business and local SEO, plus Seer Interactive and Directive for more strategic or premium SEO work.
Our data capture for this prompt stored only the first 4,000 characters of each response, while the true responses ran slightly longer, 4,046 and 4,105 characters. Run 1 was cut off right as it introduced its agencies section, and run 2 was cut off before reaching one at all. We are disclosing this as a limit rather than filling in a guess.
No, and that's the point of flagging the capture limit. Even run 3's list may not be complete, since its stored text also ends mid-word, cut off partway through describing Directive. Treat every name here as a partial, not exhaustive, list.
All three runs warned against agencies that guarantee number-one rankings, describe secret or unexplained methods, can't produce case studies or references, or lock a business into a long contract with no exit terms.
Sources
SearchPod sells SEO services, which gives it a direct interest in how AI tools describe agency selection, and SearchPod did not turn up among the agencies named. The transcripts above come from a single model on a single date, partially cut short by our own capture limit, and none of the named agencies are SearchPod or a SearchPod client.
Want to know how your business shows up?
Get a free, no-obligation proposal within one business day. We look at your site and your market and tell you plainly what we would do, and what we would not.
Get your free proposal