How We Used AI to Surface the Right Answer Before the Agent Asked

A support agent at a two-monitor desk with a case open and a panel of suggested articles

The AI in support that gets all the attention is the autonomous agent that resolves a customer issue end to end with no human involved. That’s not what I built at Q4 Inc, and it isn’t where most of the real value is right now. What we built was a system that surfaced the right knowledge-base article in front of an agent before they’d thought to look for it. Not self-service. Informed service. I’d call it an AI knowledge base for support agents, and it changed more than I expected.

The distinction matters more than most people think. This is the follow-up to my practitioner’s reality check on agentic AI. That post covered the landscape, and this one covers what we actually put in and what happened, including the rough first months that vendor case studies skip.

What problem were we actually solving?

When I joined Q4, our support operation ran on Salesforce Service Cloud, a powerful platform that was frankly a machine gun where we needed a knife. We had an exceptional Salesforce admin who made it work in ways it probably wasn’t designed for. The knowledge base was comprehensive, with articles for most of what agents ran into. The content wasn’t the problem. Retrieval was.

An agent in a live interaction, under SLA pressure, with a customer waiting, doesn’t have time to work a search interface, compare four results and find the right section of the right article. In practice they either knew the answer from memory or they escalated. The knowledge base was for onboarding and for quiet moments, and during live work it was mostly ignored. I’ve seen this at every organisation I’ve worked in. The investment in the knowledge base is real and ongoing, and its use at the moment of need is close to zero.

We connected Forethought, an AI platform built for support, to our Salesforce environment through its Assist connector. It read our existing knowledge base and our case history. The agent saw a sidebar widget next to the case, nothing to learn and no change to how they worked. As soon as a customer enquiry arrived, the widget ranked the most relevant articles for that specific interaction.

Before, the agent read the enquiry, opened the knowledge base, searched, evaluated the results, found the article and located the section. Slow, inconsistent, and dependent on how well that individual knew the structure. After, the agent read the enquiry, glanced at the sidebar and saw the three most relevant articles already there. Their job became evaluation and response instead of search. It sounds incremental. The impact wasn’t.

Two colleagues sketching a workflow of boxes and arrows on a whiteboard

Were the first months any good?

No, and I’d rather say so plainly. For the first three to six months the model kept surfacing wrong or marginally relevant articles. Agents tried it, found it unreliable and went back to their old habits. Adoption was low because trust was low, and trust was low because accuracy was low.

Getting it to a useful state took active work. We retrained the model on our data, cleaned up tagging and metadata on the articles, and built a feedback loop so agents could flag bad suggestions. This isn’t plug-and-play. It learns from your data, so the quality of that data and the discipline of your training decide what you get. After about six months accuracy settled around 80%. Not perfect, but at 80% the agent gets a useful suggestion most of the time, and when it’s wrong they know to look elsewhere. Eighty percent accurate and instant beats a perfect answer nobody has time to find.

I’d tell anyone starting this to warn their team up front. Agents who try a tool in month two and find it unreliable form a bad mental model that survives long after the tool improves. Call the first three to six months a training period, for the AI and for the workflow, and don’t treat it as a pass-or-fail pilot.

What actually changed?

The impact showed up in first contact resolution. We tracked repeat enquiries, where a customer contacted us more than once about the same underlying issue, and saw a meaningful drop. That fits what Matthew Dixon and his colleagues found in their research on customer effort: what erodes loyalty is effort, and the customers who churn are the ones who contact you repeatedly or get partial fixes that need follow-up.

At its best, the surfacing let agents answer what the customer would ask next, not just what they’d asked. Someone resetting a password almost always wants to know what happens to their saved preferences. A billing query usually leads to a question about renewal terms. An agent with the full-scope article already open could cover both in one conversation, and the customer never had to call back.

It was also part of a larger redesign that included skill-based routing, which assigns cases by the complexity of the issue and the capability of the agent. Simpler issues went to newer agents, who were better equipped because the AI was feeding them articles, and complex cases stayed with senior people who had the institutional knowledge to handle them unprompted. The drop in first response time from 17 hours to 2 came from the routing redesign. The drop in repeat contacts, the better FCR and the shorter ramp-up for new agents came largely from the knowledge surfacing. We designed them together and they worked together. It also showed us the volume at each of our seven skill levels, so we could staff to actual demand. I wrote about the routing side in skill-based routing.

Why does retrieval matter more than content?

It’s a pattern I’ve come to recognise. Organisations pour effort into writing knowledge, and very little into getting it to the person at the moment they need it. A good article that takes ninety seconds to find is, for practical purposes, an article that doesn’t exist when a customer is waiting. Looking back, the biggest gain from the surfacing wasn’t any particular article. It was that the best information stopped being something agents had to go and fetch, and became something that was simply in front of them.

That shifts what you ask of your knowledge base. Instead of asking whether the content is complete, I started asking whether it’s findable at the point of use, and who notices when it isn’t. The feedback button on a bad suggestion turned out to be as valuable as the model, because it told us where the articles were thin, mislabelled or out of date. A tool that shows you where your content is weak is giving you a second benefit nobody put in the brochure.

What would I do differently?

Start with the knowledge base, not the AI. What Forethought could surface depended directly on the quality of our articles, their tagging, structure and coverage, and we spent real time fixing the content as part of the rollout. Garbage in, garbage out applies here as hard as anywhere. I’d also build the agent feedback mechanism on day one, not a few weeks in, because an earlier loop would have shortened the rough patch.

Measure outcomes, not utilisation. We didn’t treat article views as the success metric. We tracked repeat contacts and FCR, where the business value actually shows up. And be clear about what you’ve built. We didn’t build a system that resolved issues on its own, and customers were still talking to human agents. The AI did the retrieval and the agent did the judgment: deciding what applied to this customer, saying it clearly, and handling the human side of the conversation. I think that’s where most support organisations should put their AI budget right now. Autonomous resolution is hard to run reliably at scale beyond the simplest query types, and the blended model may be the right endpoint for a large share of interactions, not a stepping stone. If you’re weighing vendors, my notes on writing an RFP for a support tool will save you some pain.

Looking back, the part I remember best is the quiet stretch in the middle, when agents had stopped trusting the sidebar and I hadn’t yet earned that trust back. It would’ve been easy to call it a failed experiment in month three. We kept going because the retrieval problem was real and the fix, rough as it was, pointed the right way.

Leave a Comment

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.