Milena Traikovich has spent the last several years at the intersection of high-stakes performance marketing and emerging automation. As a leading expert in demand generation and MarTech optimization, she has navigated the transition from basic scripts to the sophisticated AI agents that define the PPC landscape in 2026. Her approach moves beyond the simple question of whether an algorithm can be trusted, focusing instead on the structural scaffolding required to keep live ad accounts profitable and secure. By specializing in lead quality and performance analytics, she has developed a rigorous framework for integrating artificial intelligence into systems where real-time spending requires absolute precision.
In this discussion, we explore the critical layers of AI safety for modern advertising accounts, moving through the necessity of deep data grounding and the implementation of rigid policy gates. The conversation covers why a “human-in-the-loop” must be a technical requirement rather than a vague intention and how comprehensive audit trails are transforming account management from reactive troubleshooting into a disciplined science. We also examine the specific data points—from GAQL integration to multi-platform benchmarks—that prevent AI agents from “going off the rails” and hallucinating results that could jeopardize a brand’s entire annual budget.
When an AI agent operates on a limited data layer, it often fabricates answers with total confidence. How do you ensure that these agents are properly grounded with full GAQL and GA4 data to prevent them from guessing?
Grounding isn’t just a convenience; it is the most fundamental safety feature we have because a blind agent is a dangerous agent. When an agent only sees a thin slice of an account, it will answer your questions fluently and immediately, but it won’t tell you it’s guessing because it doesn’t even know it’s guessing. To solve this, we integrate the full query layer for Google Ads, utilizing real GAQL to access every resource, field, and metric the API exposes, rather than just a curated summary. We also pull GA4 data alongside the ads data so the agent understands what happened after the click, preventing it from making wild assumptions across the boundary of different tools. By consolidating negative keywords across all four levels—account, shared lists, campaign, and ad group—the agent can deterministically check if a query is blocked instead of reasoning wrongly about the account’s state.
Many marketers struggle with the concept of trusting an automated system with their live budgets. How should we shift the conversation from “Do we trust the AI?” to “How can we build a framework that makes the AI trustworthy?”
The shift happens when you start evaluating an AI agent the same way you would evaluate a human collaborator or a new agency. You don’t just “trust” an agency in the abstract; you define what they can see, what they are allowed to change without permission, and who is responsible for reviewing their work. We apply these same three questions to agents to build a structural environment of safety: what can it see, what is it structurally prevented from doing, and who signs off? By moving away from the binary of trust versus no trust, we focus on building layers—grounding, gating, and reviewing—that each pay for themselves by closing different failure modes. When these layers compound, grounding makes the agent’s proposals worth reviewing, and policies filter out the obvious non-starters so the human workload remains manageable.
You’ve mentioned the importance of “automation layering” in your recent work. Can you explain how a separate policy layer acts as a guardrail that even the AI cannot override?
A policy layer is about writing down the “never do this” rules in a way that is entirely independent of the AI model itself. If you put a rule in a prompt, a clever user or a long session can talk the model out of it, but a policy layer lives on the account and doesn’t care who is asking. For example, we might set a rule that no bid increase can exceed 10% in a single move or that specific competitor brand terms can never be added as keywords. This creates a hard stop where a hallucination, a junior staffer with a misplaced decimal, or even an agent’s drift is blocked automatically. It is a structural guardrail rather than a “speed bump made of paint,” ensuring that if anyone wants to bypass the rule, they must do so deliberately and on the record with their name attached.
The “human-in-the-loop” concept is often cited but rarely enforced in a meaningful way. How does a formal change request sequence turn good intentions into a functional technical requirement?
Most people say they are in the loop, but if you ask them which screen or queue they use to enforce it, they usually admit they just check the change history after the money has already been spent. We’ve implemented a sequence where the agent proposes a change—like adjusting a target ROAS from 200% to 220%—and that proposal becomes a draft change request that never reaches the ad platform on its own. These requests are then evaluated against account policies, marked with verdicts, and presented to a human who can see the exact deterministic changes and the reasoning behind them. This process mimics engineering standards where no code is pushed to production without a peer review, ensuring that nothing changes in the ads account until a person specifically confirms the action. It turns the “loop” into a physical checkpoint that no agent can fabricate or bypass.
Beyond immediate safety, you’ve suggested that these safety layers actually create a superior form of account documentation. How does having a complete record of intent change the way agencies interact with their clients?
The change request queue has accidentally become the best account documentation tool we’ve ever had because it preserves the “why” behind every single action. Usually, if a client asks in November why a target CPA was moved back in March, you’re left digging through a group chat or a dry change history that shows the “what” but none of the context. With this audit trail, we have the original proposal, the agent’s rationale, the data used for the recommendation, and the name of the person who finally approved it. This builds immense credibility for an agency because you aren’t just showing a list of edits; you’re showing a disciplined, documented strategy that accounts for every dollar spent. It transforms the relationship from one based on mystery to one based on transparent, verifiable intent.
When dealing with competitive landscapes, hallucinations can be particularly hard to catch. How do vertical benchmarks and auction insights help ground an agent’s reasoning in reality?
Competitive questions are high-risk because you often have no independent way to verify if an agent’s answer is accurate or just a “fluently wrong” guess. We ground the agent by providing Auction Insights with drill-downs into specific competitor domains, allowing the AI to see your own performance on every keyword you share with them. We also pull in vertical benchmarks so the agent evaluates your CTR, CPC, and conversion rates as a percentile against real-time industry data rather than an outdated blog post average from years ago. This ensures that when the agent suggests a budget shift or a strategy change based on “market trends,” it is doing so based on hard data points from Google, Microsoft, Meta, and other platforms. It keeps the agent’s suggestions “boring” in the best way possible—predictable, data-backed, and devoid of the “excitement” that comes from unexpected errors.
What is your forecast for the evolution of agentic PPC over the next few years?
I believe we are moving toward a standard where “bare” AI agents—those operating without a policy and review layer—will be considered a professional liability. By 2028, the industry will likely have matured to a point where the “Model Context Protocol” (MCP) is the baseline for any serious ad operation, allowing agents to see across every platform from Amazon and TikTok to LinkedIn and OpenAI. We will see a shift where the primary job of a PPC expert isn’t to pull the levers themselves, but to manage the policy engine that governs how the AI pulls those levers. The most successful marketers won’t be the ones who “trust” AI the most, but the ones who build the most robust environments for AI to fail safely. Ultimately, this will lead to accounts that are not only more efficient but significantly better documented and more strategic than anything we could manage manually.
