Designing the context and control layer for an AI brand agent

Overview

Brand Kit is the source of truth behind AtlasNova’s AI-generated marketing assets.

Across social posts, captions, and brand voice, I redesigned it as a shared context and control layer between human intent and agent behavior.

The design turns users’ brand intent and input into structured context, constraints, and evolving preferences that guide how AI generates and adapts. It supports teams ranging from single-store operators with limited assets to enterprise brands managing large, distributed brand libraries.

My Role

I led Brand Kit’s product and agent behavior design.

Design

  • Owned research synthesis, information architecture, interaction design, and validation.

Collaboration

  • Partnered with engineers on context and constraints, and guided an intern’s model exploration.

AI-native workflow

  • Used AI tools end-to-end to draft agent specs, test prompts, prototype interactions, and evaluate outputs.

Design Challenge

Turn evolving brand inputs into reliable context for AI agents

Users arrive with established guidelines, scattered references, or only an early sense of their brand. Their preferences also evolve through content creation. Brand Kit needs to turn these changing inputs into consistent guidance for AI agents, while keeping interpretations and updates visible, reviewable and correctable. To address this challenge, I planned, designed and validated Brand Kit through the following workflow:

  • Ideation (Day 1): explore directions through discussion with ChatGPT.
  • Planning: review features, architecture, tech stack, database, and APIs.
  • Break the plan into scoped tasks and dependencies.

Research

To understand how users should experience Brand Kit and the broader GenAI journey, and how the AI agent could build on that experience, I conducted both user research and model research. Interviews and journey mapping showed me how users started and where they needed support; prompt testing and a review of the existing flow revealed where the system could be helpful or misleading. The intensive research gave me a clearer basis: Brand Kit had to work with uneven input while keeping users involved in shaping its interpretation.

01 — Different types of users have different inputs and ideas about their brands.

Through user interviews, I found that some of our users were still discovering how to describe their brand; others arrived with extensive material and needed to decide what mattered. I mapped both ends of that range to see how one experience could guide users without forcing them through the same steps.

User Interview
User interview spreadsheet
Brand Resource Reference
Brand resource reference

02 — AI agents need specific directions and references, but users don’t know what to give

Model and prompt testing showed that limited material could lead to unsupported brand facts, while large amounts of material made brand consistency harder to achieve. A word that felt clear to a user could still leave the model room to go off course. Testing also showed that image references guided image generation better than text-only descriptions, with six to nine images working best as a reference set. These findings gave me two concrete requirements: make it easier for users to supply visual references, and keep those references focused enough to guide generation.

Partial Testing Results

Models under test

Read

Turn a website or images into a brand spec

GPT-4o mini vs GPT-4o

Website to brand profile

GPT-4o mini vision

Reference images to style spec

Generate

The outputs under test

GPT-4o mini

Captions, tone, review replies

gpt-image-1

Brand images, seed-paired

GPT Image 2 · Nano Banana 2 · Seedream 5 Pro

Production models the prompts map to; not scored here

Judge

Did the output follow the rule?

GPT-4o

Text judge, temperature 0

GPT-4o vision

Image judge

Code checks

Counts, word lists, color distance

Brand extraction

GPT-4o mini and GPT-4o · 4 prompt versions

Original promptv1
Role and worked examplev2
One rule per fieldv3
Low-information modev4

Captions, tone and replies

GPT-4o mini, judged by GPT-4o · 4,179 outputs

Control labels144
Tone: labels, bundles, persona729
Signal repair1,080
Sentence count and energy405
Relative tone grading135
Five tones in four business types720
Warm tone, paired rerun480
Review replies243
Replies with a word cap243

Images

gpt-image-1, judged by GPT-4o vision · 597 images

Eight style rules162
Judge-first rerun90
Thirteen art-direction specs150
Wording ladder135
One-shot vs two-step60

By amount · random on-brand images

1 image
0.25
3 images
0.45
9 images
0.59
14 images
0.56

By mix · nine images

Curated 9
0.66
6 + 3 same category
0.44
6 + 3 other category
0.35
Everything, 23
0.35

Reading visuals is easier

System direction68%
Image prompt51%
Generic prompt37%

Control contribution

Composition+18
Image treatment+14
Palette+09
Subject+07

03 — Users need to automate the loop of evolving and updating the brand

Working with the marketing team and founders, we mapped out the user journey and discussed how to create a complete Brand Kit loop for our end users, from both UX and implementation perspectives. To keep all AI-generated content on track, the next challenge was to keep brand guidance consistent while making the reach of each change clear to users.

Feature Map
User Journey Mapping
User Comments on Maintenance and Updates
User comments on maintaining and updating the brand

Agentic Design Ideation

By learning about context and model capabilities, I defined the main design challenge: Brand Kit needed to do more behind the scenes without asking users to trust decisions they could not inspect. A higher level of automation means a better user experience with Brand Kit. Learning from both users and model capabilities, I began exploring how users interact with the interface and how to make the harness foundation trustworthy.

01 — Build the harness so agents guide the content creation process precisely

Drawing on user research, model testing, and the mapped user journey, I wrote brand-kit-agent.md—a behavior specification for the Brand Kit agent. It defines the harness rules for how Brand Kit processes customer inputs, prepares instructions for generation, and updates the brand profile. These rules made the intended behavior explicit and provided a basis for test cases.

Role of the agent
The Brand Kit agent writes profile versions and supplies context. Owners edit and accept changes; reviewers approve or request changes. Generators read context without updating the profile.
Scope
Read brand materials, posts, reviews and photo choices to maintain the profile and prepare creation guidance. Content generation and publishing stay with creation tools.
Constraints
Turn caption and reply styles into explicit writing rules, and photo readings into visual plans. Preserve product facts and supplied logos, and respect HQ restrictions.
Entry points
Route onboarding, material updates, Taste regeneration and accepted feedback to the right rules. Supply context on request; learn after ten new published posts, replies or photo choices.
Core rules
Prioritize HQ-locked values and owner corrections, select brand photos for the board, and derive palette, visual direction, caption and reply styles. Save changes as new profile versions.
Memory
Keep profile versions, the working board, corrections, source evidence, reviewer requests and event history so brand learning remains traceable.
Output format
Return a reviewable brand profile or task-specific context for captions, replies, images, video and campaigns, with explicit gaps, skipped sources, owner decisions and error states.
Definition of done
Check evidence-backed brand facts, board and Taste rules, active corrections and sample quality before saving a version. Validate each creation tool’s context before returning it.
Failure modes
Retry failed tools, skip unreadable brand files or links, and flag missing products or logos. Stop without saving when processing or validation fails.
Human escalation
Surface conflicting brand guidelines and reviewer requests that cannot map to a brand or Taste field. Give owners the alternatives and flag low-confidence values.

02 — Rapid testing informed a visual-first moodboard and an intuitive input flow

Across several rounds of rapid user testing on prototypes I built with AI tools, we found that users needed a faster way to judge brand direction than reading a text-heavy profile. Model testing added two findings: image references guided image generation better than text-only descriptions, and six to nine images worked best as a reference set.
These findings informed the agent’s board rules: select up to nine brand photos within one visual direction, keep additional photos outside the active board, and use the board to guide visual interpretation. Written brand restrictions and explicit user corrections retain priority, so the moodboard supports generation without overriding the rest of the brand profile.

Ideation canvas, part 1 of 3
Ideation canvas, part 2 of 3
Ideation canvas, part 3 of 3

03 — New choices update the profile through rules owners can inspect and correct.

A profile that stays fixed after setup would fall behind the content users actually choose. If every edit rewrote it, users would lose control. By testing agent behavior across different use cases, I defined how user behavior (selected photos and published captions or replies) could inform later guidance, which corrections retain priority, and when a new version should be saved. Each change keeps its source and reason, so Brand Kit can adapt without hiding how its direction changed.

Key Design Decisions

User research revealed how people express and refine their brands; model testing revealed what the system needed to interpret that intent reliably. I brought these findings into the Brand Kit agent’s rules for reading inputs, selecting references, and updating brand context. Together, the research and agent design informed three interface decisions: an adaptable starting point, a visual-first moodboard, and a traceable update loop.

01 — One starting point adapts to different levels of brand readiness

User research showed that some users were still discovering their brand, while others arrived with extensive materials. Model testing added another constraint: limited evidence could produce confident but unsupported brand facts. I translated both findings into agent rules for handling different inputs, identifying missing evidence, and preserving user corrections. This informed one starting point where users could begin with what they had, while the agent handled the materials behind the scenes. Users remained involved in choosing what represented their brand and checking uncertain details.

02 — Make brand direction visible and steerable through a visual-first moodboard

User testing showed that people needed a faster way to judge brand direction than reading a text-heavy profile. Model testing showed that image references guided image generation better than text-only descriptions, with six to nine images working best as a reference set. I translated these findings into agent rules that selected up to nine brand photos within one visual direction and kept additional photos outside the active board. Together, these findings and rules informed a visual-first moodboard at the top of Brand Kit. Users could review and change the references, then regenerate the Taste Panel to inspect the resulting direction, while written brand restrictions and explicit corrections remained in force.

03 — Let brand context evolve without losing user control

Journey mapping showed that brand decisions continued beyond setup, through the photos users selected and the content they published. The agent design defined how these signals could inform later guidance, while preserving explicit corrections and recording each change’s source and reason. Together, these findings informed a traceable update loop: Brand Kit could learn from everyday work, and users could inspect what changed, correct its interpretation, or restore an earlier version. Adaptation became part of the workflow while deliberate brand decisions stayed in the user’s hands.

Evaluation

Prompt and Spec Evaluation: testing what the interface could promise

I treated AI model behavior as part of the UX design. For each control, I defined the expected output difference, held other variables constant, generated controlled variants, and compared the results. AI generated the variants and provided a first-pass blind evaluation; I audited the evaluator, reviewed failure patterns, and made the product decisions.

Evaluation Metrics
Controlled prompt variants for image tone, video direction, caption voice, and review replies
Signal specification table comparing metrics, level granularity, concrete prompts, and measured effectiveness

Impact

Impact across user control, model behavior, and system reuse

Brand Kit launched in July. The following outcomes reflect the design’s impact on user control, model behavior and reuse across content creation workflows.

User

  • Brand decisions became visible and correctable, with fewer prompt revisions needed to refine each post.

Model

  • Controlled evaluation improved refined-tone compliance from 33% to 93%, with 93–98% generation stability in validated tests.

System

  • One shared brand context could carry reviewed decisions into campaign generation, captions, and AI review replies.
8/10
Users satisfied with AI-generated content
User-reported GenAI output quality
50%
Fewer prompt iterations per post
Before: 6 iterations → After: 3
93%
Refined-tone compliance
33% → 93% in controlled evaluation
80%
Brand-system creation satisfaction
Users satisfied with creating their brand system
Reusable
Brand decisions across creation workflows
Shared context for campaigns, captions, and AI review replies

Takeaways

Design the capability system, not just the model

Designing an AI-native workflow requires understanding two different capability layers: what the model can reliably infer or generate, and what the product can collect, store, verify, control, and recover from. I learned to start with the user’s goal, then decide which parts should be handled by AI, deterministic product logic, or human judgment. The model should expand the workflow, not define its limits. Good AI UX uses AI where it is strong, builds controls and fallbacks where it is weak, and makes those boundaries understandable to users.

Field Research
Field research: merchant interview
Field research: observation with a merchant
Field research: reviewing the interface in context