
Lessons from Alex Albert
Alex Albert is a research PM at Anthropic. In his talks, he describes how developers can evaluate models, revisit old prompts and scaffolding, and build with Claude as its capabilities change. — Capability Curve Talk.
Part 1: The Art of Prompt Engineering
- Separate Parts of a Prompt: Albert recommends XML tags to mark distinct instructions, examples, and input text in Claude prompts. — Anthropic Prompting Tips.
- Use Long Context Deliberately: Albert points to Claude’s long context as a way to include substantial source material in a prompt. — Anthropic Prompting Tips.
- Show Examples: Albert advises giving a range of examples to show Claude the desired behavior. — Anthropic Prompting Tips.
- Let the Model Think: Albert recommends allowing thinking steps before a final answer on complex tasks. — Anthropic Prompting Tips.
- Iterate and Measure: Albert describes prompting as rounds of optimization and advocates testing candidate prompts against benchmarks. — AI Engineer Talk.
- System Prompts Can Leak: In a 2023 test, Albert reports eliciting the rules of a system message he wrote, showing why developers should not treat that message as a secret security boundary. — Cognitive Revolution Interview.
- Specify the Desired Output: Albert’s prompting advice starts with clear, direct, specific instructions about the task and desired behavior. — Anthropic Prompting Tips.
- Test Prompt Changes: Albert describes an empirical process: run prompt variants against a benchmark rather than trusting one successful example. — Anthropic Prompting Tips.
Part 2: Jailbreaking and AI Security
- Make Testing Shareable: Albert built Jailbreak Chat to gather scattered attack prompts and let people compare, rate, and iterate on them in one place. — Cognitive Revolution Interview.
- Roleplay Can Circumvent Rules: Albert describes early jailbreaks that assigned the model a new persona and later variations that revived the tactic. — Cognitive Revolution Interview.
- Helpfulness and Safety: Albert discusses the tension between restrictive responses and useful creative assistance, while recognizing broad safety boundaries. — Cognitive Revolution Interview.
- Defenses and Attacks Coevolve: Albert reports that providers patched some jailbreaks while new variants continued to appear; he treats this as a continuing testing problem. — Cognitive Revolution Interview.
- Community Red Teaming: Albert argues that public jailbreak experiments can reveal model failure modes and give labs useful examples to examine. — Cognitive Revolution Interview.
- The DAN Family: Albert describes early persona-based ‘do anything’ jailbreaks and later variations of the tactic. — Cognitive Revolution Interview.
- Token Smuggling: Albert describes combining earlier ideas into a token-smuggling jailbreak and is explicit that he could not be sure which model mechanism it exploited. — Cognitive Revolution Interview.
- Use Findings to Educate: Albert proposes using jailbreak findings to explain model boundaries and failure modes to the public rather than ignoring them. — Cognitive Revolution Interview.
Part 3: Developer Relations and Community
- Bring User Feedback into Research: In a later interview, Albert discusses how feedback from Claude users informs model development and the capabilities his team prioritizes. — Behind the Craft Interview.
- Organize a Community Feedback Loop: Jailbreak Chat provided a common place to share working prompts and compare new variants, shortening the feedback loop. — Cognitive Revolution Interview.
- Translate Capabilities into Products: Albert’s AI Engineer talk moves from model benchmarks to concrete product interfaces, API tools, and cost considerations for builders. — AI Engineer Talk.
Part 4: Building with Claude
- Artifacts as a Work Surface: Albert demonstrates how Artifacts separate generated apps, documents, and graphics from the chat used to create them. — AI Engineer Talk.
- Internal Repository Context: Albert says Anthropic engineers upload repositories and documentation to Claude Projects as context for their work. — AI Engineer Talk.
- Ground Work in Project Knowledge: Albert presents Projects as a way to supply style guides, codebases, transcripts, and prior work as reusable context. — AI Engineer Talk.
- Share Context with Teammates: He describes sharing Projects, chats, and Artifacts so teams can carry context forward alongside the resulting work. — AI Engineer Talk.
- Connect Claude to Functions: Albert explains that tool use lets developers expose client-side functions for Claude to invoke, connecting the model to application behavior. — AI Engineer Talk.
- Test Coding on Real Tasks: Albert favors pull-request evaluations for multi-step coding because a model can attempt, write, and test a defined change. — AI Engineer Talk.
- Give Reasoning Room: Albert advises allowing Claude time to think through a complex task before producing a final answer. — Anthropic Prompting Tips.
Part 5: The Evolution of Large Language Models
- Measure the Capability Curve: Albert uses coding evaluations to illustrate how quickly model capabilities changed across recent releases. — Capability Curve Talk.
- Instruction Tuning Changes Usability: Albert gives a simplified explanation of how supervised examples and human preference ranking shape a chat model into a usable assistant. — Cognitive Revolution Interview.
- Vision Extends the Interface: Albert demonstrates extracting a table from an image and turning visual input into structured text, while noting the model’s vision opens additional uses. — AI Engineer Talk.
- Keep Evals Informative: Albert warns that saturated benchmarks stop differentiating newer models and urges evaluations near the product’s actual task distribution. — Capability Curve Talk.
- Reasoning Changes Across Models: Albert contrasts older coding agents’ tendency to act before planning with newer models’ stronger planning and error recovery. — Capability Curve Talk.
- Balance Model Capability and Cost: Albert stresses that intelligence, latency, and price together determine whether a model is economical for an application. — AI Engineer Talk.
- Latency Is a Product Constraint: Albert treats lower latency alongside higher intelligence and lower cost as a product-design consideration. — AI Engineer Talk.
- Cascade Between Models: Albert sketches a possible design in which a smaller local model handles simple work and calls a more capable cloud model for harder requests. — Cognitive Revolution Interview.
Part 6: AI Tooling and Agentic Workflows
- Tool-Enabled Workflows: Albert describes exposing functions to Claude so model output can become application actions rather than text alone. — AI Engineer Talk.
- Agents Need Room and Tools: Albert recommends giving agents useful tools and enough autonomy to plan, act, inspect results, and iterate within controlled access. — Capability Curve Talk.
- Control Risky Actions: Albert describes a mode that checks proposed tool calls and seeks human approval for actions classified as needing it. — Capability Curve Talk.
- Sustain Attention Over Long Runs: Albert says newer coding models can keep instructions and task context coherent over longer agent runs than earlier versions. — Capability Curve Talk.
- Recover from Failures: He contrasts earlier agents’ repetitive failure loops with newer models that can backtrack and try a different approach. — Capability Curve Talk.
- Inject Relevant Project Knowledge: Albert describes using Projects to ground Claude in codebases, style guides, and other task-specific knowledge. — AI Engineer Talk.
- Revisit Scaffolding: Albert argues that scaffolding useful for older models can impede newer ones; test whether simplifying prompts and workflows improves results. — Capability Curve Talk.
Part 7: Product Management in AI
- Prototype for Feedback: Albert urges builders to create quick prototypes to validate ideas and start a feedback loop. — AI Engineer Talk.
- Ask Builders What Matters: Albert invites developer questions and, in a later interview, discusses using feedback to shape model-development priorities. — Behind the Craft Interview.
- Use the Product Internally: Albert reports Anthropic engineers using Claude Projects with their repositories and documentation. — AI Engineer Talk.
- Affordability Enables Uses: Albert says a faster, less expensive model was more practical to embed in applications; the exact 2024 launch prices are historical, not current pricing advice. — AI Engineer Talk.
- Design Around Capabilities: Albert uses Artifacts and Projects to argue that better product interfaces can matter as much as model scores for a usable experience. — AI Engineer Talk.
- Treat Prompt Leakage as Product Risk: Albert’s jailbreak testing found that a system message could be exposed, which he connects to developer reputation and security risks. — Cognitive Revolution Interview.
- Evaluate Your Real Task: Albert urges developers to build evaluations resembling their users’ tasks and to keep them challenging as models improve. — Capability Curve Talk.
Part 8: The Future of AI Development
- Compare Working Software, Not Just Demos: Albert compares the same app-building task across two model generations, emphasizing functional behavior and end-to-end results over an attractive UI alone. — Capability Curve Talk.
- Describe a Task, Then Verify: In Albert’s coding demonstration, a natural-language request can drive substantial app construction, but the resulting software still needs testing. — Capability Curve Talk.
- Personal AI as a Possibility: Albert imagines individualized AI help across education and daily life. — Cognitive Revolution Interview.
- Build AI-Native Interfaces: Albert argues for redesigning products around what models can do, rather than adding an AI button to an unchanged interface. — AI Engineer Talk.
- Keep Reassessing Models: Albert advises retesting applications on new model releases and pruning prompts or scaffolding that no longer help. — Capability Curve Talk.
- Alignment Remains Open: Albert says he is optimistic about alignment while doubting that the methods available in 2023 were sufficient on their own. — Cognitive Revolution Interview.