Introducing Adversary Intelligence: LLM Tradecraft

Read Time

7 mins

Published

Sep 30, 2026

Share

TL;DR: Adversary Intelligence: LLM Tradecraft is a new on-demand course covering how modern LLMs and agentic systems work, how they can be exploited, and how to secure them in practice. 

In July of this year, we released our first on-demand courses as part of the SpecterOps Tradecraft Academy. Today, we’re excited to introduce our new course for the academy developed in partnership with OpenAI: Adversary Intelligence: LLM Tradecraft.

Security practitioners are increasingly expected to use, assess, and secure LLMs and agents, often before they’ve had the opportunity to build a solid understanding of how these systems work. We’ve built this course to close that gap. Across ten standalone, on-demand modules with hosted labs, we cover everything from machine-learning fundamentals to LLMs, agents, and their application to security work, along with how to attack and secure the systems built around agentic architectures.

This course is built for security practitioners who want to understand the concepts behind modern LLMs + agents and their application to infosec, but don’t have an extensive background in machine learning. We cover the underlying mechanics as we go, connecting unfamiliar concepts to security work rather than expecting you to arrive knowing what an embedding is or how an agent manages its context.

An Infosec Approach to AI

Back in 2022, our Learning Machine Learning blog series worked through model training, feature selection, and evaluation using obfuscated PowerShell as a case study. Our goal with that series was to build intuition for fundamental machine learning concepts (without drowning you in math!), digging into the details while keeping everything grounded in realistic infosec scenarios.

A lot has changed in the last several years with the explosion of LLMs and the rapid evolution of their capabilities, but we’re taking a similar approach with this course. There’s a substantial amount of machinery between training a language model and using an agent to perform security work, and understanding that machinery makes it much easier to decide where to apply these tools. We start with the underlying concepts, then build up to systems that can retrieve information, execute tools, maintain state, and work through a task over multiple steps. From there, we explore how to use Codex for practical security work, including incident investigation, threat modeling, reverse engineering, and more.

One throughline in the course is the interplay between the model and the system around it. The model is only part of an agent; the surrounding software, or harness, determines what information it receives, which tools it can use, how its state is preserved, and how execution proceeds:

That same foundation carries over to security assessments. An LLM application still has identities, APIs, dependencies, and data-access decisions, but now model context can influence how those components are used. We cover the familiar application-security problems alongside the additional attack paths introduced by prompts, retrieval, memory, and tools. We think learning to both build and assess these systems, with a solid grasp of the underlying concepts, helps you make better decisions about where to use them, how much to trust them, and how to test their limits.

What We Cover

The course is divided into the following ten modules:

AI Foundations – We start with machine learning through security problems such as phishing detection and malware classification. Topics include data and feature engineering, common model families, training versus inference, and evaluation. We also cover why apparently good results can fall apart because of class imbalance, drift, or thresholds that produce more alerts than anyone can investigate.

Large Language Models – We follow how an LLM produces an answer, from tokenization and embeddings through attention and token generation, then look at training, reasoning models, and their limitations. This module also covers embedding models and retrieval-augmented generation (RAG), including how to supply private or current information and the quality and security problems that retrieval introduces.

Prompting – We cover message roles, sampling, zero-shot and few-shot prompting, task decomposition, and context management, along with structured outputs and validation. The emphasis is on understanding how these techniques affect behavior so you can adapt them to your own work.

AI Agents – We break down the model-plus-harness architecture and build up the execution loop: tool calls, observations, planning, memory, and stopping conditions. We then compare fixed workflows, bounded agents, and multi-agent designs, including the practical tradeoffs around context, tool design, autonomy, cost, and human involvement.

Coding Agents and Codex – This module focuses on using Codex for security tasks beyond ordinary software development and turning those tasks into reusable workflows. We cover repository guidance through `AGENTS.md`, skills, integrations, subagents, configuration, sandboxing, and approvals, as well as ways to inspect, steer, and automate the work.

Model Context Protocol – We work through how MCP connects clients to tools and data, and how tool definitions and results ultimately enter the model’s context. From there, we examine local and remote server risks, tool-description injection, dynamic definition changes, and tool hijacking, alongside familiar API problems such as missing authorization, injection, confused-deputy behavior, and SSRF.

Observability, Evaluation, and Optimization – We use traces to examine agent behavior and evaluations to compare changes to prompts, models, and harnesses. Topics include MLflow, deterministic verifiers, LLM-as-judge, repeated runs, and automatic prompt and harness optimization. We also look at evaluator bias and reward hacking: improving a score doesn’t necessarily mean we’ve improved the system.

The LLM Threat Model – We assess the complete application, including its model provider, data flows, identities, tools, persistent state, and downstream consumers. The material connects conventional application security with LLM-specific risks, expands supply-chain review beyond software packages, and uses an AI-assisted support application as a case study for locating trust boundaries and assessing potential impact.

Jailbreaks, Prompt Injection, and Defenses – We distinguish attacks on model safety behavior from attacks on application instructions and hidden context, then examine how they work. This includes direct and indirect injection, prompt extraction, documented attack chains, repeatable testing, and practical defenses, with particular attention to which protections rely on model behavior and which the application can enforce.

Reverse Engineering with LLMs – We introduce the relevant Ghidra workflows, function identification, and reusable type information before exploring ways to integrate LLM assistance. The module covers both improving the analysis supplied to the model and checking its interpretations, including how deterministic tools and tests can complement the parts of reversing where flexible reasoning is useful.

Every module has a set of hosted labs that ground the concepts in security-focused scenarios. You’ll build classic ML models, work through in-depth prompting exercises, create agentic workflows, evaluate agent runs with MLflow, use Codex to reverse malware, and more!

Taking the Course

We’ve designed the course modules to stand on their own so you can work through the full progression or concentrate on the areas relevant to you. Someone getting started with LLMs can begin with the foundations, while someone already building agents may be more interested in evaluation, threat modeling, or MCP security. The reversing material provides another practical entry point for people interested in applying these tools to binary analysis.

There is very little math in the course, but we do spend time on the concepts needed to reason about the systems we’re using. The aim is to give you enough background and hands-on experience to adapt these techniques to your own security work, rather than needing a new recipe every time the model, framework, or task changes.

Enrollment opens today at https://academy.specterops.io/adversary-intelligence-llm-tradecraft

Our first cohort kicks off October 15, with new cohorts opening monthly. Each cohort includes 30 days of access to all course materials, hands-on labs, and a ChatGPT Pro subscription from OpenAI.

Will Schroeder

Principal Security Researcher

Will Schroeder (@harmj0y) is a Principal Security Researcher at SpecterOps specializing in machine learning and offensive development. He has co-authored numerous projects ranging from BloodHound to the “Certified Pre-Owned” white paper.

Ready to get started?

Book a Demo