Data-driven strategy for Fabio ArΓ£o β’ US/EU Senior Markets
Live System
IAM + AI β’ 2 pills/day
π IAM Β· OktaOct 1, 2026
Okta Entitlement Management: The Assignment Method Is the Access Model, and Switching It Empties the Grant
π‘ Key Concept
A group answers one question: which app. An entitlement answers the question an auditor actually asks: which permission inside the app. Okta's definition is deliberately narrow β per the Okta Entitlements documentation, Oct 2026, "an entitlement is a permission, privilege, or access level that allows users to take specific actions within a third-party app," and, per Okta's Governance API reference, it only exists for apps where you have turned the capability on: "an app must have entitlement management enabled before it can be used as an entitlement resource." Everything else in Okta Identity Governance consumes this model. Access requests request entitlements. Certification campaigns certify entitlements. Reports count entitlements. So the modelling decisions you make here are not a configuration detail β they are the vocabulary your whole governance programme will speak for years.
The object model is small. An entitlement has values. An entitlement bundle β "entitlement bundles allow you to grant multiple entitlements simultaneously to your users," per the Okta entitlement bundles documentation, Oct 2026 β is the virtual role: a named set of entitlement values scoped to exactly one app. In the Okta Governance API reference for Entitlement Bundles, Oct 2026 that is literally the shape: a required target of type APPLICATION whose externalId is the Okta app.id, plus an entitlements array of ids and value ids, a status of ACTIVE or DELETED, and an orn. The first operational surprise is in the same docs: "you can't assign entitlement bundles directly to your users from the Admin Console." A bundle reaches a human through Access Requests or through the API β never through the click-path an admin reaches for in an incident. That is a defensible design choice and a staffing fact: if your break-glass procedure assumes an admin can hand someone a role in the console, bundles are not where that role lives.
The second surprise is the one that changes architecture, and it is why this pill exists. The assignment method is not a label on a grant; it is a mode on the user. Per Okta's own Unlocking Okta: Understanding Entitlements and Assignment Behavior, Oct 2026, "when assigned a custom entitlement via the UI, the user will automatically be assigned to the Custom Assignment method type and no longer be eligible for Policy-rule-based entitlements." One manual grant, made once, in a hurry, silently removes that user from policy-driven governance for that app β and the way back is destructive: on reverting to policy "the user will lose all entitlements in their existing grants and will then be granted entitlements based on Policy." Okta's stated mitigation is an ordering rule, not a feature: "start with a Policy Assignment type." Layer on the data-type behaviour and you have two opposite failure modes living in one tenant. "By default, <string> data type entitlements allow a user to hold only one value at a time in the assignment" β granting a new value replaces the old one, which is greyed out β while "string array-based entitlements are additive." Whether grant means add or replace is therefore decided by the target app's schema, not by your policy, not by your request workflow, and not by anything visible on the approval screen.
Cross-pollination: today's AI pill is about barge-in in a voice agent, where the server keeps believing it said the whole sentence unless the client reports the millisecond at which the caller actually cut it off. Same failure class as a single-value entitlement: a later write silently supersedes an earlier one, nothing in the record marks the supersession, and the system's history and reality quietly stop matching.
π¬ Deep Dive
Practitioner trap: you will audit bundles with the API, get a clean-looking list, and conclude your bundles are empty. The list endpoint hides their contents by design β the reference states that the include filter "adds additional properties that are available in the retrieve an entitlement bundle operation, but are omitted from the list response normally," with the single documented value full_entitlements. So GET /governance/api/v1/entitlement-bundles returns names, orn, status and timestamps, and your "what is in each bundle" report comes back structurally correct and semantically blank unless you ask. Two more edges on the same path. The default page size is 20 with a documented range of 1 to 200, cursor-paginated through after (a 20-character id matching enb[0-9a-zA-Z]+), and the rate-limit headers in the reference illustrate a ceiling of 60 β so a naive "export every bundle in the tenant" loop is both truncated and throttled, and the first version of that script will quietly under-report. And deletion is a lifecycle, not an event: a bundle moves to statusDELETED, but "when the entitlement bundle is purged, it isn't returned in a GET operation" β meaning the API surface is not a durable record of who was once in which bundle. If your evidence story for an auditor is "we can reconstruct historical bundle membership from the governance API," test that assumption against a purged bundle before you write it into a control description.
Staff-level framing: read Okta's limit changes as design signals, because they tell you which modelling style the product is now built for. The Okta Identity Governance API release notes, Sep 2026 record that in the monthly 2026.09.0 release (10 Sep 2026) the ceiling for "entitlement bundles in an access level condition" went to 1,000 from 100, groups in the same condition to 1,000 from 500, users per task or question to 25 from 10, request type configurations to 250 per org from 100, and request types per org to 750 from 500. A tenfold increase in bundles-per-condition is not a rounding fix; it is Okta conceding that customers model entitlements per team, per region and per job family, and that the previous ceiling blocked real designs. Which hands you the governance risk in the same breath: bundles are the new groups, and they will sprawl the same way, except they are invisible in the Admin Console assignment path and each one is pinned to a single app. Decide now who owns bundle naming, what the maximum per app is, and how a bundle gets retired β because the API will happily let you build 1,000 of them and the certification campaign that reviews them is a separate product surface. The review side did move in your favour this quarter: per the same notes, Governance Analyzer reached GA in production in the 2026.08.0 release, service account certification arrived with the campaign resource types OKTA_SERVICE_ACCOUNT and APP_SERVICE_ACCOUNT, and in September AI agent resource connections became certifiable in identity campaigns. Non-human identities are now first-class in the review flow; entitlement bundles for them are still yours to model.
The boundary of the model β and it is the boundary an auditor will find. Entitlement Management governs what Okta discovered and what Okta assigned. A permission granted by someone with local admin rights inside the SaaS app itself is, structurally, outside that grant model: Okta can see it as a discovered value, but there is no Okta-side rule, request or approval that explains it. That gap is precisely what Okta is now selling against. Its product innovations announcement of 9 Sep 2026 lists Automated Drift Detection & Remediation to "quickly identify unapproved access changes and revoke rogue permissions in real time," and Advanced Entitlement Management for AWS to "centralize fine-grained permission tracking across developers, workloads, and AI agents" β both stated as early access by Q4. Two things follow for anyone designing today. First, the honest control statement for the next two quarters is "we govern the entitlements we assign and we detect, with a lag, the ones we did not," not "we control all permissions in this app." Second, the AWS item is the one to watch architecturally, because it pulls entitlements out of the SaaS-app frame this product was built on and into cloud permissions, where the grant model is IAM policy and the subject may be a workload or an agent rather than a person. If you are drawing a target-state diagram, draw Okta as the authority on who should hold which entitlement and leave a labelled seam where the cloud's own authorization engine evaluates it β the seam is where every one of these integrations lives.
# Okta Governance API: the three calls an entitlement audit actually needs.
# Scopes: okta.governance.entitlements.manage (write) / .read (read)
# Admin role: APP_ADMIN. Placeholders only -- not executed in this run.
# 1. What entitlements does this app even expose, and what values do they have?
# (an app must have entitlement management enabled to be a resource at all)
curl -s -X GET \
'https://YOUR_OKTA_DOMAIN/governance/api/v1/entitlements' \
-H 'Authorization: Bearer <TOKEN>'
curl -s -X GET \
'https://YOUR_OKTA_DOMAIN/governance/api/v1/entitlements/<ENTITLEMENT_ID>/values' \
-H 'Authorization: Bearer <TOKEN>'
# 2. Create a bundle = a virtual role scoped to exactly ONE application.
# target.externalId is the Okta app.id; target.type is always APPLICATION.
curl -s -X POST \
'https://YOUR_OKTA_DOMAIN/governance/api/v1/entitlement-bundles' \
-H 'Authorization: Bearer <TOKEN>' \
-H 'Content-Type: application/json' \
-d '{
"name": "salesforce-support-tier2",
"description": "Tier 2 support: case edit, no org setup",
"target": { "externalId": "YOUR_APP_ID", "type": "APPLICATION" },
"entitlements": [
{ "id": "<ENTITLEMENT_ID>", "values": [ { "id": "<VALUE_ID>" } ] }
]
}'
# -> 201 with status: ACTIVE and an orn:
# orn:okta:idp:<ORG_ID>:entitlement-bundles:<BUNDLE_ID>
# 3. THE TRAP. Listing bundles omits their contents unless you ask for them.
# Without include=full_entitlements your inventory looks empty but valid.
curl -s -X GET \
'https://YOUR_OKTA_DOMAIN/governance/api/v1/entitlement-bundles?include=full_entitlements&limit=200' \
-H 'Authorization: Bearer <TOKEN>'
# limit range is 1..200 (default 20); page with the `after` cursor (enb...).
# Watch X-Rate-Limit-Remaining -- the reference illustrates a ceiling of 60.
# Scope the sweep to one app, or to what changed since the last run:
# ?filter=target.externalId eq "YOUR_APP_ID" AND target.type eq "APPLICATION"
# ?filter=lastUpdated gt "2026-09-01T00:00:00Z"
# ?filter=status eq "ACTIVE" (name supports eq, co and sw)
# Query parameters require percent-encoding.
Doc-verified against the Okta Governance API reference for Entitlements and Entitlement Bundles and the Okta Entitlement Management product docs, Oct 2026: endpoint paths, OAuth scopes, the APP_ADMIN role, request and response field names, status values, the ORN shape, the include=full_entitlements parameter, the 1β200 limit range with a default of 20, the after cursor pattern and the documented filter operators. Composed from the reference's own examples with placeholders; not executed against a tenant in this run.
π§ Recall
Nine days ago the Okta ISPM pill made the case that discovery is the step before governance β that ISPM reconstructs the non-human identities actually present in your apps rather than the ones you provisioned. Put the two models side by side: ISPM reports an account holding admin privilege in a governed SaaS app, and your Entitlement Management data shows that same user with no admin entitlement. Which record is wrong, and what does the answer tell you about what a certification campaign can and cannot attest?
Show answer
Neither is wrong; they are answering different questions, and that is the finding. Entitlement Management is a record of grants Okta made β a policy rule fired, a request was approved, an admin assigned a value β so its answer to "does this user have admin?" is really "did we grant admin?". ISPM answers "what does the app say is true?". A privilege created inside the app's own console by a local admin is real, effective, and has no Okta-side grant to certify, so it is invisible to the entitlement model and will not appear in a campaign for a reviewer to revoke. The practical consequence is that a certification campaign attests to the governed subset, and your control description has to say so. Two design moves follow. First, treat app-local administration as the privileged path it is: if people can create permissions inside the app, that app's local admin role belongs under privileged access with the same scrutiny as a server, otherwise your entitlement model is advisory. Second, this is exactly the gap Okta is filling with Automated Drift Detection & Remediation (early access by Q4 2026) β detection with a lag, not prevention β so the interim answer for an auditor is a reconciliation cadence you can evidence, not a claim of completeness. Write the number down: how long can an ungoverned admin grant exist inside a Tier 1 app before something notices? Today, for most tenants, that number is "until the next campaign."
πΌ Market Signal
The buying centre for this moved in September, and it moved toward fine-grained permissions rather than app assignment. Okta's product innovations announcement, Sep 2026 puts four Identity Governance capabilities at early access by Q4 2026 β Advanced Entitlement Management for AWS, Intelligent Request Recommendations ("AI-powered intelligent insights backed by request history, usage patterns, and resource sensitivity data to access request approvers"), Automated Drift Detection & Remediation, and 48-Hour Integrations β while the Privileged Access set, including Unified Database Security, Dynamic Kubernetes Protection, network device coverage and "machine-speed" authorization for CI/CD pipelines and AI agents, is stated as GA by Q1. The same announcement names the Permiso Security acquisition as the route to threat detection across human, non-human and AI identities. Two signals worth separating. The roadmap one: entitlements, drift and privileged access are being sold as one story, which means the buyer is no longer an IAM admin but whoever owns audit outcomes across cloud and SaaS. The evidence one is quieter and more reliable, because it is shipped rather than announced β the 10 Sep 2026 limit increases in the Identity Governance API release notes, Sep 2026 only matter to customers who already hit 100 bundles in a single access level condition. Ceilings get raised when paying customers are pressed against them. Career read: the scarce skill is not operating the console, it is making the modelling call β which permissions become entitlements, which entitlements become bundles, who may hold a Custom assignment and under what expiry β and then defending that model to an auditor who will ask what it does not cover. That work is a fixed-scope engagement with a written deliverable, which is the shape fractional architecture actually sells in, and the early-access window through Q4 is when design opinions are worth the most, because nobody has a reference implementation yet.
β‘ Action This Week
Pick one governed app β the one with the most entitlements, not the one with the most users β and produce a two-page entitlement model review. Pull its entitlements and values, then list every bundle with include=full_entitlements so you are reading contents rather than names. Then answer four questions in writing. Which entitlements are single-value and which are arrays, and therefore where does a grant silently replace a prior permission rather than add to it? Which users are on a Custom assignment method, when did that happen, and who decided? What is the break-glass path for this app, given that bundles cannot be assigned from the Admin Console? And what percentage of the app's discovered privileged values map to an Okta-side grant with a traceable reason? That last number is your real governance coverage, and it is almost never 100%. Keep it inside two hours by capping the sweep at one app and 200 bundles. Done = a two-page review that names the single-value entitlements where "grant" means "replace", lists every user on a Custom assignment with a date, and states one coverage percentage you are willing to show an auditor. The portfolio artifact is the model, not the export: a redacted one-page diagram of how one real app's permissions map to entitlements, bundles and policy rules β with the ungoverned seam drawn in β is a better interview exhibit than any certification, because almost nobody can draw one from memory.
Turn-Taking Is the Product: Semantic VAD, Barge-In, and the One Field Your Voice Transcript Depends On
π‘ Key Concept
Callers do not judge a voice agent on answer quality. They judge it on whether it knows when they have stopped talking. That single decision β end of turn β is the product, and in the Realtime API it is a configuration choice between two fundamentally different theories of what a turn is. Per the OpenAI Realtime API voice activity detection guide, Oct 2026, the default server_vad "automatically chunks the audio based on periods of silence," tuned by three numbers: a threshold from 0 to 1, where "a higher threshold will require louder audio to activate the model", prefix_padding_ms for the audio retained before speech is detected, and silence_duration_ms, the "duration of silence (in milliseconds) to detect speech stop." The alternative, semantic_vad, "uses a semantic classifier to detect when the user has finished speaking, based on the words they have uttered." End of turn stops being an acoustic property and becomes an inference about meaning.
Why that is architectural rather than a tuning preference: no single silence timeout is correct for the utterances one call actually contains. "Yes" ends in 300 ms. A sixteen-digit card number read aloud contains three pauses longer than most people's silence_duration_ms. A caller searching for a word pauses longer still. With server_vad you pick one number for every utterance in the session and absorb both failure modes: raise it and every exchange feels sluggish, lower it and the agent talks over people at exactly the moments that carry the information. semantic_vad replaces the number with a policy β a single knob, eagerness, taking "low" | "medium" | "high" | "auto", defaulting to "auto" which the docs equate to medium, controlling "how eager the model is to interrupt the user, tuning the maximum wait timeout." The mechanism is explicit: "when the probability is low, the model will wait for a timeout, whereas when it is high, there is no need to wait." Note what moved. The wait is now variable and content-dependent, which is what you wanted, and it sits inside your latency budget, which is what you now have to measure. One more constraint shapes designs early: create_response and interrupt_response are documented as conversation-mode only, so a team that drives its own turn loop does not get the automatic respond-and-interrupt behaviour and must rebuild it.
The other half of turn-taking is interruption, and it hides the detail that actually matters for anyone shipping regulated voice. The Realtime client events reference, Oct 2026 defines conversation.item.truncate, which "cuts assistant message audio at specified duration," carrying item_id, content_index and audio_end_ms; alongside it response.cancel stops an in-progress response, and output_audio_buffer.clear stops audio generation β documented as WebRTC and SIP only. Put those together with how the session streams: the Realtime guide for gpt-realtime-2.1, Oct 2026 notes the session "connects over WebRTC in the browser or WebSocket on the server" and emits response.output_audio.delta alongside response.output_audio_transcript.delta. The model generates a complete turn and a complete transcript of that turn. If the caller cuts in after four words, the server's conversation state β and therefore the transcript you archive β still contains the whole sentence unless your client reports where playback actually stopped, in audio_end_ms. Your client measures that number. Which means the only reconciliation between what the model said and what a human heard is a millisecond value supplied by the least trustworthy component in the stack.
Cross-pollination: today's IAM pill is about Okta entitlements where granting a single-value permission silently replaces the previous one and nothing in the record marks the replacement. Same class of defect: a later write supersedes an earlier one, the supersession is never written down, and the system's history stops describing reality.
π¬ Deep Dive
Practitioner trap: handling barge-in by stopping playback locally and calling response.cancel, without conversation.item.truncate. Everything looks right in the room β the audio stops, the caller is heard, the next turn works β and the conversation state is now wrong in a way that compounds. The server believes the assistant delivered its whole turn, so the model will not repeat the part the caller never heard, will reason as though the disclosure landed, and will produce a transcript asserting it. The fix is the event's three fields: item_id, content_index and audio_end_ms, "cuts assistant message audio at specified duration." The second half of the trap is that the same bug has two different fixes depending on transport, because output_audio_buffer.clear is documented as WebRTC and SIP only β on a server-side WebSocket session there is no server-side buffer to clear, so discarding already-streamed audio is entirely your client's job, and a team that develops on WebRTC and ships over WebSockets will reintroduce the defect at the worst possible moment. Third, the smallest and most expensive knob: prefix_padding_ms controls how much audio before detected speech is kept. Set it too low and you clip the first phoneme, which is precisely where single-syllable answers live β "yes" against "no", and the consent turn is the one utterance in the call you cannot afford to mis-hear.
Staff-level framing: turn detection is the only component of the latency budget that is a deliberate choice, so it is the one that needs an owner and a documented rationale. The OpenAI voice agents guide, Oct 2026 gives the instrumentation instruction plainly β "measure how long callers wait for a useful spoken answer. Track backend time separately to find delays" β and the separation matters because the two halves have different owners: backend time is an engineering problem, the turn-detection wait is a product policy. Make that explicit in the design review, because eagerness is not a performance setting, it is a distributional one. Raising it shortens waits for fluent, fast speakers and increases false interruptions for everyone who pauses β older callers, non-native speakers, people reading a document, people who are distressed. That is a complaint-rate and fairness question with a named parameter attached to it, which is rare and worth exploiting: you can actually A/B it per cohort and show the result. Then there is the governance half, which almost nobody designs for before their first audit. In regulated voice, the evidence that a disclosure, recording notice or consent prompt was delivered is not the model's transcript β the model always believes it finished speaking. The evidence is the truncation record, assembled from a client-measured millisecond. Treat it accordingly: log audio_end_ms against generated duration for every interrupted turn, keep it with the call recording rather than in application logs with a 14-day retention, and decide deliberately whether a turn that was cut off before the disclosure completed forces a re-read. That decision is a compliance control, and it belongs in the design, not in a prompt.
Choose the architecture from the evidence you need, not the demo quality. The voice agents guide frames the two options honestly. Speech-to-speech puts "speech, reasoning, and tool use in one session" with "one model to interpret audio, decide what to do, and respond in speech" β fewer hops, better prosody, and the model hears tone, hesitation and overlap rather than a flattened transcript. The chained pipeline gives "control over each speech and text stage" and is what you want when you "want to inspect or transform text between speech recognition, your agent, and speech generation", including the ability to "store the transcript" as a first-class artifact; the guide's chained example names a text model, gpt-6-astra, in that middle position. Read the tradeoff as an evidence tradeoff rather than a latency one. If you must redact card numbers before they reach a model, route by detected language, keep a verbatim human-readable record, or swap the reasoning model without re-certifying the voice, the seam between stages is the requirement and chaining pays for itself. If the product value is that the agent sounds like it is listening β interruptible, responsive, able to say "mm-hm" β the single session is the one that delivers it, and you buy back observability by instrumenting turn boundaries instead of text boundaries. A concrete metric set for either: time from last caller audio to first assistant audio; truncations per call; audio_end_ms as a fraction of generated audio duration, which is simultaneously your unheard-content rate and your wasted-generation rate; and false-interruption rate split by caller cohort. Those four numbers tell you more about a voice deployment than any eval suite, and none of them exist by default.
# Two theories of "the turn ended", and the event that keeps your
# transcript honest when the caller disagrees.
# A. server_vad -- end of turn is a silence timer. One number for every
# utterance in the session: "yes" and a 16-digit card number alike.
{
"type": "session.update",
"session": {
"type": "realtime",
"model": "gpt-realtime-2.1",
"audio": {
"input": {
"turn_detection": {
"type": "server_vad",
"threshold": 0.5, # 0..1; higher = needs louder audio
"prefix_padding_ms": 300, # audio kept BEFORE speech onset
"silence_duration_ms": 500 # silence that means "speech stop"
}
}
}
}
}
# B. semantic_vad -- end of turn is a classifier over what was said.
# eagerness tunes the MAXIMUM WAIT, it is not a latency setting.
{
"type": "session.update",
"session": {
"audio": {
"input": {
"turn_detection": {
"type": "semantic_vad",
"eagerness": "low" # "low" | "medium" | "high" | "auto"
} # auto (default) == medium
} # low = waits longer, interrupts less
} # high = answers sooner, cuts people off
}
}
# create_response / interrupt_response exist for BOTH types but are
# documented as conversation-mode only -- own your loop, own this logic.
# C. Barge-in. The caller cut in 1.2s into a 4.8s assistant turn.
# response.cancel alone leaves the server believing it said everything.
{ "type": "response.cancel" }
{
"type": "conversation.item.truncate",
"item_id": "item_A1",
"content_index": 0,
"audio_end_ms": 1200 # <- measured by YOUR client, from playback
}
# Now server state matches the caller's ears. Omit it and the archived
# transcript asserts a disclosure that was never heard.
# D. Transport asymmetry -- same bug, two fixes:
{ "type": "output_audio_buffer.clear" } # WebRTC / SIP only
# On a server-side WebSocket session there is no server audio buffer:
# dropping already-streamed audio is entirely your client's problem.
# Log per interrupted turn: audio_end_ms / generated_ms
# = unheard-content rate AND wasted-generation rate, in one number.
Doc-verified against the OpenAI Realtime API voice activity detection guide, the Realtime client events reference and the Realtime and voice agents guides, Oct 2026: the server_vad and semantic_vad types, every parameter name, the eagerness value set and its "auto" default, the conversation-mode-only note on create_response/interrupt_response, the model id gpt-realtime-2.1, and the event type strings with conversation.item.truncate's item_id, content_index and audio_end_ms fields plus the WebRTC/SIP-only scope of output_audio_buffer.clear. The numeric values are illustrative examples, and the session.audio.input nesting follows the GA migration note on the Realtime guide; not executed against the API in this run.
π§ Recall
Nine days ago the OpenAI Agents API pill described what you rent when you stop running your own harness: a Session as "a durable instance of that agent working on a task across turns", with the provider owning orchestration, context compaction and recovery so a task can resume tomorrow. Apply that promise to a phone call. Which part of durable-session semantics stops being meaningful in a voice agent, and what does that tell you about where the authoritative record of the conversation actually lives?
Show answer
Recovery stops being meaningful, because the expensive state is not on the server. A text agent can crash mid-task and resume because everything it had done is replayable from the session: items in, items out, checkpoint restored. A voice turn has already left the building β the audio played into a human ear, and no amount of durable state can un-hear it or re-deliver it with the same meaning. Resumption, the headline feature of a managed harness, degrades to "the caller is still on the line and will tolerate silence", which is a budget measured in seconds. That inverts where the authoritative record lives. For a text agent the server-side session is the truth. For a voice agent the truth is what the caller heard, and the server only approximates it using the one number the client sends back on interruption. Two consequences worth carrying into a design review. First, idempotency moves to the tool layer and nowhere else: a voice session that retries a transfer, a payment or an account change after a reconnect has no way to know whether the caller already heard the confirmation, so every side-effecting tool needs a caller-visible idempotency story, not just a technical one. Second, "durable session" and "auditable conversation" are different products. You can buy the first; you have to build the second, out of truncation records, call recordings and turn-boundary telemetry that your platform does not produce for you.
πΌ Market Signal
The useful market fact about voice agents is that the most-cited forecast for this technology named this year, and it is now checkable. Gartner's press release of 31 Aug 2022 stated that "by 2026, conversational artificial intelligence (AI) deployments within contact centers will reduce agent labor costs by $80 billion", against a population of "approximately 17 million contact center agents worldwide" where labour "can represent up to 95% of contact center costs". Read the second number in that release more carefully than the headline, because it is the one that describes the work: Gartner projected "one in 10 agent interactions will be automated by 2026, an increase from an estimated 1.6% of interactions today." One in ten is a modest target for a technology that demos as finished β and the release said why, in 2022, before the models were good: "implementing conversational AI requires expensive professional resources in areas such as data analytics, knowledge graphs and natural language understanding," with integration priced at "$1,000 to $1,500 per conversational AI agent, though some organizations cite costs of up to $2,000 per agent" and early adoption "primarily led by organizations with 2,500 or more agents". Four years on, the model problem is largely solved and that integration bill has not gone anywhere; it has only changed shape, from intent trees and knowledge graphs to turn-taking policy, barge-in correctness, transcript evidence and telephony integration. Career read: this is a market that is bottlenecked on exactly the judgement calls in this pill, and it has an unusual property for an AI niche β the buyer already has a hard number for what a bad call costs, because they have measured containment, handle time and complaint rates for two decades. You do not have to sell them on value, only on a design. Position on the parts vendors do not ship: the latency budget with the turn-detection wait named as a product decision, the interruption contract, the evidence trail for disclosures, and per-cohort false-interruption measurement. That is a two-week assessment with a written deliverable, in a sector that buys professional services by default.
β‘ Action This Week
Build or borrow the smallest possible voice agent and run a twenty-call experiment that produces one table. Ten calls with server_vad at its defaults, ten with semantic_vad and eagerness: "low", and keep the script fixed: a greeting, one short confirmation ("yes"), one long structured utterance read aloud with natural pauses β a card number, a policy reference, an address β and one deliberate interruption of the agent mid-sentence. Record three things per call: time from your last audio frame to the first assistant audio frame, the number of times the agent cut you off, and whether your client emitted conversation.item.truncate with a real audio_end_ms on the interruption. If you have no agent to hand, the experiment still works against any vendor demo line you can legally call β you are measuring turn behaviour, not code. Done = a two-column table of time-to-first-audio and cut-off counts for the two VAD modes on identical utterances, plus a one-line verdict on whether your own implementation reports truncation or silently lies to the server. Write up the finding, not the feature: "our agent interrupted the card-number turn 7 times out of 10 on default settings, and our transcript claimed it had read the full disclosure" is a result from a real system, and it is the kind of sentence that gets a reply from someone who owns a contact centre. The explainer about semantic VAD will not.
Deferred Token Response: OAuth Makes "Wait for the Human" a Property of the Token Endpoint, Not of a Grant
π‘ Key Concept
Every human-in-the-loop story for AI agents currently rests on choosing a grant that happens to be asynchronous. If the agent needs a person to approve a payment, you adopt CIBA. If it is a device with no browser, you adopt the Device Authorization Grant. The asynchrony is welded to the grant, so "this action needs a human" becomes an architectural commitment made months earlier, at integration time, for the whole client. On 16 Sep 2026 the IETF OAuth working group published draft-ietf-oauth-deferred-token-response-00, Sep 2026 as a WG document, and it unwelds them. Its abstract is the whole design: "In existing OAuth grants, the token endpoint either issues an access token or returns an error. DTR establishes a generic asynchronous token request mechanism that any OAuth grant may plug into." The token endpoint gains a third possible answer β not yet β and every grant you already use inherits it.
The direction of initiation is the part worth slowing down for, because it inverts who decides. The draft is explicit that "CIBA is initiated by the client. DTR is not initiated by a client at all: an existing grant is initiated by its normal means and the authorization server elects to defer the resulting token response," and likewise that "unlike the Device Authorization Grant, deferral under DTR is initiated by the authorization server, not by the client or the end-user." Read that as a control-plane change, not a convenience feature. The authorization server β your policy engine β becomes the component that decides a given token request needs out-of-band resolution, and it can decide that per request, on live signals: this agent, this scope, this amount, this hour, this risk score. Mechanically it answers the token request with error: "authorization_pending" plus a deferral_code, an expires_in and a polling interval; the client redeems the code later at the same token endpoint under the new grant type urn:ietf:params:oauth:grant-type:deferred, or is called back on a notification endpoint if one is configured.
The Okta lineage here is direct and worth knowing, because it tells you where this lands in your tenant. DTR replaces draft-parecki-oauth-jwt-grant-interaction-response, 2026 β Aaron Parecki (Okta), Brian Campbell (Ping Identity) and Dapeng Liu β which posed the same problem narrowly, as an extension to the RFC 7523 JWT authorization grant, returning "a URI that the client launches where the user can interact with the authorization server, along with a polling interval." That JWT-grant family is precisely the machinery behind Okta's Cross App Access and ID-JAG token exchange, so the original framing was "what if an agent's token exchange needs consent mid-flight". The WG generalised it away from JWT grants entirely. Today Okta ships this capability product-side rather than protocol-side: per Auth0 for AI Agents asynchronous authorization docs, Sep 2026, the agent backend calls /bc-authorize, receives an auth_req_id, polls /token, and the user approves on a Guardian push, SMS or email β CIBA, client-initiated, with the Rich Authorization Request payload supplying the consent context. DTR is the standards-track version of the same user experience with the decision moved to the server and the grant left alone.
Cross-pollination: today's AI pill is about routing inference requests to the replica that already holds the matching KV cache. Both mechanisms are the same shape β a short-lived opaque handle (deferral_code there, a prefix hash here) that binds a follow-up request to the one server holding the state for it, because stateless round-robin throws that state away.
π¬ Deep Dive
Practitioner trap: teams will treat the deferral_code as a request id and log it. The draft closes that door in one sentence: "Deferral codes are sender-constrained per [Section 10.1], inheriting the binding rules... for refresh tokens. Its threat profile is the same as a refresh token of equivalent lifetime." Equivalent lifetime is not short β the draft's own example pairs expires_in: 10800 (three hours) with interval: 60. So an agent platform that parks pending approvals in a work queue, a Redis key, a Temporal workflow payload or a CloudWatch log line has just spread refresh-token-grade credentials across four systems with different retention and different readers, to hold an authorization that is about to become a payment. Treat the pending-approval store as a credential store: encrypted at rest, no logging of the value, access audited, and the record deleted on terminal state. The second half of the trap is the polling loop. Alongside authorization_pending the spec defines slow_down ("the client is polling faster than interval allows"), expired_token and access_denied. Most agent frameworks retry on non-2xx with fixed or exponential backoff and no notion of a server-dictated floor, so the first thing your integration will do in production is get rate-limited by your own IdP β and because access_denied is also a non-2xx, a naive retry loop will keep polling a decision the human already refused.
Staff-level framing: the governance value is that approval becomes a server-side policy decision instead of a client-side code path, and the governance risk is exactly the same sentence. Once the authorization server can defer any grant, a policy change can put an interaction step in front of a client that was written years ago and never expected one. The interlock is the client's opt-in: completion_mode, "a space-separated list of completion-mode values," carrying deferred to signal the client will accept a deferred response, and if it is "absent or lacking deferred, the client requires synchronous handling." That makes your rollout plan concrete and boring in the right way β inventory which clients send it, stage policy by client rather than by scope, and discover server capability from the metadata parameter the draft registers, deferred_token_response_supported. Two more items belong in the design review. First, the kill switch: a pending authorization is cancellable at the revocation endpoint with the deferral code as token and token_type_hint of urn:ietf:params:oauth:token-type:deferral-code, after which the server "suppresses pending callbacks, and causes subsequent polling to return access_denied" β that is the incident-response primitive for "an agent is mid-approval and we want it stopped", and it needs an owner and a runbook, not a ticket. Second, the audit story an auditor will actually ask for: which human approved, against which displayed context, at what time, and how long the request sat pending β DTR carries the plumbing, your IdP's consent UI carries the evidence, and they are logged in different places.
What DTR deliberately does not answer: how the human is reached. This is the gap to plan around, and it is where the predecessor draft and the WG draft differ in a way worth reading carefully. The Okta-authored JWT-grant version returned "a URI that the client launches" β the client had somewhere to send the user. The WG's DTR-00 deferred response I verified carries deferral_code, expires_in, interval and an optional error_description; reaching the approver is left to the authorization server, and optionally back to the client through a callback authenticated with the client_notification_token the client supplied on the initial request β the server "MUST authenticate the callback request by including the client_notification_token as a Bearer credential," POSTing {"deferral_code": "..."} and expecting 204. That division is correct layering and an operational cliff: the protocol standardises waiting, not notifying. Which is why CIBA does not become obsolete β it answers the notification question concretely. Auth0's async authorization requires the CIBA grant enabled on the application and Guardian push MFA enrolled for the approving user, and surfaces the decision to agent code as a withAsyncAuthorization() wrapper around a tool call that raises AccessDeniedInterrupt on refusal. The honest reading for an architect in 2026: build the human-reachability path on what your IdP ships today, and design the waiting mechanism so the grant-agnostic version can replace it without touching your agent's tool definitions.
# Deferred Token Response, end to end. Values are the draft's own examples.
# The point: request 1 is an ORDINARY grant. Nothing about it is CIBA-shaped.
# 1. Normal token request -- plus the one new thing the client must send.
# Without completion_mode=deferred the server MUST answer synchronously.
POST /token HTTP/1.1
Host: server.example.com
Content-Type: application/x-www-form-urlencoded
grant_type=urn:ietf:params:oauth:grant-type:token-exchange
&subject_token=<AGENT_SUBJECT_TOKEN>
&scope=payments:initiate
&completion_mode=deferred
&client_notification_token=f4oirNBUlM # optional: enables the callback
# 2. The server elects to defer. HTTP error shape, but NOT a failure.
{
"error": "authorization_pending",
"deferral_code": "8d67dc78-7faa-4d41-aabd-67707b374255",
"expires_in": 10800,
"interval": 60
}
# ^ 3 hours of validity, 60s minimum between polls. Treat deferral_code
# as a refresh token: encrypted store, never logged, deleted on terminal state.
# 3. Redeem it later, at the SAME token endpoint, under the new grant type.
POST /token HTTP/1.1
Host: server.example.com
Content-Type: application/x-www-form-urlencoded
Authorization: Basic czZCaGRSa3F0MzpnWDFmQmF0M2JW
grant_type=urn:ietf:params:oauth:grant-type:deferred
&deferral_code=8d67dc78-7faa-4d41-aabd-67707b374255
# -> authorization_pending | slow_down | expired_token | access_denied
# -> or the originating grant's own success response:
{"access_token":"SlAV32hkKG","token_type":"Bearer","expires_in":3600}
# 4. Instead of polling: the server calls YOU, bearing your notification token.
POST /deferred-callback HTTP/1.1
Authorization: Bearer f4oirNBUlM
Content-Type: application/json
{"deferral_code": "8d67dc78-7faa-4d41-aabd-67707b374255"}
# expected response: 204 No Content
# 5. The kill switch. Cancel a pending approval before the human answers.
POST /revoke (token_type_hint is RECOMMENDED, not required)
token=8d67dc78-7faa-4d41-aabd-67707b374255
token_type_hint=urn:ietf:params:oauth:token-type:deferral-code
# Server cancels, suppresses pending callbacks, polling now -> access_denied.
# Discovery, before you build any of it:
# GET /.well-known/oauth-authorization-server
# -> "deferred_token_response_supported": true
Doc-verified against draft-ietf-oauth-deferred-token-response-00 (IETF OAuth WG, 16 Sep 2026): parameter names, the urn:ietf:params:oauth:grant-type:deferred and urn:ietf:params:oauth:token-type:deferral-code URNs, error codes, the callback shape and the metadata parameter, Sep 2026. Composed from the draft's examples; placeholders only, not executed in this run. This is a -00 working-group draft, not a deployable Okta feature β expect parameter churn.
π§ Recall
Five days ago the Okta disaster recovery pill argued that in a read-only failover, expiry is the only revocation primitive that still functions, because every actual revoke is a write. Apply that lens to the mechanism above: what happens to a queue of pending agent approvals when the identity control plane goes read-only, and which of DTR's four polling errors becomes the one you cannot produce?
Show answer
You lose the kill switch and keep the clock. Cancelling a pending authorization is a write β the draft routes it through the revocation endpoint and has the server transition the request to a cancelled state, suppress callbacks and return access_denied on the next poll. In a mutation freeze none of that can happen, so access_denied is precisely the outcome you can no longer cause, while expired_token still arrives on its own because expiry needs nobody's permission. That is the same asymmetry the DR pill described one layer up, and it has a design consequence you can act on now: the expires_in you choose for a deferral code is not a UX parameter, it is the worst-case window during which a pending approval you have decided to stop cannot be stopped. Three hours of validity means three hours of exposure in any state β regional failover, IdP degradation, a revocation endpoint returning 500 β where your only remaining control is waiting. Pick the shortest lifetime the approving human can realistically meet, and treat any pending-approval TTL longer than your incident-response target as a finding.
πΌ Market Signal
The vendor side of this is now shipping, not previewing. At Oktane on 22 Sep 2026, per Okta's Oktane 2026 AI announcements, Sep 2026, Agent SSO, Agent-to-Agent Connections and Resource Access Certifications went generally available, with Agent Gateway and Shadow AI Agent Discovery for Endpoints slated for Q3 2026 and a Configuration Designer plus an expanded runtime-layer kill switch for Q4 2026 β and a multi-vendor Blueprint Alliance organised around four questions that read like a job description: where are my agents, what can they do, what are they doing, how do I respond. Note what the roadmap implies about sequence: the runtime kill switch is the last thing to arrive, which is the same gap DTR's cancellation path fills at the protocol layer, and it is the gap most agent deployments have right now. The demand signal underneath is unflattering and useful: Okta's own materials cite Gravitee's State of AI Agent Security 2026 finding that 88% of organizations report suspected or confirmed AI agent security incidents while only 22% treat AI agents as independent, identity-bearing entities (Okta citing Gravitee, Feb 2026). That 66-point gap is the market. Career read: the scarce profile is not "knows CIBA" β it is the person who can stand in front of a risk committee and explain, with wire-level specifics, what happens between an agent deciding to act and a human authorising it: who is asked, how long the request may sit, who can cancel it, what is logged, and what degrades when the IdP is unhealthy. That conversation is unbuyable as a certification and directly billable as fractional architecture work, because the teams shipping agents in 2026 have a product roadmap for it and no design.
β‘ Action This Week
Pick one agent or automation in your estate that acts with a human's authority β a provisioning bot, an approval assistant, anything that moves money or grants access β and write its one-page pending authorization contract. Five lines, nothing more: who approves (a named role, and the fallback when they are asleep); what the approval prompt displays (the concrete binding message, not "approve action"); the maximum time a request may sit pending, justified against your incident-response target; who can cancel a pending request and through what interface; and what the agent does on denial and on timeout β distinct paths, because they usually are not. Then do the cheap empirical half: grep your agent's code for its retry logic and check whether it can honour a server-supplied polling interval and stop on a terminal denial. If you have an Auth0 tenant, wrap exactly one tool call in withAsyncAuthorization() with a real bindingMessage and watch what the Guardian push actually says to the approver. Done = a one-page contract whose pending-approval TTL is a number you can defend, and a one-line finding on whether your agent's retry loop would survive slow_down. The LinkedIn post writes itself from the finding, not the feature: "our agent could ask for approval and could not be told to stop asking" is a result from your own estate, which outperforms any explainer of a -00 draft.
Round-Robin Is Deleting Your KV Cache: Inference-Aware Routing With InferencePool
π‘ Key Concept
A Kubernetes Service treats your model-server replicas as interchangeable. They are not. Each replica holds a KV cache, and that cache is the single most valuable thing in your inference fleet: it is the difference between re-running prefill over a 4,000-token system prompt and RAG context, or skipping it. When turn two of a conversation lands on a different replica than turn one, nothing is broken β you simply pay full price for work you already did, and you pay it in the phase that is compute-bound. With N replicas and connection-level balancing, the probability of landing on the replica that holds your prefix is roughly 1/N. Scale the fleet to reduce latency and you make cache affinity worse. That is the structural bug the Kubernetes Gateway API Inference Extension, Sep 2026 exists to fix, and it fixes it by changing what a backend is.
The resource is InferencePool, at apiVersion: inference.networking.k8s.io/v1. It takes the place of a Service in the Gateway API model β an HTTPRoute names it directly in backendRefs with group inference.networking.k8s.io and kind InferencePool β and it carries three things: a selector whose labels "must exactly match the labels applied to your model server Pods", targetPorts, and the field that does the actual work, endpointPickerRef. That reference points at an Endpoint Picker (EPP), a service that "monitors key metrics from model servers within the InferencePool" and is consulted per request to choose the pod. Routing stops being a property of the load balancer's connection table and becomes a scheduling decision made with live knowledge of each replica's internal state. The companion resource InferencePoolImport extends the same idea across clusters, as "a cluster-local, controller-managed representation of an imported InferencePool from another cluster".
What the picker scores on is where the architecture becomes concrete, and the implementations agree. Google's GKE Inference Gateway documentation, Sep 2026 names three signals for inference-optimized load balancing β "KV cache hits: the number of successful lookups in the key-value (KV) cache", "GPU or TPU utilization", and "Request queue length: the number of requests waiting to be processed" β and describes prefix-cache-aware routing as sending requests that share context to the same replica by "maximizing cache hits", which "dramatically reduces redundant computations and improves Time-to-First-Token", singling out conversational AI and RAG as the workloads that benefit. Istio's implementation describes the same picker evaluating request criticality, GPU memory and KV cache utilisation, LoRA adapter affinity, prefix-cache awareness and queue depth. Note the blend: pure prefix affinity would be a hotspotting machine, sending every request sharing a popular system prompt to one overloaded pod, so the signals are weighed against each other. Affinity is a preference, not a routing key.
Cross-pollination: today's IAM pill covers OAuth's Deferred Token Response, where a deferral_code ties a later poll back to the one authorization server holding that pending decision. Same primitive, different plane β a handle that says "this request must return to the machine that already has the state", because the alternative is recomputing or losing it.
π¬ Deep Dive
Practitioner trap:failureMode on endpointPickerRef is the most consequential field in the manifest and it looks like boilerplate. The documented example ships failureMode: FailOpen. Read what that buys you: if the EPP is unreachable, traffic still flows β and flows without inference-aware scoring, which means you have silently reverted to the 1/N behaviour this whole mechanism exists to avoid. Nothing pages. Pods are Ready, the gateway returns 200s, error-rate dashboards are flat, and the only symptom is that TTFT and accelerator spend get quietly worse, which is exactly the class of regression that survives a whole quarter. FailClose has the opposite failure: your inference pool becomes unavailable because a sidecar-shaped scheduler died. There is no free choice here, only an explicit one, and both choices oblige you to monitor the EPP as a first-class dependency with its own SLO and its own alert β cache hit rate and TTFT percentiles are the signals that actually detect a degraded picker, not pod readiness. While you are in there, note the second-order cost: prefix-cache-aware routing works by "analyzing the request body", so your gateway now parses and hashes request bodies on the hot path. Body size limits, buffering behaviour and streaming semantics move from ingress trivia to capacity-planning inputs.
Staff-level framing: this reframes inference cost as a routing problem, and that is a conversation you can have with a CFO. Most teams treat unit economics as a model-selection and quantisation exercise β smaller model, fewer bits, cheaper token. Cache affinity attacks a different term: the prefill you never run at all. For a RAG or agent workload where a long system prompt and retrieved context dominate the input, the shared prefix can be the overwhelming majority of input tokens per request, so the routing decision determines whether you pay for those tokens once per conversation or once per turn β with no change to the model, the hardware or the prompt. Two more items belong in the design doc. First, autoscaling: GKE's docs describe "optimized autoscaling using HPA with model server metrics", and that phrasing is the point β scaling an LLM fleet on CPU or even raw accelerator utilisation is a category error, because a GPU at 95% with an empty queue is healthy while a GPU at 60% with a growing queue is failing your SLO. Scale on queue depth and KV cache utilisation, the same signals the picker routes on. Second, priority as a governance lever rather than a tuning knob: the same docs describe "model-specific serving Priority" to separate latency-sensitive traffic from batch. That is what lets you put nightly eval runs and interactive users on one expensive accelerator pool and be able to say, in writing, which one gets shed under contention. Concentration of a fleet behind a shared pool without an explicit priority policy is how an eval job becomes a customer-facing incident.
The standard is the reason to care, not any one vendor's gateway. The pattern is converging on one API across implementations that compete on everything else. Istio shipped support in 1.27, Jul 2025, routing to InferencePool backends through an ext_proc EPP while keeping mutual TLS, access policies and distributed tracing intact β which matters more than it sounds, because a bespoke Python router in front of your model servers usually means opting out of your service mesh's security and observability posture at exactly the layer auditors ask about. Google productised it as GKE Inference Gateway, adding predicted latency-based routing and Model Armor safety integration on the same substrate. And llm-d, a CNCF Sandbox project, Sep 2026, builds directly on it β "a high-performance L7 proxy conformant with the Gateway API Inference Extension (GAIE)" β describing its EPP as consulting "real-time metrics, KV-cache affinity, and configured policies", offering "heuristic and precise techniques to maximize cache hits", and layering prefill/decode disaggregation on top so "a single inference request is split into multiple phases... handled by specialized workers". The architectural reading: cache affinity and phase disaggregation are the same insight applied twice. Stop pretending inference is stateless, then route according to where the state actually lives. The portable skill is knowing which signals a scheduler needs and what each one costs to collect β that transfers across every implementation; a vendor's YAML does not.
# The whole mechanism is two objects. Note what is NOT here: no Service.
---
apiVersion: inference.networking.k8s.io/v1
kind: InferencePool
metadata:
name: vllm-qwen3-32b
spec:
selector:
matchLabels:
app: vllm-qwen3-32b # MUST match the model-server Pod labels
targetPorts:
- number: 8000 # the port on the model server pods
endpointPickerRef:
name: vllm-qwen3-32b-epp # the scheduler consulted per request
port:
number: 9002
failureMode: FailOpen # <-- decide this deliberately:
# FailOpen = traffic flows unscored
# (silent loss of affinity)
# FailClose = pool unavailable if EPP dies
---
# An HTTPRoute points at the pool the way it would point at a Service.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: llm-route
spec:
parentRefs:
- name: inference-gateway
rules:
- backendRefs:
- group: inference.networking.k8s.io
kind: InferencePool
name: vllm-qwen3-32b
# A route MAY reference multiple InferencePools as backendRefs.
# What to alert on once this is live -- NOT pod readiness:
# 1. EPP availability (FailOpen hides its death completely)
# 2. KV cache hit rate (the thing you bought; it can silently drop)
# 3. TTFT p50/p99 (where a lost cache hit actually shows up)
# 4. request queue length (the HPA signal, not accelerator utilization)
Doc-verified against the Gateway API Inference Extension InferencePool API documentation β inference.networking.k8s.io/v1, the selector/targetPorts/endpointPickerRef fields and the HTTPRoute backendRefs group and kind, Sep 2026. The HTTPRoute wrapper is assembled from that documented reference pattern; not applied to a cluster in this run.
π§ Recall
Five days ago the MCP Skills pill covered SEP-2640, where a user's approval of a skill is bound to the exact set of file digests seen at approval time β but the spec notes a host "need not poll for changes", so a revocation only takes effect the next time the host re-fetches. Today's pill has a component that also acts on a snapshot of state it does not own. Name it, and say what the equivalent staleness failure looks like.
Show answer
The Endpoint Picker. It scores replicas on metrics it scrapes from the model servers β KV cache hits, queue length, accelerator utilisation β and every one of those is a reading from a moment slightly in the past. The staleness failure is a stampede: if the picker's view of queue depth is a few seconds old, a burst of arrivals all get scored against the same snapshot, all prefer the same "idle" replica, and the picker cheerfully builds the hotspot it exists to prevent β then over-corrects when the next scrape lands. Both cases share a shape worth carrying around: a control that is correct with respect to the state it last observed, and wrong with respect to the state that now exists, with the error growing in proportion to the observation interval. The design responses rhyme too. SEP-2640's answer is to shorten the gap between fetches; a picker's answer is to shorten the scrape interval, blend several signals so no single stale one dominates, and add hysteresis so decisions do not all swing together. The generalisable Staff-level question to ask of any scheduler, authorization cache or feature flag: how old is the newest fact this thing is allowed to act on, and what happens in the window before it learns better?
πΌ Market Signal
Watch where the competing implementations landed, because vendor convergence on a Kubernetes API is the most reliable adoption signal in infrastructure. Istio made the extension officially available in 1.27, Jul 2025; Google ships it as a managed product with GKE Inference Gateway, Sep 2026, extending it with predicted latency-based routing, dynamic LoRA serving on shared accelerators, HPA integration and Apigee-based API management; and llm-d entered the CNCF Sandbox, Sep 2026 as a GAIE-conformant stack adding prefill/decode disaggregation. When a service mesh, a hyperscaler and a CNCF project all implement the same InferencePool contract rather than shipping proprietary routers, the skill stops being vendor trivia and becomes a platform competency β the same trajectory ingress took a decade ago, when knowing the Ingress and later Gateway API mattered more than knowing any one controller. Career read: this is the seam where "ML engineer" and "platform engineer" stop being separate roles, and very few people sit in it. Most ML practitioners can fine-tune and evaluate but cannot explain why scaling on GPU utilisation misprices an LLM fleet; most platform engineers know Gateway API cold but have never thought about a KV cache. Being credibly fluent in both β able to walk into a team burning six figures a month on accelerators and show that its routing layer is discarding cache hits β is a Staff-level positioning that maps directly onto fractional and consulting work, because the finding pays for the engagement in the first week.
β‘ Action This Week
Quantify what your current routing throws away β no cluster changes, under two hours, and it works on logs you already have. Take one multi-turn or RAG workload and pull a day of request logs grouped by conversation or session id. Compute two numbers. First, shared prefix ratio: for each request after the first in a session, how many input tokens are identical to the previous request's prefix (system prompt plus retrieved context plus history) as a fraction of that request's input tokens β a rough token count is fine, precision is not the point. Second, affinity loss: with your current replica count N and connection-level balancing, the expected share of those follow-up requests that land on a replica holding the prefix is about 1/N, so multiply the shared tokens by (1 β 1/N) to get prefill tokens recomputed for nothing per day. Price it with your accelerator hourly rate and measured prefill throughput. Done = one sentence with three numbers in it: "X% of our input tokens are shared prefix, with N replicas we recompute about Y of them per day, costing roughly Z." Bring that sentence, not an architecture diagram, to the next platform review β and it is a strong portfolio artifact precisely because it is measured rather than argued: a short write-up titled "we are paying twice for the same prefill, here is the number" travels much further on LinkedIn than an explainer about a Kubernetes CRD.
Okta Disaster Recovery: Identity Survives the Outage Read-Only, and Read-Only Is the Whole Problem
π‘ Key Concept
Every pill in this dashboard so far has treated the Okta tenant as a thing that is there β you write policy into it, mint tokens from it, revoke through it. Disaster recovery is the day that assumption is suspended, and almost nobody has read what the suspension actually looks like. Per Okta's disaster recovery documentation, every customer already has Standard DR: an active-active-active deployment across three availability zones per region, with automatic failover to a paired secondary region and a recovery time objective of roughly one hour after Okta identifies the outage. The pairings are published and geographic β OK1 runs North Virginia to Oregon, EU1 runs Frankfurt to Ireland, OK16 runs Tokyo to Osaka β which is also your data-residency answer when someone asks where identity goes during a regional failure. Enhanced Disaster Recovery is the paid add-on on top: it compresses failover to five minutes and, crucially, hands the trigger to you through a self-service admin app and an API instead of leaving it to Okta's judgement.
Here is the part that changes how you design: both tiers land in the same state β read-only. In the DR region, users continue to authenticate into their apps, and admins get the Admin Console back with read-only access. What does not work is documented and specific: users cannot reset their passwords, and authentication through an external identity provider fails, because those providers rely on redirect URIs pointing at the primary domain. Admins authenticating with locally-sourced Okta credentials still get in. Okta's GA announcement for the Enhanced DR add-on, Mar 2024, adds the number that matters more than the five minutes: full read-write access returns within 24 hours. So the honest way to read the product is that Enhanced DR buys you a faster path into a degraded mode you may sit in for a day.
Read-only is a mutation freeze, and once you say it that way the blast radius reorganises itself. Authentication β the read path β keeps running and keeps issuing sessions and tokens. Every control you own that is a write stops: revoking a token, deactivating a service app or an API token, rotating a secret, tightening a sign-on policy, deprovisioning a compromised account, onboarding an incident responder. For up to 24 hours your identity plane will keep letting principals in and will not let you take anything away. If your incident playbook's first move is "revoke in Okta", that playbook has an unstated dependency on the control plane being writable, and DR is precisely the window where it is not.
Cross-pollination: today's AI pill covers the MCP Skills extension, whose spec binds a user's approval of a skill to the exact set of file digests seen at approval time β but explicitly notes a host "need not poll for changes", so revocation only takes effect the next time the host re-fetches. Same failure class in a different plane: a control that exists on paper, executes only when something else chooses to look, and is therefore absent exactly when you reach for it.
π¬ Deep Dive
Practitioner trap: the organisation that did federation best is the one that loses the Admin Console. Mature Okta tenants often federate administrative sign-in to an upstream IdP β Entra ID, a smartcard provider, a separate admin tenant β because standing local credentials are a liability. But Okta's DR documentation is explicit that external IdP authentication fails during failover, since those providers depend on redirect URIs pointing at the primary domain, while locally-sourced Okta credentials continue to authenticate admins. If 100% of your super admins sign in through an external IdP, your break-glass path and your failed path are the same path. The fix is unglamorous and needs doing before you need it: at least two break-glass super admins with Okta-sourced credentials and phishing-resistant local authenticators, excluded from the routing rule that sends admins to the external IdP, stored in the physical safe rather than the SSO-protected password manager, and exercised on a calendar. Note the second-order trap while you are there: users cannot reset passwords during failover either, so a helpdesk runbook whose first step is a password reset has no first step.
Staff-level framing: do not present this as "should we buy Enhanced DR". Present the 24-hour read-only window as a business continuity question and let the answer size the spend. The artefact is a one-page table: for each Tier-0 business process, what happens when authentication works but no identity can be created, changed or revoked for a day. Onboarding stops β acceptable. Contractor offboarding stops β now it is a legal and audit question. Incident response loses revocation β now it is a security question, and the honest mitigation is not a DR feature at all, it is short token lifetimes, because expiry is the only revocation primitive that still functions in a read-only plane. Then be precise about what the money buys: the delta between Standard and Enhanced is 55 minutes of degraded-but-authenticating versus degraded-but-authenticating, plus the right to pull the trigger yourself. That is worth real money to a trading platform and close to nothing to a company whose workforce signs in once a morning. Two more constraints belong in the paper: Enhanced DR is available on Production orgs only β not Preview, not Integrator Free β so you cannot rehearse the mechanism in a sandbox, only the runbook around it; and the audit trail lives in the System Log as system.dr.failover and system.dr.failback, with email to super admins, which is what you will hand a regulator asking who declared the disaster and when.
DR is a regional-infrastructure control, and your likeliest identity outage is not regional. Okta's documentation lists what disaster recovery does not cover, and the list is the uncomfortable one: DDoS attacks, third-party integration issues, data deletion or modification by bad actors, and configuration errors. Rank your actual incident history against that. A mis-scoped sign-on policy pushed at 17:00, a Universal Directory mapping that blanks an attribute every downstream app authorises on, an admin whose account is taken over and starts deleting groups β none of these are fixed by moving to Oregon, and failing over during one of them would make it worse by freezing the tenant read-only so you cannot roll the change back. That is the real argument for the identity-as-code posture: config errors need version control, review and a revert path, not a second region. Failover and rollback are different instruments answering different failures, and conflating them is how a team ends up with an expensive add-on and still no way to undo Tuesday's policy change.
# Enhanced DR is the only Okta control plane that deliberately lives OUTSIDE
# your org's normal hostname: it answers on drapp.<your-domain> so it stays
# reachable while the org itself is serving from the DR region.
#
# Scopes: okta.dr.read (status) | okta.dr.manage (failover/failback)
# okta.logs.read is a SEPARATE scope, for the System Log audit trail.
DR="https://drapp.YOUR_OKTA_DOMAIN/api/v1/dr"
# 1. Status. Poll this on a schedule from OUTSIDE the tenant, not during the
# incident -- the whole point is that it answers when the org is degraded.
curl -sS "$DR/status" \
-H "Authorization: Bearer <TOKEN>" -H "Accept: application/json"
# {"status":[{"domain":"YOUR_OKTA_DOMAIN","isFailedOver":false}]}
# 2. Declare the disaster. The body is optional; {} is accepted.
curl -sS -X POST "$DR/failover" \
-H "Authorization: Bearer <TOKEN>" -H "Content-Type: application/json" -d '{}'
# {"results":[{"domain":"YOUR_OKTA_DOMAIN","message":"Failover was successful"}]}
# 3. Go home again, once Okta reports the primary region healthy.
curl -sS -X POST "$DR/failbackStart" \
-H "Authorization: Bearer <TOKEN>" -H "Content-Type: application/json" -d '{}'
# The alert worth building is NOT "are we failed over" -- you will know that.
# It is: isFailedOver == true AND elapsed > 1h
# because since the transition every admin write your automation attempts --
# SCIM deprovisioning, Terraform applies, token revocation, JML workflows --
# has been failing, and CI that retries quietly will hide it from you.
Doc-verified against Okta's Enhanced Disaster Recovery developer guide (base URL drapp.{yourOktaDomain}, GET /api/v1/dr/status, POST /api/v1/dr/failover, POST /api/v1/dr/failbackStart, response shapes and OAuth scopes), Sep 2026. Placeholders only; not executed in this run.
π§ Recall
Eight days ago the IPSIE pill made the case for one interoperable enterprise identity profile, so that apps and identity providers finally speak the same session and signal language. Today's pill says external IdP authentication simply does not work during an Okta failover. Why does an interoperability profile not save you here β and what does that tell you about the layer interop actually operates on?
Show answer
Because the failure is not in the protocol, it is in the routing underneath it. Okta's DR documentation attributes the breakage to external providers relying on redirect URIs that point at the primary domain β the SAML or OIDC exchange is still perfectly well-formed, it is just addressed to a place that is no longer serving writes for you. An interop profile standardises the messages: what a session assertion contains, what a revocation signal means, which authenticator classes count. It does not standardise DNS, endpoint topology, or which region answers a redirect. The architectural lesson generalises well beyond Okta: profiles buy you semantic portability, never availability, and the two are commonly confused in vendor conversations. It is also why the break-glass path must be the least interoperable thing you own β a locally-sourced Okta credential that depends on nothing external is valuable here precisely because it is not federated.
πΌ Market Signal
Identity resilience stopped being an architecture preference in the EU and became a filing obligation. EIOPA's DORA overview, Reg. (EU) 2022/2554, in application since Jan 2025, puts roughly 20 categories of financial entity β plus their ICT third-party providers β under binding requirements for ICT risk management, digital operational resilience testing, third-party risk monitoring and major-incident reporting to competent authorities. An identity provider is the textbook critical ICT dependency: in scope for the continuity plan, for the testing programme, and for the supervisory conversation about concentration risk. Meanwhile the product side has been quietly maturing for two years β Okta's Enhanced DR GA announcement, Mar 2024, shipped the add-on first to North American commercial cells with expansion to global and compliance cells (FedRAMP, HIPAA) planned through 2025, which tells you the regulated segment was the target market from the start. Career read: this is a rare IAM topic that a CFO, a CISO and an auditor all have a stake in, and almost no Okta specialist can speak to it β most identity CVs stop at SSO, MFA and lifecycle. Being the person who can walk a regulator through a tested identity continuity plan, name the RTO, and say honestly what degrades is Staff-level positioning that no certification supplies, and it is directly sellable into EU financial services as fractional work.
β‘ Action This Week
Write the identity read-only runbook for a tenant you administer β one page, under two hours. Three sections. First, break-glass: list every super admin and mark which ones authenticate with Okta-sourced credentials versus an external IdP; if the answer is zero local, that is your finding. Second, writes you would lose: walk your incident playbooks and your automation (SCIM deprovisioning, Terraform applies, JML workflows, token revocation) and mark each step write or read β the write list is what stops for up to 24 hours. Third, detection: define the alert on isFailedOver staying true past an hour, and where system.dr.failover lands in your SIEM. Done = a one-page runbook whose first line states how many break-glass admins can sign in with no external dependency, and whose second states which incident-response step you would be unable to execute. That page is the portfolio artefact: "I audited my identity provider's degraded mode and found we had N admins who could still get in, and we could not revoke" is a far stronger LinkedIn post than another feature explainer, because it reports a result from your own estate rather than repeating a vendor's datasheet.
MCP Skills Went Final: Your Servers Now Ship Instructions, Not Just Tools
π‘ Key Concept
An MCP server has always been able to hand an agent tools. What it never had was a standard way to hand over the instructions for using them: the workflow, the order of operations, the things you must never do. That gap closed on 13 Sep 2026, when SEP-2640, the Skills Extension, was merged Final as the official extension io.modelcontextprotocol/skills. The design is deliberately unambitious at the transport layer: a skill is a directory whose files are served as ordinary MCP resources, conventionally under a skill:// URI, and the extension adds exactly three methods β skills/list and skills/get, both mandatory for any server declaring the extension, plus an optional resources/directory/read gated behind a directoryRead capability flag. The skill format is not redefined at all; it is delegated wholesale to the Agent Skills specification, so a host that already treats MCP resources as a virtual filesystem consumes a remote skill exactly as it consumes a local one.
The motivation is a context-budget problem before it is anything else. Server instructions ride in the instructions field of the discovery result and are practically bounded in size; the SEP cites an 875-line real-world skill that simply does not fit that model. Agent Skills solves size with progressive disclosure in three stages β roughly 100 tokens of name and description loaded at startup for every skill, the full SKILL.md body (recommended under 5,000 tokens) only once the skill is activated, and supporting files in references/, scripts/ or assets/ read strictly on demand. The extension preserves those stages over the wire: a host assembles its registry from listings alone and must not fetch any SKILL.md at that stage, and a virtual mount must resolve reads on access rather than pre-populating files. What arrives in context is what the task actually reached for.
Then comes the part worth your attention as an architect, and it is most of the document: the security model. The SEP is blunt that skill content is instructional text a model acts on, which makes it a prompt-injection surface, and it ranks an MCP-served skill as a higher-risk surface than a remote tool call β because unlike a tool invocation, a skill can put server-authored bytes on your filesystem and then tell the model to execute them with host-side tools. So hosts MUST treat the content as untrusted input, MUST tag it with its originating server at the point it enters context, MUST NOT allow host-side code execution from it without explicit per-skill user approval, and MUST ignore the Agent Skills allowed-tools field for MCP-origin skills unless the user approved that grant specifically. The spec's framing of that last one is the sentence to remember in design review: a remote server filling in allowed-tools is requesting elevated access on your machine, not describing its own.
Cross-pollination: today's IAM pill covers Okta's disaster-recovery mode, where authentication keeps issuing sessions while every administrative write β revocation included β is frozen for up to 24 hours. Same shape here: approval is bound to the digest set observed when the user approved, but the host need not poll, so a revocation exists in the spec and fires only when something re-fetches.
π¬ Deep Dive
Practitioner trap: assuming that fetching a skill is the same as loading it. The SEP separates the two on purpose. resources/read is transport β it returns bytes to whoever asked, a resource browser, a generic read tool, a curious user β and the spec states that hosts MUST NOT treat a SKILL.md arriving by that route as a load: it grants no approval, opens no window, and confers no standing on the skill's supporting files. Activation happens only through the host's own skill-loading path, the one that verifies content against the entry's digests and applies user approval. Get this wrong in a client you are building and you have written the vulnerability the extension was designed to prevent: a model that reads a resource and is thereby steered by it, with no approval gate and no origin tag. The same trap has a nested variant β a SKILL.md sitting inside an approved skill's directory is ordinary markdown, and hosts MUST NOT act on its frontmatter; activating it needs fresh consent, or a server can ride new instructions and new permission requests in on an approval the user already gave.
Staff-level framing: treat connecting an MCP server as a software supply-chain decision, and write the policy before the first server ships a skill. Three properties from the spec should become three lines in your standard. Identity: resolve skill names per-origin using a host-assigned label, never the server's self-reported name, because names collide and a malicious server can publish under a popular skill's name and count on your host resolving its way. Integrity: digests confirm the listing matches what you fetched, and nothing more β they are unsigned and come from the same server as the content, so they defend an approval after it is granted but cannot tell you the content deserved approval. Containment: a resource-read surface driven by skill content is a cross-server confused-deputy vector, so reads must be bound to the skill's originating server and any cross-origin read gated on per-call approval naming both servers. The governance cost is real and worth stating in the paper: per-skill, content-bound approval means a server that rotates one supporting file revokes the approval and re-prompts your users. Teams will feel that as friction and ask you to disable it. The answer is a curated allowlist of servers, not a weaker binding.
The extension is small because the July core did the heavy lifting. Skills rides entirely on Resources and adds almost no new surface β which is only possible because the 2026-07-28 revision had already made MCP's core stateless: no initialize handshake, no Mcp-Session-Id, protocol version and client identity travelling in request _meta, so a server runs on any instance behind an ordinary load balancer with no shared storage. That revision also gave list results cache attributes (ttlMs, cacheScope), which skills/list inherits, and moved header-based routing into Mcp-Method and Mcp-Name so a gateway can route and meter without parsing JSON bodies. Read together, the direction is unmistakable and it should shape your build: MCP is becoming infrastructure you put a proxy in front of and govern centrally, not a stateful session you hold open per client. If your internal MCP estate still assumes sticky sessions, the Skills rollout is the forcing function to fix that first.
// 1. The server declares the extension (SEP-2133 negotiation). Declaring it
// COMMITS the server to BOTH skills/list and skills/get, and it must also
// declare `resources`. directoryRead is the one optional setting.
{
"capabilities": {
"resources": {},
"extensions": {
"io.modelcontextprotocol/skills": { "directoryRead": true }
}
}
}
// 2. skills/list -- the authoritative record of what this server publishes.
{ "jsonrpc": "2.0", "id": 4, "method": "skills/list", "params": {} }
// 3. The result. `resources` is the governance payload, not the content:
// every file, its sha256 digest and its size -- and NOT the bytes.
{
"jsonrpc": "2.0",
"id": 4,
"result": {
"resultType": "complete",
"skills": [
{
"uri": "skill://acme/billing/refunds/SKILL.md",
"frontmatter": {
"name": "refunds",
"description": "Process customer refund requests per company policy"
},
"resources": [
{ "uri": "skill://acme/billing/refunds/SKILL.md",
"digest": "sha256:b2c3d4e5", "size": 3871 },
{ "uri": "skill://acme/billing/refunds/examples/email.md",
"digest": "sha256:c3d4e5f6", "size": 962 }
]
}
]
}
}
// Read the URI, not the frontmatter: the LAST path segment is always the
// skill name ("refunds"); "acme/billing" is a server-chosen prefix. So the
// name is recoverable without fetching a byte -- which is exactly why the
// registry is built from listings alone.
//
// Three review questions for any server you are about to allow:
// resultType != "complete" -> catalog is partial; skills/get by URI still
// works, so "not listed" is not "not served".
// "resources": "dynamic" -> generated content; it CANNOT be content-bound.
// A persisted approval covers nothing. Decline.
// digest set changed -> prior approval is revoked; you MUST re-prompt.
Syntax-checked β the three JSON objects above were parsed with json.load in this session; shapes and field names follow the SEP-2640 examples (capabilities.extensions, skills/list request and result), Sep 2026. Digests truncated for legibility.
π§ Recall
Eight days ago the AG-UI pill covered a protocol carrying typed events and state deltas from an agent out to the user interface. Today's extension carries instructions the other way, from a server into the agent. Only one of the two makes origin tagging a hard requirement of the specification. Which, and what is the principle that decides it?
Show answer
Skills. The deciding principle is who or what consumes the payload and whether it can be steered by it. AG-UI's payload terminates at a renderer: its events and state deltas are turned into pixels for a human, so its risk surface is rendering and state correctness β real problems, but a human reader is not reprogrammed by a malformed event. A skill's payload terminates in the model's context as instructions the model will act on, which makes it a prompt-injection surface by construction. That is why SEP-2640 requires hosts to tag MCP-served skill content with its originating server at the point it enters context and forbids presenting it as indistinguishable from a local filesystem skill: the model, not the host, is the layer that ultimately decides whether to follow the instructions, so withholding provenance from it makes the untrusted-input rule unenforceable where it matters. Generalise it into a design rule: any channel whose content reaches the model as instructions needs provenance carried all the way to the point of use, not just checked at the edge.
πΌ Market Signal
The context economics behind this extension are measurable, not theoretical. Datadog's State of AI Engineering, Jul 2026, drawn from over 1,000 customers with data through March 2026, found that system prompts account for 69% of all input tokens while only 28% of LLM calls show cached-read tokens β most teams are paying full price, on every call, for instructions the request never needed. It also found agent-framework adoption nearly doubling from 9% to 18% year over year, over 70% of organisations running three or more models, and rate-limit errors making up roughly a third of all LLM failures (8.4 million in March 2026 alone). Progressive disclosure is the direct answer to the first pair of numbers and the governance model is the tax on the answer. The governance itself is telling: the Skills Over MCP Working Group charter, Sep 2026 lists leads from Nordstrom, Anthropic and Bloomberg with participants from Google, Databricks, GitHub, AWS, Saxo Bank, Stacklok, Astronomer and Arcade.dev β an enterprise-and-vendor mix, with conformance scenarios merged 11 Sep 2026 and SDK implementations in Python, C# and Go. Career read: the scarce skill is no longer wiring an MCP server up, it is writing the policy that decides which servers may put instructions in your agents' context and on what terms. That is a security-architecture conversation an AI engineer is usually not in the room for, and identity people usually lack the protocol depth to hold β the overlap is a narrow, well-paid seam.
β‘ Action This Week
Two hours, two halves. First, author one real skill for work you actually repeat β a deploy checklist, an incident triage sequence, a house style for reviews. Obey the format: name lowercase-and-hyphens matching the directory and under 64 characters, a description that says what it does and when to use it, body under 500 lines, detail pushed into references/. Validate it with skills-ref validate ./my-skill. Second, do the governance half, which is the part almost nobody will publish: for every MCP server your team currently connects, answer three questions β who operates it, would you let it place instructions into an agent's context, and does your host show you the origin of skill content when it does. Done = one validated skill directory, plus a short table of your connected servers with an allow-or-deny call and one sentence of reasoning each. The table is the artefact worth writing up: "we audited the MCP servers we connect against SEP-2640's trust rules and denied N of M" is a concrete governance result, and it will read as considerably more senior than another explainer of what MCP is.
Okta ISPM: Discovering the Non-Human Identities β and Shadow AI Agents β Nobody Provisioned
π‘ Key Concept
Every previous agentic-identity pill assumed you knew the agent existed: you minted it a scoped token, you exchanged on its behalf, you put it in a registry. Discovery is the step before all of that, and it is the one nobody owns. Okta's Identity Security Posture Management attacks it from the other end β instead of governing what was provisioned, it continuously pulls from the apps and identity providers integrated with it and reconstructs the identities that are actually there. Its non-human identity model splits that population into four labels: service accounts (human- or system-created identities behind automations, integrations and shared access), keys and tokens (AWS access keys, Okta API tokens, Snowflake key pairs, GitHub PATs, SSH keys), workload identities (apps and services that authenticate programmatically with no human sign-in), and AI agents. A fifth flag β admin or super admin β cuts across all four and is the one that turns an inventory row into an incident.
On top of that inventory ISPM runs 19 issue detections carrying an NHI label, mapped to the OWASP Non-Human Identity Top 10: improper offboarding, overprivilege, insecure authentication, long-lived secrets. The interesting movement is recent and specifically about agents. Per the ISPM release notes, on 1 Sep 2026 AI agent discovery stopped being exclusive to Okta for AI Agents and became available to all ISPM customers, Microsoft Copilot Studio agent discovery went GA, and endpoint shadow-AI-agent discovery through CrowdStrike entered Early Access. Read the three together and the product thesis is explicit: sanctioned agent platforms get inventoried by API, and the unsanctioned ones get caught on the endpoint and in the browser β the same posture split Okta describes in its four-stage agent plan, Feb 2026 (know the crown-jewel platforms, discover the unknown ones, harden the NHI layer, consolidate into one policy plane).
The uncomfortable part for an architect is that discovery is a posture product, not a control plane: ISPM tells you the identity exists and why it is risky, then hands remediation to event hooks plus Okta Workflows for automation, or to the Admin Console under an Okta Privileged Access subscription for manual action. Visibility and authority are sold, and staffed, separately β which is exactly where most NHI programmes stall.
Cross-pollination: today's AI pill covers the OpenAI Agents API, where every durable session is a server-side agent holding an API key and MCP credentials. That agent is not a Copilot Studio agent and will not be discovered as one β it lands in your estate as a key, which is precisely why OpenAI shipped org- and project-level API key governance controls five days after the Agents API beta.
π¬ Deep Dive
Practitioner trap: the Okta API token is the NHI whose lifecycle rules people get wrong most often, and it is sitting in your own tenant. Per Okta's token documentation, an SSWS token is valid for 30 days and renews on every API request β so a token called daily by a CI job never expires, and "30-day token" is not a rotation story. Worse for governance: the token carries the permissions of the admin who created it and those permissions track that user, so promoting the creator silently promotes every automation they ever wired up; deactivating them silently breaks it, because tokens from a deactivated user are rejected. That is improper offboarding and overprivilege in a single object, and it is why ISPM reports the human owner next to the machine credential rather than on a separate screen.
Staff-level framing: do not sell NHI discovery as a security project β sell it as the prerequisite for the agent rollout leadership already approved. The credible artefact is a two-column table: discovered NHIs per category versus NHIs with a named accountable owner, an expiry and a rotation path. The delta is the programme. Blast radius is the second column: every admin-labelled service account or admin-created API token is an identity whose compromise is indistinguishable from a compromised super admin, with no MFA in the path. Then be honest about the operating model β ISPM detects, and then you need Workflows to act, or a Privileged Access subscription to act by hand. Budget the remediation arm before you buy the telescope, or you will own a dashboard of findings nobody is authorised to close.
Coverage is plane-bound, so name the plane in every claim. ISPM discovers AI agents where it has a connector or a sensor: Salesforce Agentforce and Copilot Studio by API, unsanctioned platforms through the browser plugin, endpoint agents through CrowdStrike in Early Access. An agent your own team ships β a container with an OpenAI or Anthropic key and three MCP servers β matches none of those detectors as an agent. It will surface, if at all, as a key or a workload identity, with no link to the prompt, the tools or the blast radius. "We have agent discovery" is therefore always an incomplete sentence; the useful version is "we discover agents on these four planes, and here is the fifth we do not cover yet."
# The cheapest NHI inventory you can run today needs no ISPM licence: every
# SSWS token in your own tenant is a non-human identity with a human owner.
# Scope: okta.apiTokens.read (okta.apiTokens.manage to revoke)
curl -sS -X GET "https://YOUR_OKTA_DOMAIN/api/v1/api-tokens" \
-H "Accept: application/json" \
-H "Authorization: SSWS ${OKTA_API_TOKEN}" \
| jq -r '.[] | [ .id, .name, .clientName, .userId, .created, .expiresAt,
.tokenWindow, (.network.connection // "ANYWHERE") ] | @tsv'
# 00Tabcdefg1234567890 ci-provisioner Okta API TEST_USER_ID
# 2024-03-02T09:11:00.000Z 2026-10-22T09:11:00.000Z P30D ANYWHERE
# Three questions per row β all of them identity questions, not token questions:
# .userId -> whose permissions is this really running with today?
# .tokenWindow -> ISO-8601 window that RENEWS on use; daily use = immortal
# .network.connection -> ANYWHERE means a machine credential with no IP zone
# Then revoke what no one will claim (scope: okta.apiTokens.manage):
curl -sS -X DELETE \
"https://YOUR_OKTA_DOMAIN/api/v1/api-tokens/00Tabcdefg1234567890" \
-H "Authorization: SSWS ${OKTA_API_TOKEN}" # 204 No Content
Doc-verified against the Okta Admin Management API β API Tokens reference (GET /api/v1/api-tokens, DELETE /api/v1/api-tokens/{apiTokenId}, response fields and scopes), Sep 2026. Placeholders only; not executed in this run.
π§ Recall
Eight days ago the Rich Authorization Requests pill bound an agent's token to one specific transaction through authorization_details instead of a coarse scope. Today's pill inventories long-lived SSWS tokens that grant whatever their creating admin can do. Why does RAR not fix the tokens you just listed β and what does that tell you about the order of the work?
Show answer
RAR is a property of an OAuth authorization flow: an authorization server mints a token carrying the granted authorization_details, and a resource server enforces it. An SSWS API token is not issued by that flow at all β it is a static credential minted in the Admin Console that inherits the creator's admin permissions wholesale, with no consent step, no audience and nothing to constrain per call. You cannot narrow it; you can only replace it with an OAuth service app and scoped grants, or revoke it. Hence the order: discovery first, because fine-grained authorization only governs the identities that came through the front door, and the ungoverned population is exactly the one that did not.
πΌ Market Signal
Okta is pricing this thesis into its own results. In Q2 FY2027, Okta Newsroom, Aug 2026 the company reported revenue of $805M (up 11% year over year), subscription revenue of $793M (up 12%), RPO of $4.858B (up 17%) and cRPO of $2.585B (up 14%), with free cash flow of $227M at a 28% margin β growth in the low teens, but a backlog growing faster than revenue, which is what a platform sells when customers are signing multi-year for capabilities not yet deployed. CEO Todd McKinnon framed the driver as agents needing "a trusted identity and clear controls over what it can access". The shipping cadence backs it: the ISPM release notes, Sep 2026 show agent and workload-identity features landing roughly monthly through 2026. Career read: NHI discovery is where identity work stops being SSO configuration and starts being estate-wide governance β the one place an Okta specialist can speak credibly to a CISO about AI risk without leaving their own domain, and the part of the agentic story a generalist security engineer cannot fake.
β‘ Action This Week
Run the token inventory above against an Okta developer org (or any tenant you administer), then widen it by hand to one other system you control β a GitHub org's PATs, or your cloud account's access keys. Classify every row into the four ISPM labels (service account, key/token, workload identity, AI agent), and add the two columns that matter: named human owner and can it expire. Count how many rows have neither. Done = a single table of your real non-human identities with an owner-and-expiry gap count at the bottom. That number is the whole post: "I inventoried the non-human identities in two systems I own; N of M have no owner and cannot expire" is a portfolio artefact that reads as governance practice rather than vendor commentary β and it is the same table you would put in front of a client in week one.
The OpenAI Agents API: What You Actually Rent When You Stop Running Your Own Harness
π‘ Key Concept
A production agent is mostly not the model. It is the harness around it: the loop that decides when to call a tool, the compaction that keeps a long task inside the context window, the checkpoint that survives a crash, the sandbox the agent writes files in, and the bookkeeping that lets a task resume tomorrow. Every serious team has built that layer β and every one of them has rebuilt it. On 10 Sep 2026 OpenAI put its own version behind an API: the changelog records the Agents API entering public beta with the managed Codex harness, OpenAI handling session orchestration, context compaction and recovery, while you supply tools and pick where code runs.
The object model is small enough to hold in your head, which is the point. Per the Agents API overview there are four concepts: an Agent (model, instructions, tools, MCP servers), an Environment (an optional sandbox or computer for files and commands), a Session β a durable instance of that agent working on a task across turns β and events and items, the inputs sent in and outputs produced. You create one with POST /v1/agents/sessions behind the header OpenAI-Beta: agents=v1; the body nests agent, environment and input. Tools are declared by type β programmatic_tool_calling, mcp (with a server_label and a transport), web_search β and fan-out is a two-field switch: multi_agent: { enabled: true, max_concurrent_subagents: 4 }. Execution is none, openai_hosted, or self_hosted, the last pointing at your own compute through a workspace_directory and capability_directories.
Read it as a pricing of your own engineering rather than as a framework. Nothing here is capability you could not build; all of it is capability you would otherwise maintain. The architecture question is therefore not "is the harness good" but "which half of my agent stack am I willing to stop being able to inspect" β and the answer changes the moment a regulator, an incident review or a data-residency clause asks what the agent saw at step 40.
Cross-pollination: today's IAM pill is the other side of this bill β a durable server-side session holding your API key and MCP credentials is a non-human identity that no agent-discovery connector will recognise as an agent, which is why OpenAI shipped org- and project-level API key governance controls on 15 Sep, five days after this beta.
π¬ Deep Dive
Practitioner trap: the blocker is in the fine print, not the API. The overview states plainly that the Agents API is US-only for data residency and does not support Zero Data Retention. For an EU-regulated workload that is not a configuration detail you negotiate later β it disqualifies the managed harness for the exact enterprise use cases whose long-running, resumable tasks make the harness attractive in the first place. Two smaller traps ride along: OpenAI-Beta: agents=v1 is a beta contract, and the durability you are buying means your session state now lives on someone else's side of that contract. Choosing environment.type: "self_hosted" relocates the sandbox, not the session.
Staff-level framing: write the boundary down before the spike, because "we saved six weeks of harness work" and "we cannot explain what the agent did" are the same decision seen from two quarters apart. Three columns decide it. Audit trail: compaction is a lossy rewrite of the agent's own history, and a managed compaction policy is one you cannot version, diff or replay β if your incident review needs step 40 verbatim, you need the transcript, not the summary. Cost: there is no separate Agents API fee, so spend appears as model tokens, tool calls and hosted container rates β an agent that silently retries for an hour bills like one, with no line item named "harness". Exit: tools and MCP servers are portable because MCP is a standard; sessions, items and the compaction behaviour your prompts were tuned against are not. Rent the harness for tasks that are long and boring; keep your own for the ones a regulator will read.
The features that actually change your design are the two you would build last.programmatic_tool_calling lets the agent compose calls in code rather than one round trip per tool, which collapses the fan-out latency that makes hand-rolled loops feel slow; multi_agent.max_concurrent_subagents turns delegation into a capacity dial instead of a topology you hand-wire β last month's supervisor-orchestration pill was largely about doing that by hand. Note what remains yours either way: the instructions, the tool contracts, the credentials those tools carry, and every eval. The harness makes a bad agent fail faster and more durably; nothing in this API tells you whether the output was right.
curl -sS -X POST "https://api.openai.com/v1/agents/sessions" \
-H "OpenAI-Beta: agents=v1" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"agent": {
"model": "gpt-6-astra",
"instructions": "Use the OpenAI documentation MCP and web search to answer technical questions accurately. Delegate independent research tasks to subagents when useful.",
"tools": [
{ "type": "programmatic_tool_calling" },
{ "type": "mcp",
"server_label": "openai_docs",
"transport": { "type": "http",
"server_url": "https://developers.openai.com/mcp" } },
{ "type": "web_search" }
],
"multi_agent": { "enabled": true, "max_concurrent_subagents": 4 }
},
"environment": {
"type": "self_hosted",
"workspace_directory": "/workspace",
"capability_directories": ["/workspace/capabilities/skills"]
},
"input": [
{ "role": "user",
"content": [ { "type": "input_text",
"text": "Research how to connect an MCP server to an OpenAI agent, check for recent updates, and summarize the recommended setup." } ] }
]
}'
# Read it as a boundary, not a payload. Everything under "agent" is yours to
# version and eval. Everything the session does between turns - the loop, the
# summarisation, the retry after a crash - is the part you stopped owning.
Doc-verified against the OpenAI Agents API overview (create-session example, environment types, tool types and multi_agent fields), Sep 2026. Reformatted for width; keys and nesting unchanged. Not executed in this run.
π§ Recall
Six days ago the AG-UI pill argued that MCP standardises agentβtool and A2A agentβagent, leaving the agentβhuman channel to be hand-rolled per product. The Agents API streams session progress back to you. Does that close the AG-UI gap?
Show answer
No β it moves it. A managed session gives you a vendor-specific stream of events and items describing what the harness is doing server-side; AG-UI is an open, transport-agnostic contract for what the front-end renders, with typed events and JSON-Patch state deltas that any compliant agent backend and any compliant UI can speak. Consuming OpenAI's stream still means writing a translation layer into your own components, and that layer is tied to this vendor's event shape. The two live at different seams: the Agents API standardises the agent's execution, AG-UI standardises its conversation with the human. Renting one does not relieve you of the other.
πΌ Market Signal
The managed harness is arriving into a market whose measured weakness it does not address. LangChain's State of Agent Engineering, Jun 2026 β 1,340 respondents β found 57% with agents in production and 89% running observability, but only 52.4% running offline evals and 37.3% running online evals, with quality named the top blocker to production by 32% (latency second at 20%) and more than 75% running multiple models in production. That is the shape of the opportunity: orchestration is being commoditised by the platforms while evaluation, the thing that decides whether an agent is fit to ship, stays a team's own problem β and multi-model estates mean a vendor harness is a partial answer by construction. Pricing reinforces it; per the OpenAI docs, Sep 2026, the Agents API carries no separate fee, only model, tool and container rates β platforms give the harness away because the differentiated work was never there. Career read: as harnesses commoditise, "I built an agent loop" stops being a credential. Eval design, cost and latency budgets, and the residency and audit judgement calls in this pill are what survive the commoditisation.
β‘ Action This Week
Take one agent you already run and fill a six-row boundary table β session durability, context compaction, crash recovery, tool calling, sandboxed execution, subagent fan-out β with three columns: who owns it today, what renting it would cost you in inspectability, and whether US-only residency and the absence of Zero Data Retention rule it out for that workload. If you have an API key, spend fifteen minutes creating one session with "environment": {"type": "none"} and a single MCP server so you are judging the real event stream rather than the marketing. Done = a one-page table whose last row is an explicit rent-or-build decision with the reason written next to it. The honest version of that page β "we rent the harness for batch research, we keep ours for anything an auditor reads" β is a stronger LinkedIn post than another launch summary, because almost nobody publishes the boundary they chose.
IPSIE: Turning Six Enterprise Identity Standards Into One Interoperable Profile
π‘ Key Concept
Enterprise identity is never one protocol. A single SaaS app is federated in over SAML or OIDC, provisioned over SCIM, told about risk and session changes over the Shared Signals Framework, and de-provisioned by yet another path. Each of those specs was written to cover many contexts, so each is full of optionality β and optionality is the enemy of interoperability: two vendors can both be "SCIM-compliant" and still not talk without a bespoke connector. At Oktane 2024 Okta proposed, and with the OpenID Foundation founded, the Interoperability Profile for Secure Identity in the Enterprise (IPSIE) working group to fix exactly this β not by inventing a new protocol, but by profiling the ones we already have down to a secure-by-default, interoperable subset.
IPSIE is organized by the dimensions of the enterprise identity lifecycle rather than by protocol: Single Sign-On (an OpenID Connect profile), User Lifecycle Management and Entitlements (a SCIM profile), Risk Signal Sharing (a Shared Signals Framework profile β CAEP for session change, RISC for account change), and Logout / Token Revocation. The founding table is deliberately cross-vendor β Okta and the OpenID Foundation with Ping Identity, Microsoft, Capital One, SGNL and Beyond Identity β because a profile only matters if the other side implements it too. Status matters for how you talk about it: the group meets weekly and has a v1 requirements draft, but no final specification or certification program yet.
For an architect the value is that IPSIE turns "we integrated app X" from artisanal connector work into a conformance target, and it makes the hardest, least-built dimension β killing a session across every downstream app in seconds when someone is terminated or a device falls out of compliance β a named, testable requirement instead of a hope. Okta, unsurprisingly, frames its own stack as the on-ramp: Auth0 for the SSO and lifecycle profiles, Okta FGA for the entitlements dimension, native Shared Signals for risk.
Cross-pollination: the agent-stack protocols in today's AI pill (AG-UI, A2A, MCP) standardize how agents talk to users, each other and tools β but those agent sessions still ride on the enterprise SSO, provisioning and session-revocation plumbing that IPSIE is trying to make interoperable.
π¬ Deep Dive
Practitioner trap: IPSIE is a v1 requirements draft, not a ratified spec β there is no conformance suite or certification to be "IPSIE-compliant" against yet. Do not put "IPSIE-compliant" on a slide; put "IPSIE-aligned across N of 6 dimensions." And beware the easy-20% / hard-80% split: SSO is the dimension every vendor already ships, so a readiness claim that means "we do SSO" is theatre β the session-revocation and risk-signal dimensions (a working Shared Signals receiver) are the unbuilt, valuable work. The entitlements dimension is the least mature of all: SCIM never standardized fine-grained entitlements, which is exactly why Okta routes that dimension through FGA / ReBAC rather than pure SCIM.
Staff-level framing: use IPSIE as a gap-analysis scaffold, not a product to buy. Score each crown-jewel SaaS app against the six dimensions; the red cells are your integration debt and your blast-radius narrative in one table. The governance question "when we terminate an employee, how many seconds until every downstream session is dead?" stops being a hope and becomes a measured property of your Shared Signals receiver coverage. Rollout without boiling the ocean: adopt CAEP session-revocation receipt on the two or three apps that hold real data first β highest security payoff, lowest current vendor coverage, best story to a CISO.
Compose, don't confuse: IPSIE is not SSF and not SCIM β it profiles them. It pins the optional knobs so an independent transmitter and receiver actually line up: SSO is an OIDC profile; lifecycle and entitlements a SCIM profile; risk plus session termination a Shared Signals profile (CAEP session events, RISC account events). The thing IPSIE adds over "we support SCIM" is the removal of optionality β the reason two conformant implementations interoperate without a hand-written connector.
# IPSIE's hardest dimension is real-time session revocation. It rides on the
# Shared Signals Framework: the IdP (transmitter) PUSHES a CAEP Security Event
# Token to each app (receiver) the instant a session must die.
# 1) Receiver registers a push stream (RFC 8935) with the transmitter:
POST https://YOUR_OKTA_DOMAIN/ssf/streams
Authorization: Bearer <TOKEN>
{
"delivery": { "method": "urn:ietf:rfc:8935",
"endpoint_url": "https://app.example.com/ssf/events" },
"events_requested": [
"https://schemas.openid.net/secevent/caep/event-type/session-revoked"
]
}
# 2) On termination the transmitter POSTs a signed SET (a JWT) to the receiver:
{
"iss": "https://YOUR_OKTA_DOMAIN", "aud": "https://app.example.com",
"iat": 1789900000, "jti": "a1b2c3",
"events": {
"https://schemas.openid.net/secevent/caep/event-type/session-revoked": {
"subject": { "format": "email", "email": "TEST_USER_ID@example.com" },
"event_timestamp": 1789900000,
"reason_admin": { "en": "User terminated in HR system" }
}
}
}
# The receiver validates the JWT and kills the local session now β no polling,
# no waiting for the access token to expire. That is the dimension IPSIE names.
Illustrative, modeled on the Shared Signals Framework push delivery (RFC 8935) and the OpenID CAEP session-revoked event type, Sep 2026. Placeholders only; not executed in this run.
πΌ Market Signal
IPSIE is a vendor-backed bet, and the backers are the market. Okta launched it at Oktane 2024 with the OpenID Foundation as co-founders and Ping Identity, Microsoft, Capital One, SGNL and Beyond Identity, Oct 2024 at the table; the OpenID Foundation working group, 2026 now meets weekly to profile OIDC, SCIM and the Shared Signals Framework. Okta is already selling the on-ramp: its own guidance, Okta 2026 maps IPSIE's dimensions onto Auth0 and Okta FGA and tells developers that building on that stack now puts them "well on the way to achieving the IPSIE standard once it is finalized." For a career: the profile turns a fuzzy "identity integration" role into a concrete, checkable competency β the architect who can run a six-dimension IPSIE gap analysis across a SaaS estate, and knows the session-revocation dimension is the hard one, is selling exactly the identity-fabric skill the standard is manufacturing demand for.
β‘ Action This Week
Build an IPSIE dimension scorecard for eight SaaS apps you know well: columns = SSO, SCIM lifecycle, entitlements, Shared Signals risk-in, session termination, token revocation; rate each red / amber / green. Then take your single most sensitive app and check one concrete thing β can it receive a CAEP session-revoked signal today, or does a terminated user's session only die when the token expires? Done = a one-page scorecard plus a two-line note naming the highest-value gap you found. It becomes a LinkedIn post that reads as architecture, not vendor news β "I scored eight SaaS apps against the six IPSIE dimensions; seven do SSO, one can kill a session in real time" β the gap is the insight.
AG-UI: The Missing Protocol for the AgentβUser Channel MCP and A2A Leave Out
π‘ Key Concept
The agent ecosystem converged on two protocols fast: MCP standardized how an agent reaches tools, A2A how an agent talks to other agents. The third edge β how an agent talks to the human inside a real application β was left to every team to hand-roll: a bespoke SSE or WebSocket stream shuttling token deltas, tool-call status and UI state between a Python agent and a React front-end, re-invented and re-debugged per product. AG-UI (Agent-User Interaction Protocol) is an open, event-based protocol, MIT-licensed on GitHub, that standardizes that channel as a stream of typed events.
The backend answers a RunAgentInput by emitting a stream of typed events over a transport-agnostic layer β Server-Sent Events, WebSockets, webhooks or plain HTTP. The reference defines 34 standard event types in clean groups: lifecycle (RUN_STARTED, RUN_FINISHED, RUN_ERROR, STEP_STARTED/FINISHED), text (TEXT_MESSAGE_START/CONTENT/END), tool calls (TOOL_CALL_START/ARGS/END/RESULT), reasoning, and β the part that makes it more than a chat stream β shared state: STATE_SNAPSHOT ships the whole agent state once, then STATE_DELTA ships incremental changes as JSON Patch (RFC 6902) operations, so a collaborative UI and the agent stay in lockstep without resending the world every turn.
The payoff is the decoupling MCP brought to tools: the front-end binds to the event contract, not to your agent framework. Swap LangGraph for CrewAI, Pydantic AI or Google's ADK behind the same stream and the React app does not change β which is why those frameworks and the major clouds moved to emit AG-UI natively.
Cross-pollination: AG-UI carries no identity β it is the presentation contract, not the trust boundary. The human's session still needs enterprise SSO and per-action token scoping (today's Okta / IPSIE IAM pill); AG-UI decides what the user sees, not what the agent is allowed to do.
π¬ Deep Dive
Practitioner trap: AG-UI is a wire contract, not a UI kit β it defines the events crossing the boundary and renders nothing. CopilotKit is the reference renderer, but the protocol is not CopilotKit; keep them separate or you couple your product to one vendor's React components. Two: STATE_DELTA is JSON Patch (RFC 6902) applied sequentially β drop or reorder one delta on a flaky connection and client state silently diverges from the agent, so you need periodic STATE_SNAPSHOT / MESSAGES_SNAPSHOT resyncs and idempotent application, not just optimistic patching. Three: it is a de-facto standard by adoption, stewarded by a single company (CopilotKit) rather than a neutral body, with deliberately loose event-format matching β pin your SDK version and never assume every producer emits every event.
Staff-level framing: the real win is a stable seam between the agent runtime and the product surface. Adopt the event schema as your internal contract even if you never ship the library β it forces you to name every state the agent exposes to a user (tool calls in flight, reasoning, partial output, editable shared state), which is exactly the list your design and security reviews need. Observability: that typed stream is already an audit and telemetry tap β every tool call and state change is an event, so fan it into your OpenTelemetry GenAI spans (a prior pill) instead of inventing a second event log. Cost and latency at scale: streaming STATE_DELTAs instead of re-sending snapshots is the difference between a toy chat box and a collaborative surface that stays cheap under load.
MCP + A2A + AG-UI are one agent's three edges β compose them. A user edits a document in the browser (AG-UI shared state) → the agent retrieves context through a tool (MCP) → hands a subtask to a specialist agent (A2A) → streams the answer back as TEXT_MESSAGE_CONTENT plus a STATE_DELTA the UI applies. Each protocol owns exactly one boundary; adopting the three retires the three bespoke integrations most teams still maintain by hand.
# AG-UI is a stream of typed events over SSE. One agent run, wire `type` values:
data: {"type":"RUN_STARTED","threadId":"th_1","runId":"run_1"}
data: {"type":"TEXT_MESSAGE_START","messageId":"m1","role":"assistant"}
data: {"type":"TEXT_MESSAGE_CONTENT","messageId":"m1","delta":"Booking your "}
data: {"type":"TEXT_MESSAGE_CONTENT","messageId":"m1","delta":"flight..."}
data: {"type":"TEXT_MESSAGE_END","messageId":"m1"}
data: {"type":"TOOL_CALL_START","toolCallId":"tc1","toolCallName":"search_flights"}
data: {"type":"TOOL_CALL_ARGS","toolCallId":"tc1","delta":"{\"to\":\"LIS\"}"}
data: {"type":"TOOL_CALL_END","toolCallId":"tc1"}
# Shared state: incremental JSON Patch (RFC 6902), applied in order on the client
data: {"type":"STATE_DELTA","delta":[
{"op":"add","path":"/itinerary/0","value":{"flight":"TP123"}},
{"op":"replace","path":"/status","value":"awaiting_confirmation"}]}
data: {"type":"RUN_FINISHED","threadId":"th_1","runId":"run_1"}
Doc-verified against the AG-UI event reference (34 event types; STATE_DELTA = JSON Patch, RFC 6902), Sep 2026. Wire type values shown in SCREAMING_SNAKE_CASE (the SDK enum members are PascalCase). Placeholders only; not executed in this run.
πΌ Market Signal
The adoption curve is steep and cross-vendor. In May 2026 CopilotKit, AG-UI's steward, raised a $27M Series A, May 2026 (Glilot Capital, NfX, SignalFire), citing over 40,000 GitHub stars and 4M+ weekly downloads across CopilotKit and AG-UI and enterprise users including DocuSign, S&P Global, Cisco and Deutsche Telekom. The protocol is already emitted natively by Google, Amazon, Microsoft, Oracle and LangChain, Sep 2026, plus Mastra, Pydantic AI, Agno, AG2 and LlamaIndex β the same "the big clouds all shipped it" pattern that made MCP a default. For a career: "full-stack AI engineer" is quietly bifurcating into people who can wire an agent runtime to a real, stateful UI and people who can only return a chat string. AG-UI is the concrete artifact that proves you are in the first group.
β‘ Action This Week
Put an AG-UI event stream in front of an agent you already have. A ~50-line SSE endpoint that emits RUN_STARTED, streams two TEXT_MESSAGE_CONTENT deltas, fires one TOOL_CALL_START/ARGS/END trio, and pushes one STATE_DELTA (a JSON Patch the browser applies), then RUN_FINISHED β with a tiny HTML page that renders the text and applies the patch to a state object β proves the whole contract. Done = a curl -N capture of the event stream plus a screenshot of the browser updating from a STATE_DELTA. Ship it as a LinkedIn post: "MCP connects agents to tools, A2A to other agents β AG-UI connects them to your UI. Here is the third protocol wired end-to-end in 90 minutes," with the event log as the evidence.
Rich Authorization Requests: Binding an Agent's Token to One Action, Not a Coarse Scope
π‘ Key Concept
OAuth scope is a flat list of strings β payments:write, mail.modify. It answers "what class of thing may this token touch," never "which resource, which action, up to what limit." For a human clicking a button once, that coarseness is tolerable. For an autonomous agent that will call the API thousands of times on its own initiative, a coarse scope is a standing grant of everything in the class: the token that may "transfer money" may transfer any amount, to any account, any number of times. Rich Authorization Requests (RFC 9396, May 2023) replaces that flat string with a structured authorization_details JSON array, so the token is minted for a specific, bounded transaction instead of a broad capability.
Mechanically, authorization_details is a JSON array of objects, each with a REQUIRED type that names the schema, plus the RFC's reusable common fields β locations (resource-server URIs), actions, datatypes, identifier, and privileges. The authorization server shows the detail on the consent screen and, critically, "MUST also return the authorization_details as granted by the resource owner and assigned to the respective access token." In Auth0's RAR implementation (Auth0 is Okta) you register the type, push it to the /oauth/par endpoint so the AS validates the type early, and a custom consent screen renders the request; the token response then carries the granted authorization_details β e.g. a money_transfer object with instructedAmount, sourceAccount, and destinationAccount β beside the access token.
This turns authorization from a capability into a contract: the resource server enforces authority from the token's authorization_details ("a transfer of exactly 2500 USD from β¦1234 to β¦9876, nothing else"), not from a broad scope claim. The blast radius of a stolen or misbehaving agent token collapses from "everything in the scope" to "one already-consented transaction."
Cross-pollination: an agent post-trained with RLVR (today's AI pill) optimizes relentlessly toward its reward and will exploit any unconstrained path β which is exactly why its credential must be bounded at the protocol layer. RAR bounds the action, not merely the resource class.
π¬ Deep Dive
Practitioner trap: the detail is only as strong as your enforcement. Auth0 validates the type but, in its own words, "you must implement validation for the JSON objects in authorization_details" β the AS does not understand your instructedAmount. If your resource server does not re-read the granted authorization_details from the access token and check it against the actual request body, you have shipped rich requests with coarse enforcement, which is worse than scopes because it looks safe. Two more: RAR must go through Pushed Authorization Requests (/oauth/par) β putting authorization_details on the front-channel URL is a leak and a length bomb; and JWE-encrypted access tokens surface only the type, so never design the RS to read the full detail out of an encrypted token client-side.
Staff-level framing:authorization_details is a per-transaction, machine-readable authorization record β the governance win. An access review can now answer "what exactly was this agent authorized to do," not "which broad scopes does it hold," and that granted detail is the artifact your SIEM and your auditor actually want. Blast radius: a leaked coarse-scope token is a capability; a leaked RAR token is the shape of a single, already-consented transaction. Rollout without boiling the ocean: define one type per high-risk action, leave coarse scopes on read-only paths, and gate the RAR consent behind asynchronous approval (CIBA) for the genuinely dangerous ones so a human authorizes before the agent moves money or deletes data.
RAR, scopes, and fine-grained authorization solve different layers β compose them. Scopes are coarse, static, per-client capability. RAR is per-request, transaction-shaped authority consented at issuance and carried in the token. Relationship-based fine-grained authorization (ReBAC/FGA) is a runtime check of "may this subject touch this specific object right now." An agent platform needs RAR at the token boundary and FGA at the data boundary; neither replaces the other, and shipping only scopes leaves both gaps open.
# 1) RAR must be pushed via PAR; the JSON below is url-encoded in the real request.
POST https://YOUR_OKTA_DOMAIN/oauth/par
Content-Type: application/x-www-form-urlencoded
client_id=YOUR_CLIENT_ID
response_type=code
redirect_uri=https://app.example.com/callback
authorization_details=[
{
"type": "payment_initiation",
"actions": ["initiate"],
"locations": ["https://api.example.com/payments"],
"instructedAmount": { "currency": "USD", "amount": 2500 },
"sourceAccount": "acct_...1234",
"destinationAccount": "acct_...9876"
}
]
# 2) /oauth/token response echoes the GRANTED detail, bound to the token:
# { "access_token": "<TOKEN>",
# "authorization_details": [ { "type": "payment_initiation",
# "instructedAmount": {"currency":"USD","amount":2500}, ... } ] }
# The resource server authorizes THIS transaction from the detail β not a scope.
Doc-verified against RFC 9396 Β§2 (authorization_details, required type, common fields, granted detail bound to the token) and the Auth0 RAR docs (PAR requirement, echoed token response), Sep 2026. Placeholders only; not executed in this run.
πΌ Market Signal
Okta's own agent story is built on this shape: Auth0 for AI Agents lists "Access Management" as fine-grained permissions and short-lived tokens that "restrict agent actions to provisioned tasks," with explicit user-consent gates for sensitive operations β that is RAR-shaped authority, not scopes. And the demand driver is quantified: Gartner (Jun 2025) projects 33% of enterprise applications will include agentic AI by 2028 (up from under 1% in 2024), yet predicts over 40% of agentic AI projects will be canceled by the end of 2027 for "escalating costs, unclear business value or inadequate risk controls." The risk-controls half of that sentence is where transaction-scoped authorization lives β the identity architect who can bound an agent's authority at the token, not just at the prompt, is on the surviving side of that statistic.
β‘ Action This Week
On a free Auth0 dev tenant, stand up RAR end-to-end in under two hours: register one authorization_details type (payment_initiation), enable PAR on a test app, run the authorization-code flow pushing a detail with a bounded amount, and confirm the /oauth/token response echoes the granted detail bound to the access token. Then add a five-line resource-server check that compares the token's authorization_details to the request, submit a larger amount, and watch it reject. Done = a screenshot of the token response containing authorization_details, plus a two-line note stating where your resource server enforces it. Post it to LinkedIn β "Why an AI agent's OAuth token should authorize one transfer, not 'payments' β a 90-minute RAR walkthrough" β the enforcement point reads as senior judgment, not tool trivia.
GRPO + Verifiable Rewards: Post-Training a Reasoning Model With No Critic and No Reward Model
π‘ Key Concept
Supervised fine-tuning teaches a model to imitate one gold trajectory; it cannot teach a model to find a better one. Reinforcement Learning with Verifiable Rewards (RLVR) flips that: for any task whose answer can be checked programmatically β the math answer matches, the unit tests pass, the JSON validates, the database ends in the right state β you drop the learned reward model entirely and give reward only when an automated verifier says the output is correct. The model explores many reasoning paths and reinforces the ones that pass. This is the recipe behind DeepSeek-R1, published in Nature (Sep 2025): DeepSeek-R1-Zero reached 77.9% pass@1 on AIME 2024 (from a 15.6% base), 86.7% with self-consistency β trained by pure RL, with no SFT stage at all.
Group Relative Policy Optimization (GRPO) is the algorithm that made this cheap enough to matter. Classic PPO needs a separate critic/value network roughly the size of the policy, doubling training memory. GRPO deletes the critic: for each prompt it samples a group of G completions, scores each with the verifier, and sets each completion's advantage to its reward normalized within the group β Γ = (r β mean(r)) / std(r). Above-average samples get pushed up, below-average down; the group is its own baseline. In Hugging Face TRL that is GRPOConfig(num_generations=G) plus one or more reward functions with the signature def reward_func(prompts, completions, **kwargs) -> list[float].
The reward design is the entire game. DeepSeek used rule-based rewards only β accuracy (a deterministic check: answer matching, code test cases) plus format (reasoning wrapped in <think>β¦</think>) β and deliberately refused neural reward models: "We abstain from applying neural reward models β¦ [they are] susceptible to reward hacking during large-scale RL." That is the lesson in one sentence β a model optimizing an imperfect reward finds the exploit, not the intent.
Cross-pollination: that same relentless-optimizer property is why an RLVR-trained agent's runtime authority must be bounded at the authorization layer (today's IAM pill on Rich Authorization Requests) β you cannot assume a capable agent stays within the spirit of a coarse scope.
π¬ Deep Dive
Practitioner trap: your verifier is the spec, and the model will read it literally. If the accuracy check can be satisfied without solving the task β a regex that matches a substring, a unit test with a trivial passing path, a format reward the model games by emitting empty <think></think> β GRPO finds that exploit within a few hundred steps, because it optimizes the reward you wrote, not the one you meant. Second trap is reward variance: if all G samples in a group earn the same reward, std = 0, the advantage is zero, and the step contributes no gradient. You need prompts that sit at the edge of the model's ability so the group spreads out β a dataset the model already aces teaches it nothing.
Staff-level framing: a GRPO run is a serving problem, not a backprop problem. Cost is dominated by generation β you sample num_generations completions (commonly 8β16) per prompt per step β so budget GPU-hours around rollout throughput and put a fast inference backend behind the sampler, not around the optimizer. Governance: a verifiable reward is a deterministic function you can inspect, version, and diff β auditable in a way a learned RLHF reward model is not, which matters in regulated settings. Blast radius: RL runs fail silently far from your dashboards (mode collapse, KL blow-up, reward hacking), so gate promotion on a held-out eval the reward function never touched β never on rising training reward, which is the number most likely to be lying to you.
Pick the post-training tool by whether correctness is checkable. SFT when you have gold outputs and want imitation or a format cold-start. DPO/RLHF when quality is subjective β helpfulness, tone, style β and you have preference pairs. RLVR/GRPO when correctness is programmatically verifiable β math, code, structured extraction, tool-call success β because then a verifier beats a learned reward model and needs zero human labels. Most production reasoning stacks are not either/or: they SFT a cold-start, then run GRPO on verifiable tasks, which is exactly DeepSeek-R1's pipeline after R1-Zero proved the pure-RL point.
from trl import GRPOTrainer, GRPOConfig
# Verifiable reward: no reward model, just deterministic checks.
def accuracy_reward(prompts, completions, ground_truth, **kwargs):
return [1.0 if extract_answer(c) == gt else 0.0
for c, gt in zip(completions, ground_truth)]
def format_reward(completions, **kwargs): # reasoning must be inside <think>...</think>
import re
pat = re.compile(r"<think>.+?</think>", re.DOTALL)
return [0.2 if pat.search(c) else 0.0 for c in completions]
cfg = GRPOConfig(
num_generations=8, # G: group size β the baseline IS the group mean
reward_weights=[1.0, 0.2], # accuracy dominates; format is a small nudge
scale_rewards="group", # advantage = (r - mean) / std within the group
beta=0.0, # KL-to-reference term off by default in current TRL
)
trainer = GRPOTrainer(
model="Qwen/Qwen2.5-3B-Instruct",
reward_funcs=[accuracy_reward, format_reward], # summed with reward_weights
args=cfg, train_dataset=ds,
)
trainer.train()
Doc-verified against the TRL GRPOTrainer docs (v1.13.0): reward-function signature, num_generations, reward_weights, scale_rewards, and beta=0.0 default β Sep 2026. extract_answer is an illustrative helper; not executed in this run.
πΌ Market Signal
RLVR is no longer a frontier-lab secret: DeepSeek-R1 shipped as an open-weight model with its recipe in a peer-reviewed Nature paper (Sep 2025), and the trainer (TRL) is open source β so the differentiator has moved from access to reward design and eval discipline. The compensation follows the scarcity: per Second Talent (Sep 2026), AI/ML engineers sit around a $189K US median and agent-orchestration roles near $209.5K, with "AI skills [paying] a 62% wage premium" (PwC 2026) and agentic-AI skills jumping from 0.06% to 0.23% of US job postings in a single year. The engineers who can shape model behavior with an RL loop β not just prompt one β are the ones pricing at the top of that band.
β‘ Action This Week
Run one GRPO loop on a small instruct model (Qwen2.5-0.5B or 3B) over a tiny verifiable task β GSM8K-style arithmetic where the reward is exact-answer-match plus a <think> format check β for a few hundred steps on a single GPU or a free Colab, logging training reward against a held-out accuracy. Then deliberately weaken the reward (accept any digit that appears anywhere) and watch training reward climb while held-out accuracy flatlines: reward hacking, reproduced in one chart. Done = a reward-vs-step curve beside the held-out accuracy line, plus a one-paragraph note on the reward function you wrote and the exploit you induced. Ship it to LinkedIn β "I trained a model to game my own reward function β here's why your verifier is your real spec" β which demonstrates RL judgment, not just RL vocabulary.
Token Vault: Letting an Agent Act on the User's Behalf Without Ever Holding the Credential
π‘ Key Concept
The moment an AI agent has to do something for a user β read their calendar, open a Jira ticket, post to Slack β it needs a token for a third-party API that the user, not the agent, owns. The naive answer is to have the agent store the user's Google or GitHub refresh token. That is the whole security problem in one sentence: a long-lived, broadly-scoped credential sitting inside prompt-reachable agent code.
Token Vault (part of Auth0 for AI Agents, now Okta-owned) removes the credential from the agent entirely: Auth0 stores the third-party access and refresh tokens in a dedicated vault and hands the agent a short-lived, per-connection token only at the moment of the call.
The mechanism underneath is standards-plain: OAuth 2.0 Token Exchange (RFC 8693). First the user runs a Connected Accounts flow β the agent redirects them to authenticate with the external provider (a standard OAuth consent), and on approval Auth0 links that external account to the user profile and stores the resulting refresh token in the vault. Afterwards the agent presents its own identity plus the user context and exchanges them for a freshly-minted, narrowly-scoped access token against that connection. The agent never sees, stores, or refreshes the raw credential; Auth0 owns the lifecycle.
What makes this an identity story rather than a secrets-management one is the shape of the exchange. In RFC 8693 the subject_token represents the human being acted for, and an optional actor_token represents the agent β a non-human identity β doing the acting. Present both and you get delegation: the issued JWT carries an act claim naming the agent, so every downstream call is attributable to "agent X acting for user Y," not to an anonymous service account. That act claim is the audit primitive that turns "the bot did it" into an answerable question.
Cross-pollination: today's AI pill covers A2A, which hands a task from one agent to another but says nothing about whose credentials the callee uses β on-behalf-of token exchange is exactly the identity layer that keeps that hop least-privilege and auditable.
π¬ Deep Dive
Practitioner trap: the blast radius is set at consent time, not call time. If the Connected Accounts flow asks for full gmail.modify when the agent only ever reads, every future exchange inherits that scope, and the vault will happily mint a token that can delete mail. Request the narrowest scopes the tool needs, per connection. Second trap: the token from getAccessTokenForConnection is the third-party token (Google's), not your app's session token β never log it, never send it anywhere but the target API. Third: a user revoking access at Google silently invalidates the vault entry, so exchanges start failing β you must handle the re-consent path, not assume the token is always there.
Staff-level framing: this is non-human-identity governance in miniature. The controls that matter for a platform review: (1) the agent holds no standing credential β tokens are short-lived and centrally revocable, so an over-permissioned or compromised agent is contained to its consented scopes; (2) the act claim gives you a per-agent, per-user audit trail instead of a shared service account no one can attribute; (3) high-risk actions get gated by pairing the vault with asynchronous authorization (CIBA) so a human approves before the agent moves money or deletes data. Identity, not a secrets manager, is what makes agent access reviewable.
Delegation vs impersonation is a design decision, not a default. Present only subject_token and the agent effectively becomes the user (impersonation) β cheaper, but the audit trail loses the agent. Present subject_token and actor_token and you get delegation with the act claim intact. For anything that touches production data, choose delegation and make the reviewer able to answer "which agent, acting for whom."
π οΈ Artifact β the exchange, and the SDK that hides it
POST https://YOUR_OKTA_DOMAIN/oauth/token
Content-Type: application/x-www-form-urlencoded
grant_type=urn:ietf:params:oauth:grant-type:token-exchange
subject_token=USER_REFRESH_TOKEN # the human the agent acts for
subject_token_type=urn:ietf:params:oauth:token-type:refresh_token
actor_token=AGENT_ACCESS_TOKEN # the agent's own (non-human) identity
actor_token_type=urn:ietf:params:oauth:token-type:access_token
requested_token_type=urn:ietf:params:oauth:token-type:access_token
# Auth0 selects the federated connection (google-oauth2 / github / slack) and
# returns a third-party-scoped access token whose JWT carries an "act" claim.
Doc-verified against RFC 8693 Β§2.1 (token exchange) and the Auth0 Token Vault docs (auth0.com, Sep 2026). Placeholders only; not executed in this run.
// Agent tool: read the user's calendar without ever seeing a Google credential
const googleToken = await auth0.getAccessTokenForConnection({
connection: "google-oauth2",
scopes: ["https://www.googleapis.com/auth/calendar.readonly"]
});
// Not connected yet? The SDK raises a Token Vault interrupt β surface a
// Connect Account / re-consent prompt to the user; do not swallow the error.
const res = await fetch("https://www.googleapis.com/calendar/v3/calendars/primary/events", {
headers: { Authorization: "Bearer " + googleToken }
});
Illustrative β Auth0 AI SDK surface (getAccessTokenForConnection, connection ids per Auth0 Token Vault docs, Sep 2026); not executed.
π§ Recall
A week ago the MCP authorization pill made your MCP server an OAuth protected resource β inbound. Token Vault is the mirror image. Which direction does it run, and why can't the agent just reuse the token it already presented to your MCP server?
Show answer
MCP authorization protects an inbound resource: the agent presents a token to your server, whose aud (audience) is your server. Token Vault is outbound: the agent needs a token for a third party (Google, Slack), whose audience is that external API and which Google will only trust if it was minted from a user-consented refresh token via RFC 8693 exchange. The inbound token is useless for the outbound call β wrong audience, wrong issuer, wrong scopes β which is exactly why a federated exchange against the vault exists instead of token passthrough.
πΌ Market Signal
Identity-for-agents has shipped, not just shipped-a-blog-post. Auth0 for AI Agents went generally available on Nov 19, 2025, structured around four pillars β user authentication, Token Vault, asynchronous authorization (CIBA), and fine-grained authorization for RAG β with 35+ prebuilt app integrations (GitHub, Slack, Google Drive, Jira) and first-class SDKs for LangChain, LlamaIndex, the Vercel AI SDK, and Cloudflare Agents. Read against this pill, the hiring signal is specific: the frameworks now ship the token plumbing, so the differentiated skill is designing the delegation model β which connections, which scopes, delegation vs impersonation, and where CIBA gates a high-risk action. That is identity-architecture scope for agentic systems, and it is exactly the non-human-identity governance work that generalist "prompt engineers" do not touch.
β‘ Action This Week
Stand up one agent tool that reads your own Google Calendar through Token Vault: configure a Google social connection, run the Connect Account flow once, then call getAccessTokenForConnection({ connection: "google-oauth2", scopes: ["...calendar.readonly"] }) and list your next event β and deliberately trigger the not-connected path to see the re-consent interrupt fire. Request read-only scope on purpose. Done = a screenshot of the agent naming your next meeting, plus a one-line note stating the agent code never held a Google credential (only a short-lived, vault-exchanged token). Ship it as a LinkedIn post β "My agent read my calendar and never touched my Google password β here's the on-behalf-of token exchange that makes that true" β the least-privilege framing reads as senior judgment.
A2A v1.0: How Opaque Agents Discover Each Other and Delegate Whole Tasks
π‘ Key Concept
MCP connects one agent down to its tools and data. A2A (Agent2Agent) connects agents across to each other, so a client agent can hand an entire task to a remote agent it did not build and cannot see inside. The word doing the work in the spec is opaque: the remote agent advertises what it can do, never its tools, models, or internal reasoning. You delegate an outcome and consume updates; you do not orchestrate someone else's internals.
Discovery is a flat file at a well-known URL. Per the A2A agent-discovery docs, every agent publishes an Agent Card at /.well-known/agent-card.json (RFC 8615), declaring its name, service url, capabilities (streaming, pushNotifications), securitySchemes, and a list of skills β each with an id, inputModes, and outputModes. Fetch the card, pick a skill, and you know how to authenticate and what content types to send, with no shared SDK or prior integration.
The other half is that a real task is not request/response. A2A models a Task with an explicit lifecycle β submitted β working β completed, with input-required and auth-required as first-class interruptions and failed/canceled/rejected as terminal states. Transport is pluggable β JSON-RPC 2.0, gRPC, or HTTP+JSON/REST β with progress delivered as Server-Sent Events or pushed to a webhook for long jobs.
Cross-pollination: A2A moves the task but is deliberately silent on whose credentials the callee uses to touch a real system β pair it with today's IAM pill (on-behalf-of token exchange) to keep each hop least-privilege and attributable.
π¬ Deep Dive
Practitioner trap: treating message/send as a synchronous call. The Agent Card advertises skills, not tools β if you find yourself trying to introspect the remote agent's functions you have misread the protocol. A real delegated task goes submitted β working and may sit in input-required (the remote agent needs more from you) or auth-required before it ever reaches completed. Consume the SSE stream or poll the task; a client that reads only the first response and moves on will drop long-running work on the floor.
Staff-level framing: A2A and MCP are complementary layers, not competitors β MCP is vertical (agentβtools), A2A is horizontal (agentβagent), and a remote A2A agent may use MCP internally. The governance surface is the Agent Card: its securitySchemes (OAuth2, OpenID Connect, mutual TLS) is where you enforce authentication at the boundary, but the card is also a supply-chain artifact. An unsigned card fetched from a well-known URL is trust-on-first-use; v1.0's move to cryptographically signed cards exists precisely so that "which agent am I actually delegating to" has a verifiable answer. Pin protocolVersion, verify signatures, and treat every remote agent as an untrusted boundary.
For long jobs, do not hold the stream open. Register a push-notification config and A2A will HTTP POST task updates to your webhook, so a 20-minute reconciliation does not depend on a live SSE connection surviving. But a webhook that accepts unauthenticated POSTs is an SSRF and spoofing sink β use the config's authenticated-callback support and validate the sender, or you have traded a dropped connection for a forged "task completed."
π οΈ Artifact β an Agent Card and the task it starts
Structure per A2A spec v1.0 and the agent-discovery docs (a2a-protocol.org, Sep 2026); field names doc-verified, values illustrative.
// Client delegates a task, then consumes streamed state β not request/response
POST /a2a (JSON-RPC 2.0) method: "message/send"
-> { "kind": "task", "id": "t-9", "status": { "state": "submitted" } }
event: status-update state: "working"
event: status-update state: "input-required" # remote agent pauses to ask
event: status-update state: "completed" # terminal; artifact attached
Illustrative β task states doc-verified against A2A spec v1.0; JSON-RPC method and event shapes simplified for space, not executed.
π§ Recall
A week ago we covered MCP authorization. If MCP already lets an agent reach tools and data, why does A2A exist at all β aren't they solving the same problem?
Show answer
Different axes. MCP is vertical: it standardizes how one agent reaches its tools, resources, and context. A2A is horizontal: it standardizes how an agent delegates a whole task to another, opaque, agent β one it cannot see inside and did not build. They compose rather than compete: the A2A remote agent you delegate to may itself speak MCP to its own tools. A2A adds the pieces MCP's tool-call model does not address β agent discovery via Agent Cards, a Task lifecycle with human-in-the-loop input-required states, and cross-organization trust boundaries.
πΌ Market Signal
A2A has crossed from Google project to industry substrate. Per the Linux Foundation (Apr 9, 2026), one year after donation the protocol has 150+ supporting organizations and 22,000+ GitHub stars, and is wired into Azure AI Foundry and Copilot Studio, AWS Bedrock AgentCore Runtime, and Google Cloud β the same LF announcement notes a companion Agent Payments Protocol (AP2) already backed by 60+ organizations. Google donated A2A to the Linux Foundation on Jun 23, 2025 with AWS, Cisco, Microsoft, Salesforce, SAP, and ServiceNow as founding partners. The career read: multi-agent systems are becoming a cross-vendor integration discipline, and the scarce skill is designing the Agent Card contract, the task-lifecycle handling, and the trust boundary between agents β platform work, not prompt work.
β‘ Action This Week
Publish a minimal Agent Card at /.well-known/agent-card.json for a toy agent with exactly one skill, then stand up a message/send handler that returns a Task and streams working β completed over SSE. Validate the card against an A2A SDK (Python, JS, Go, Java, .NET, or Rust) so you know the shape is real. Done = a curl of your well-known Agent Card plus a screenshot of one task moving submitted β working β completed β and, for extra credit, a deliberately triggered input-required pause. Ship it as a LinkedIn post β "I made two of my agents talk over A2A; here's the Agent Card and the task lifecycle that made it work" β a runnable card is a stronger portfolio artifact than any diagram.
Okta as Code: Governing the Tenant with Terraform β and the Blast Radius the Provider Won't Protect You From
π‘ Key Concept
Clicking through the Okta Admin Console is how identity configuration rots: nobody can say who changed a sign-on rule, when, or why, and there is no way to roll back a bad edit except to remember what it used to be. Identity-as-Code replaces the console as the source of truth with the official okta/okta Terraform provider, so every group, app, authenticator, and policy rule lives in version-controlled HCL. Okta's own product team frames the payoff bluntly: managing configuration this way turns it "from a manually maintained configuration into a version-controlled, peer-reviewed, continuously deployable system" (Okta, Apr 23, 2026).
The three levers that matter are all governance, not syntax. Version control means every change to a network zone or risk-response rule is a commit β a full audit trail plus rollback. Peer review routes a change to your MFA policy through the same pull-request approval, CI, and deploy pipeline as application code, so a two-person rule on production identity is enforced by the tooling, not by hope. And drift detection falls out of the plan/apply cycle: terraform plan compares declared intent against the live tenant and shows any out-of-band change. The provider now covers even Identity Threat Protection β okta_network_zone, okta_entity_risk_policy_rule, okta_session_violation_policy_rule, okta_security_events_provider β so your detection-and-response posture is reviewable code, not console tribal knowledge.
The catch that separates an architect from a scripter: Terraform is a deployment tool, not a safety net. Its drift detection only sees resources it manages, and only when you run a plan; a snapshot of state is not a backup; and there is no sandbox-clone or tenant-restore primitive in the provider. A single mis-declared sign-on policy applied to production can lock out your entire workforce, and Terraform will happily do it. Cross-pollination: today's AI pill instruments the agent side with OpenTelemetry β identity-as-code (who/what may act) and telemetry-as-code (what they did) are the two halves of an auditable agentic platform.
π¬ Deep Dive
Practitioner trap β pin the provider; one bad version bricks your applies. The provider maintainers currently warn "DO NOT UPGRADE to v6.14.0. The current highest recommended version is v6.15.0." An unpinned required_providers block silently resolves to whatever is latest at terraform init time, so a routine CI run can pull a broken release and fail β or worse, misbehave β against production identity. Pin with version = "~> 6.15" and bump deliberately in a reviewed PR, never implicitly.
Practitioner trap β drift detection is blind to what you don't manage, and only fires when you look. Per an acsense analysis (Nov 25, 2025), "Terraform drift detection works only for resources it manages and only when running terraform plan." An admin who edits a policy in the console β or creates a brand-new app Terraform never imported β produces drift your state file cannot see. The fix is operational, not declarative: run terraform plan -detailed-exitcode on a schedule (exit code 2 = drift) and alert on it, and reconcile the whole tenant into state so "unmanaged" is a deliberate exception, not a blind spot.
Authenticate the pipeline as a scoped OAuth service app, not a static API token. A long-lived Okta API token in CI is a super-admin bearer secret with no least-privilege story. Use the provider's OAuth 2.0 mode β a service app with a private key and only the scopes the config touches (e.g. okta.policies.manage, okta.apps.manage) β sourced from a secrets manager, never committed. This is also what keeps the validator's secret-scan and any security review from flagging your repo.
Staff-level framing β Terraform ships the change; it does not own recovery, so design the safety layer around it. As acsense puts it, "Terraform state is not a backup" β there is no snapshot, no tenant rollback, no standby tenant, and "a single policy mistake can lock out your entire workforce." The architect's deliverable is a change-control operating model: a non-prod Okta org that the same code applies to first, a mandatory plan-review gate, a break-glass admin excluded from the risky policies, an out-of-band config export for point-in-time recovery, and blast-radius limits (never let one apply touch every authentication policy at once). Config-as-code is the control plane; the recovery plane is yours to build.
π οΈ Artifact β scoped provider, pinned version, drift as a CI gate
terraform {
required_providers {
okta = { source = "okta/okta", version = "~> 6.15" } # pin: v6.14.0 is broken, do not float
}
backend "s3" { # remote, locked, encrypted β state is shared truth, not a laptop file
bucket = "acme-tfstate"
key = "okta/prod.tfstate"
dynamodb_table = "tf-locks"
}
}
provider "okta" { # OAuth 2.0 service app + private key, NOT a static API token
org_name = "YOUR_OKTA_ORG"
base_url = "okta.com"
client_id = var.okta_client_id
scopes = ["okta.policies.manage", "okta.apps.manage"] # least privilege
private_key = var.okta_private_key # injected from Vault/CI secret; never in git
}
# Security posture as reviewed code: a phishing-resistant app sign-on rule
resource "okta_app_signon_policy_rule" "require_hardware_mfa" {
policy_id = okta_app_signon_policy.mcp_tools.id
name = "require-phishing-resistant"
factor_mode = "2FA"
# ignore_changes ONLY for attributes intentionally edited out-of-band β each one is a documented drift blind spot
# lifecycle { ignore_changes = [users_excluded] }
}
Syntax-checked against the okta/okta provider registry docs (registry.terraform.io, Sep 2026). Placeholders (YOUR_OKTA_ORG, var.*) β no real org, key, or token values.
# Scheduled drift alarm (CI): exit 0 = in sync, 2 = drift, 1 = error
terraform plan -detailed-exitcode -lock-timeout=120s || \
test $? -eq 2 && echo "DRIFT: live Okta tenant differs from declared config" >&2
Illustrative CI snippet; -detailed-exitcode is documented Terraform behavior.
π§ Recall
Five days ago you set up Okta as the authorization server for MCP, where the custom authz server's audience must equal the MCP server's canonical URI. Why is defining that custom authorization server in Terraform a governance win, not just convenience?
Show answer
Because the audience binding is the security control, and a control that lives only in a console click is un-reviewable and un-auditable. Declaring the okta_auth_server and its audience/scopes in HCL means a change to the audience β the exact value that stops the confused-deputy attack β becomes a diff in a pull request that a second engineer must approve, with a commit trail showing who changed it and when. It also makes the whole agent-authz setup reproducible across non-prod and prod, so you test the token flow in a sandbox org before it ever touches production. Code turns a silent, high-blast-radius setting into a governed change.
πΌ Market Signal
Identity-as-code sits at the intersection of two premium skills β IAM depth and platform engineering β and the comp reflects the scarcity. Per the Start with Identity IAM Salary Guide (updated Jun 24, 2026), US base ranges run ~$150Kβ$200K for a Senior IAM Engineer and ~$180Kβ$240K+ for an Identity Architect, with Head-of-Identity roles at ~$200Kβ$300K+ β and total comp "runs meaningfully above base at larger and venture-backed companies." The guide's stated reason maps exactly to this pill: identity pays a premium because "the work is business-critical, increasingly regulated, and the experienced talent pool is small." The engineer who can put an Okta tenant under Terraform and articulate the recovery/blast-radius model Terraform doesn't give you is arguing directly for the architect band, not the admin band.
β‘ Action This Week
In a non-production Okta org, terraform import one existing sign-on or app policy into state, then make a benign change in the Admin Console and run terraform plan -detailed-exitcode to watch it surface as drift. Pin the provider to ~> 6.15 and wire the plan into a scheduled CI job that alerts on exit code 2. Done = a screenshot of the plan output showing the console edit as a proposed diff, plus a one-paragraph note on what your tenant's recovery plan is outside Terraform. That note β "here's what config-as-code does and doesn't protect against" β is a strong LinkedIn post or interview talking point that signals architect-level judgment, not just tool familiarity.
Instrumenting the Agent: OpenTelemetry's GenAI Conventions β Powerful, and Still Not Stable
π‘ Key Concept
An LLM call is a black box priced by the token, and an agent is a tree of them β model calls, tool executions, retrievals β where cost and latency hide in the aggregate. OpenTelemetry's GenAI semantic conventions make that tree observable with the same tracing standard your backend already uses, so an LLM span carries typed attributes instead of ad-hoc logs. Per the OpenTelemetry GenAI observability guide (2026), a call records gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, and gen_ai.response.finish_reasons, while two metrics β gen_ai.client.operation.duration and gen_ai.client.token.usage β let you "estimate per-request cost, catch token-hungry prompts before they hit production, [and] detect latency regressions."
The 2026 conventions model a whole agent run as a span tree, not a single call: gen_ai.operation.name spans the full lifecycle β invoke_agent, execute_tool, retrieval, invoke_workflow β so you can attribute the cost and latency of an entire agent trajectory to the user request that triggered it, and see which tool call in the tree blew the budget. Because it is vendor-neutral OTel, the same spans flow to Jaeger, Grafana, or any of the LLM-native backends without re-instrumenting.
The trap that will bite a team that treats this as "just add the SDK": the GenAI conventions are still Development status β not stable. As of mid-2026 no GenAI span, metric, or attribute is marked Stable, and the whole namespace was carved into a dedicated repo with no tagged release to pin to. Attribute names have already churned across versions. Build on it, but wrap it. Cross-pollination: pair these gen_ai.* traces with Okta's System Log (today's IAM pill manages Okta as code) β the token that authorized the agent and the tool calls it then made become one queryable audit trail.
π¬ Deep Dive
Practitioner trap β the attribute names move under you; a copy-pasted dashboard rots. The conventions have already renamed core fields: token attributes went from prompt_tokens/completion_tokens to gen_ai.usage.input_tokens/output_tokens, and gen_ai.system became gen_ai.provider.name. Per a July 2026 status review, the whole gen_ai.* namespace was moved to a dedicated repo with no tagged release yet β so there is no stable schema URL to pin. A Grafana panel hard-coded to last year's attribute names silently goes blank after an SDK bump. Pin your instrumentation library version and translate into an internal, stable data model you control.
Cost is a first-class signal, not an invoice surprise. The gen_ai.client.token.usage histogram is dimensioned by gen_ai.token.type (input vs output) and gen_ai.request.model β multiply by per-model rates and you have live per-request, per-model cost, plus the ability to spot the one prompt whose input tokens 10Γ'd after a template change. Because output tokens usually dominate cost, splitting the histogram by token type is what turns "our bill went up" into "this endpoint's completions grew 3Γ."
Staff-level framing β your trace store is now a regulated data surface; govern content capture like the blast radius it is. By design, OTel captures no prompt or tool-argument content by default because it "can contain sensitive data"; capturing gen_ai.input.messages/gen_ai.output.messages is an explicit opt-in. Flip it on org-wide and you have just piped user PII, secrets, and copyrighted inputs into a telemetry backend with its own retention, access, and export story β a compliance incident waiting to happen. The architect's move is a policy: content capture off by default, opt-in only for scoped debugging with redaction and short retention, and the trace pipeline treated as in-scope for the same data-governance review as any other PII store. Reliability tooling that leaks data is a net negative.
π οΈ Artifact β a GenAI-conventions span, with content capture deliberately off
from opentelemetry import trace, metrics
tracer = trace.get_tracer("agent")
meter = metrics.get_meter("agent")
tok = meter.create_histogram("gen_ai.client.token.usage") # Development-status names β pin your semconv version
with tracer.start_as_current_span("chat gpt-4o-mini") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.provider.name", "openai") # was gen_ai.system before v1.37
span.set_attribute("gen_ai.request.model", "gpt-4o-mini")
resp = client.chat.completions.create(model="gpt-4o-mini", messages=msgs)
it, ot = resp.usage.prompt_tokens, resp.usage.completion_tokens
span.set_attribute("gen_ai.usage.input_tokens", it)
span.set_attribute("gen_ai.usage.output_tokens", ot)
span.set_attribute("gen_ai.response.finish_reasons", [resp.choices[0].finish_reason])
tok.record(it, {"gen_ai.token.type": "input", "gen_ai.request.model": "gpt-4o-mini"})
tok.record(ot, {"gen_ai.token.type": "output", "gen_ai.request.model": "gpt-4o-mini"})
# DO NOT set gen_ai.input.messages / gen_ai.output.messages unless content capture is
# explicitly opted in and redacted β prompts/completions can carry PII and secrets.
Syntax-checked against the OTel GenAI semantic conventions (Development status, opentelemetry.io, 2026). Illustrative model/SDK values; not executed in this run.
π§ Recall
Five days ago, constrained decoding guaranteed a model's JSON always parses. If the shape is already guaranteed, what does production observability add that constrained decoding can't?
Show answer
Constrained decoding guarantees syntax, never truth: a schema-valid route: "billing" can still be the wrong route, and a valid object can carry a hallucinated account_id. Observability is where you catch the wrong-but-valid outputs β by tracing finish_reasons (truncation vs clean stop), watching token/latency/cost regressions after a prompt change, and, when content capture is scoped-on, sampling actual outputs for eval. Decoding makes the failure mode "invalid JSON" impossible; observability is how you see the failure modes that remain. They're complementary layers, not substitutes.
πΌ Market Signal
LLM observability has crossed from nice-to-have to infrastructure. Per The Business Research Company's LLM Observability Platform report (Feb 2026), the market grows from $1.97B in 2025 to $2.69B in 2026 (36.3% CAGR) and on to $9.26B by 2030, with forward drivers named as "agentic workflows, stricter AI governance requirements, token analytics for cost optimization, and integration of observability with devops toolchains." Read against this pill: the differentiator is no longer wiring up a dashboard β vendors ship that β but owning the vendor-neutral OTel instrumentation, the cost model, and the data-governance policy for what traces may capture. That is a platform/Staff scope, and it is exactly the "integrate observability with the devops toolchain" work the report says is driving spend.
β‘ Action This Week
Instrument one real LLM call with the gen_ai.* attributes above, export the span to a local Jaeger (or any OTLP collector), and record the token.usage metric. Then compute a cost number from the token counts and per-model rates, and build one panel: cost-per-request over time, split by gen_ai.request.model. Keep content capture off. Done = a screenshot of a trace showing gen_ai.usage.input_tokens/output_tokens and a derived $/request figure, plus a one-line note on which semconv/library version you pinned. Ship it as a LinkedIn post β "I put my agent's dollar cost on a dashboard with OpenTelemetry (and left the prompts out on purpose)" β the privacy-by-default choice reads as senior judgment.
Okta as the Authorization Server for MCP: Tokens That Scope an Agent to One Tool
π‘ Key Concept
The June 2025 revision of the MCP authorization spec made a decision that quietly hands the whole agentic-tool market to identity vendors: an MCP server is an OAuth 2.1 resource server, not an authorization server. It stops minting its own tokens and delegates authentication to an external authorization server β which is exactly the slot Okta fills. The agent's host (the "MCP client") gets a token from Okta and presents it to the MCP server, which only has to validate it.
The discovery handshake is fully mechanical. The client calls the MCP server with no token and gets a 401 whose WWW-Authenticate header carries a resource_metadata URL. It fetches that Protected Resource Metadata (RFC 9728) document from /.well-known/oauth-protected-resource, learns which authorization server the resource trusts (your Okta org), runs standard AS metadata discovery, and does a PKCE authorization-code flow. The non-negotiable detail: per RFC 8707 Resource Indicators, the client MUST send a resource parameter β the canonical MCP server URI β on both the authorization and token requests, and the server MUST validate that the token's aud is itself. That single parameter is what turns a generic "valid Okta token" into a token scoped to one specific tool server.
Okta's own answer goes a step past the base spec. Its Cross App Access (XAA) and the Identity Assertion Grant (ID-JAG) exchange the user's session for short-lived, audience-restricted tokens tied to specific MCP tools, propagating sub, azp, aud and an act claim so on-behalf-of chains stay auditable β the agent never holds a broad standing token. Cross-pollination: today's AI pill covers constrained decoding β the token authorizes the agent to reach the tool, but only schema-valid arguments get past the MCP server's input contract.
π¬ Deep Dive
Practitioner trap β a token that is "valid" is not a token that is "for you". The spec calls out the confused-deputy attack explicitly. If your Okta custom authorization server mints a generic access token and the MCP server accepts any signed, unexpired Okta token, a hostile MCP server can replay a token it was handed to a different MCP server and act as the user. The only defense is audience binding: the client sends resource (RFC 8707), Okta stamps aud, and every resource server rejects tokens whose aud is not its own canonical URI. Skipping aud validation leaves you exploitable while every log line reads "authorized".
The resource parameter is mandatory even when the AS ignores it. The client MUST send resource on both the authorization and token requests, and MUST send it regardless of whether the authorization server advertises support. On the Okta side this means a custom authorization server whose audience equals the MCP server's canonical URI (e.g. https://mcp.example.com/mcp) β not the default api://default. Get the audience wrong and every call 401s with a token that is otherwise perfectly valid; this is the number-one silent misconfiguration in MCP-behind-Okta setups.
Least privilege is enforced by the 403, not the 401. A missing token is a 401; a token missing a scope is a 403 insufficient_scope whose WWW-Authenticate names the scope needed. That lets you issue narrow tokens (tools:read) and step-up only when a specific tool needs more β so a compromised agent token can invoke one tool, not the whole server. Don't front-load every scope "to be safe"; that rebuilds the standing-privilege problem you were trying to kill.
Staff-level framing β treat agent tokens as non-human identities with a blast radius, not as API convenience. The governance levers are token TTL (minutes, not hours β Okta's own guidance for agent tokens), per-tool audience/scope, and the act claim for on-behalf-of audit. Wire the MCP server's token-validation failures and Okta's issuance events into your System Log stream so "which agent called which tool for which user" is queryable. The architect's deliverable is a token lifecycle policy for agents β issuance, audience, TTL, revocation, audit β not a working demo. That policy is what a security review actually asks for.
# 1) Agent calls the MCP server with no token -> resource server challenges
HTTP/1.1 401 Unauthorized
WWW-Authenticate: Bearer resource_metadata="https://mcp.example.com/.well-known/oauth-protected-resource",
scope="tools:read"
# 2) Protected Resource Metadata (RFC 9728) points the client at Okta as the AS
GET https://mcp.example.com/.well-known/oauth-protected-resource
{
"resource": "https://mcp.example.com/mcp",
"authorization_servers": ["https://YOUR_OKTA_DOMAIN/oauth2/ausMcpToolsId"],
"scopes_supported": ["tools:read", "tools:call"],
"bearer_methods_supported": ["header"]
}
# 3) Token request to Okta MUST carry the RFC 8707 resource param (audience binding)
curl -X POST https://YOUR_OKTA_DOMAIN/oauth2/ausMcpToolsId/v1/token \
-d grant_type=authorization_code -d code=$CODE -d code_verifier=$PKCE \
-d client_id=$CLIENT_ID \
--data-urlencode resource=https://mcp.example.com/mcp \
--data-urlencode scope="tools:read"
# 4) MCP server validates aud == its own canonical URI; wrong audience -> 401.
# Omitting this check is the confused-deputy hole the spec warns about:
# a token minted for resource=https://other.example.com/mcp MUST be rejected here.
Doc-verified against the MCP authorization spec (modelcontextprotocol.io, Aug 2026) and Okta custom authorization server docs. Placeholders (YOUR_OKTA_DOMAIN, ausMcpToolsId) β no real tenant identifiers.
π§ Recall
Nine days ago you looked at the Okta agent registry + credential vault. If the registry already establishes who an agent is and holds its secrets, why does MCP still need a per-call, audience-bound OAuth token?
Show answer
Different layers. The registry/vault answers identity and credential custody β this agent exists, here are its stored secrets. The MCP OAuth token answers authorization for one action β this agent may call this tool server, for this user, for the next few minutes. Registry identity is standing; the token is per-request and audience-restricted so a leak can't be replayed against another resource. You need both: the vault to hold the client credential that obtains tokens, and RFC 8707 audience binding to keep each issued token scoped to a single MCP server.
πΌ Market Signal
Agentic identity is a net-new specialization, and the vendor timeline shows it forming in real time. Okta announced Cross App Access (XAA) β a standards-based protocol extending OAuth to govern agent and app-to-app access β on June 23, 2025 (Okta newsroom), targeting availability for select Okta Platform customers in Q3 2025, and its engineering guidance (Sep 23, 2025) is already prescribing minutes-not-hours tokens bound to specific MCP tools. Translation for a Staff/Architect target: "secure AI agents accessing tools via MCP + OAuth" is a job description that essentially did not exist 18 months ago. Being the person who can whiteboard the 401βPRMβaudience-bound-token flow β and name the confused-deputy failure mode β is first-mover positioning while most IAM engineers are still framing this as "just SSO for a script".
β‘ Action This Week
Stand up a toy MCP server (any of the reference SDKs) behind an Okta custom authorization server whose audience is the server's canonical URI. Walk the full loop and, crucially, prove the negative case: mint a token with resource pointing at a different URI and confirm the server 401s it. Done = a terminal capture (or three screenshots) showing (1) the Protected Resource Metadata JSON, (2) a successful call with a correctly-audienced token, and (3) a rejected call with a wrong-audience token. Post the negative case as a LinkedIn diagram β "Why a valid Okta token still gets your agent a 401 from an MCP server" β it reads as protocol depth, not admin trivia.
Constrained Decoding: JSON That Always Parses β and Why That Isn't Enough
π‘ Key Concept
Prompting "return JSON" fails a nonzero fraction of the time β a prose preamble, a trailing comma, a markdown fence β and in a multi-step agent one malformed tool call halts the whole run. Constrained (guided) decoding removes the syntactic failure mode at the sampling layer instead of hoping the model behaves. At every decode step the engine consults a grammar/finite-state machine compiled from your schema and, per the Red Hat vLLM analysis, "masks invalid tokens, ensuring only tokens that comply with the defined constraints remain candidates for sampling." The output is guaranteed to parse against the schema.
This is now first-class in the serving stack, not a bolt-on library. In vLLM you pass structured_outputs with a nested json, choice, regex or grammar key (or the OpenAI-compatible response_format: {"type":"json_schema"}), backed by xgrammar or guidance. The grammar backend caches compiled schemas, which is why repeated calls with a stable schema pay almost nothing per token.
Here is the trap that separates a Staff engineer from someone who just discovered the flag: constraining syntax is not constraining semantics. A schema forces a valid enum but cannot force the right enum, and forcing the model to emit the final object before it can reason measurably degrades answers. The production pattern is reason-first-then-constrain β a free-text reasoning field ahead of the payload, or two calls. Cross-pollination: this is the other half of today's IAM pill β the OAuth token proves an agent may call an MCP tool; constrained decoding guarantees the arguments it sends are schema-valid so the MCP server parses them instead of returning a 400.
π¬ Deep Dive
Practitioner trap β guided_json is deprecated; copy-pasted examples silently rot. vLLM moved the API: as of v0.12.0 the old guided_json / guided_choice / guided_regex parameters are deprecated in favor of the nested structured_outputs object (or response_format). Tutorials and Stack Overflow answers written in 2024β2025 still use the old names, so a pipeline built on them warns now and breaks later. Pin your vLLM version and grep your codebase for guided_ before an upgrade.
Practitioner trap β grammar compilation is a TTFT tax on cold schemas. Per Red Hat's benchmarking, xgrammar"caches well, excels at long generations" and shines "when grammars are reused"; the flip side is that a never-before-seen schema pays a compile cost on the first request. So per-request dynamically generated schemas and unbounded string fields (which can run to max_tokens) are the latency foot-guns. Use xgrammar for stable, reused schemas; reach for guidance when schemas are unpredictable or one-off, since it offers "lower latency per request".
Staff-level framing β structured outputs are a reliability contract at a system boundary, so put the boundary in the right place. Constrained decoding turns "the LLM emits JSON" from a probabilistic hope into an invariant, which is precisely what lets you delete a defensive JSON-repair/retry layer and build multi-step agents that don't wedge on a bad tool call. But keep value validation downstream: the grammar guarantees shape, not truth β a constrained enum can still be the wrong choice, and an emitted account_id can still not exist. Operationally, log the masked-token rate as a model-health signal: a spike means schema drift or a model that can no longer satisfy the contract, and that is your early warning, not the 3am pager.
π οΈ Artifact β reason-first, then constrain the payload
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
# vLLM >= 0.12.0: use structured_outputs (guided_json is deprecated).
# Note the reasoning field FIRST so the model can think before it is constrained.
schema = {
"type": "object",
"properties": {
"reasoning": {"type": "string"}, # scratchpad, unconstrained
"route": {"type": "string", "enum": ["billing", "tech", "sales"]},
"confidence": {"type": "number", "minimum": 0, "maximum": 1},
},
"required": ["reasoning", "route", "confidence"],
"additionalProperties": False,
}
resp = client.chat.completions.create(
model="Qwen/Qwen3-8B",
messages=[{"role": "user", "content": "Route this ticket: 'my card was double charged'"}],
extra_body={"structured_outputs": {"json": schema}},
)
# Output is GUARANTEED to parse against `schema`. Still validate `route` against
# business rules downstream: the grammar enforces one-of-three, not the CORRECT one.
Syntax-checked against the vLLM structured-outputs docs (docs.vllm.ai, Aug 2026). Illustrative model/endpoint values; not executed in this run.
π§ Recall
Nine days ago you covered the multi-agent supervisor pattern. Why is constrained decoding especially load-bearing for the supervisor's routing/handoff decision?
Show answer
The supervisor's handoff target must be exactly one of a fixed set of downstream agents. Free-form generation can hallucinate a target that doesn't exist ("finance_bot" when only billing/tech/sales exist), and now your router dispatches into the void. A choice/enum constraint makes an invalid target unsampleable, so the routing edge is guaranteed to land on a real node β the graph's control flow becomes type-safe. It's the difference between validating the router's output after the fact and making the invalid output impossible to emit.
πΌ Market Signal
Reliable structured output has quietly become table stakes: schema-guided decoding ships inside the dominant open inference stack (vLLM exposes it natively with xgrammar/guidance backends, per its docs), so "make the model return valid JSON" is no longer a differentiator β owning the reliability/latency trade-off is. On comp: per the KORE1 AI Engineer Salary Guide (updated Aug 5, 2026, drawing on Built In 2026, Glassdoor Feb 2026 and Levels.fyi via Exceeds AI Jan 2026), US AI engineers run $145Kβ$310K base, with senior at $180Kβ$280K base ($220Kβ$350K+ total) and Staff/Principal at $250Kβ$400K+ base; fully-remote roles land at $155Kβ$210K base. The premium sits with engineers who can reason about the serving layer β masking, grammar caching, TTFT β not just call an API.
β‘ Action This Week
Find one flaky "return JSON" prompt in your codebase and wrap it with structured_outputs (vLLM) or response_format: json_schema (OpenAI-compatible), using a schema that puts a reasoning field before the payload. Run it ~50 times before and after and record the parse-failure rate and the p50 latency delta. Done = a small before/after table showing parse-success going from <100% to 100% and the measured latency cost, plus a one-line note on whether you kept or deleted your JSON-repair layer. Ship it as a LinkedIn post β "I deleted my JSON-repair code: before/after on 50 runs" β concrete numbers travel further than opinions.
Okta Identity Governance: Access Certifications That Actually Revoke (Not Rubber-Stamp)
π‘ Key Concept
Access sprawl is the silent liability of every mature org: users accumulate entitlements through role changes, project spikes, and one-off "just give me access" tickets β and almost nothing ever takes them away. The residue is standing privilege that auditors flag (SOX, SOC 2, ISO 27001) and attackers love. Okta Identity Governance (OIG) attacks this with Access Certification campaigns: time-boxed or recurring reviews where a designated reviewer must Approve or Revoke each userβresource grant with a business justification β and, critically, a revoke triggers automated remediation that actually removes the entitlement, not a checkbox in a spreadsheet.
The unit of governance is the campaign: you define scope (which users, which apps/entitlements), a reviewer model (manager, resource owner, or a named admin), a schedule (one-off or recurring quarterly), and remediation settings. In 2026 OIG made Resource Owner a first-class reviewer, added an AI-generated access summary in the Security access reviews view (context on how a user got the access and their history), and shipped Slack notifications for reviewers β every one of those features targets the real failure mode of certifications: reviewer fatigue producing a blind "approve all."
Position it in the identity graph: certifications are the periodic counterweight to provisioning. Lifecycle Management grants access on joiner/mover events; certifications catch what lifecycle missed and what humans over-granted. Together they close the loop from "who should have access" to "who actually does β and who signed off."
π¬ Deep Dive
The reviewer model decides whether the review is real or theater. "Manager" reviewers scale but rubber-stamp β a manager rarely knows what salesforce:reports.export means. "Resource owner" reviewers understand the entitlement but not the user's job. For crown-jewel resources use resource-owner review plus the AI access summary; for broad low-risk apps, manager review is fine. Never apply one reviewer model to the whole org β scope campaigns by risk tier.
Practitioner trap β group-based (indirect) access is invisible unless you certify the group too. A user often holds an entitlement because they're in an AD/Okta group assigned to the app, not through a direct app assignment. A campaign scoped only to "app assignments" shows the user as having access but a revoke does nothing β the group still confers it. You must add the group membership (and the groupβresource assignment) as separate review items, or the campaign reports a clean result while standing privilege quietly persists.
Practitioner trap β a certification without configured remediation is an audit artifact, not a control. "Revoke" removes access only if the campaign's remediation is enabled and the entitlement is one Okta actually manages (SCIM-provisioned or Okta-group-based). For a downstream app Okta doesn't provision, "revoke" records a decision while the access lives on in the target system β you've documented the problem, not fixed it. Verify the remediation path per resource before you launch.
Staff-level framing β the metric that matters is revocation rate and time-to-remediate, not completion rate. A campaign that closes 100% on time with a 99% approve rate is almost certainly rubber-stamping; a healthy one revokes a non-trivial slice and remediates within hours. As the architect, instrument outcomes (approve/revoke ratio, reviewer response time, mean revokeβremoval latency), tier cadence by sensitivity (quarterly for crown jewels, annual for low-risk), and feed revocations back to fix the provisioning rules that over-granted. That is governance as a system, not a fire drill before the audit.
π οΈ Artifact β Governance API: a recurring campaign that closes the indirect-access gap
# Create a RECURRING access-certification campaign via the Okta Governance API
# (resource-owner reviewer, quarterly, auto-remediation on revoke)
curl -X POST https://your-org.okta.com/governance/api/v2/campaigns \
-H "Authorization: Bearer $OIG_TOKEN" -H "Content-Type: application/json" \
-d '{
"name": "Q3-CrownJewels-Salesforce",
"description": "Quarterly cert of direct AND group-based Salesforce access",
"scheduleSettings": { "type": "RECURRING", "recurrence": "P3M", "durationInDays": 14 },
"reviewerSettings": { "type": "RESOURCE_OWNER", "fallbackReviewerId": "00u_secops_lead" },
"resourceSettings": {
"targetResources": [
{ "resourceType": "APPLICATION", "resourceId": "0oa_salesforce" },
{ "resourceType": "GROUP", "resourceId": "00g_sfdc_admins" }
]
},
"remediationSettings": { "accessApproved": "NO_ACTION", "accessRevoked": "DENY" }
}'
# accessRevoked=DENY -> Okta actually removes the grant (not just logs a decision).
# The trap made concrete: the second targetResource (the GROUP) is what closes the
# indirect-access gap. Omit it and a revoke on the app grant leaves group access intact.
# Don't trust campaign status alone β prove the removal happened in the System Log:
# eventType eq "application.user_membership.remove" and target.id eq "0oa_salesforce"
Cross-pollination: the "periodic re-attestation vs. standing state" discipline maps straight onto today's AI pill β an ungoverned fine-tuned adapter in production is the model-layer equivalent of an entitlement nobody recertifies.
π§ Recall
A week ago you looked at Okta FGA (ReBAC). Certifications answer "should this user still have access?"; FGA answers "can this user access this specific object right now?". Why can't a quarterly certification substitute for FGA's runtime Check?
Show answer
Certifications are point-in-time attestations of coarse entitlements (does the user hold the app/role), run on a cadence β they can't reason about per-object, relationship-derived permissions that change continuously ("can view doc #123 because they're on the owning team"). FGA's Check is evaluated at request time against the live relationship graph. Certifications govern the coarse grant; FGA enforces the fine-grained decision on every call. Complementary layers, not substitutes.
πΌ Market Signal
Per the KORE1 IAM Engineer Salary Guide (2026), senior IAM engineers run $150Kβ$200K base, and per Start with Identity's 2026 IAM Salary Guide, governance (IGA) and CIAM specialists sit at the top of the senior band β with Okta certifications (Professional through Consultant) cited as the most-requested vendor credential in workforce-identity postings this year. Demand context: in Okta's Q2 FY2026 results (reported Aug 26, 2025), total revenue was $728M (+13% YoY) with RPO of $4.15B (+18% YoY), growth management attributed to new-product adoption β the bucket that includes Okta Identity Governance. Certification-campaign design is the IGA competency that separates a governance architect from an SSO admin.
β‘ Action This Week
In an Okta org, create a group, assign it to an app, then give user A access via that group (indirect) and user B a direct app assignment. Document the difference in access source, and write the campaign scope that certifies both the app and the group. Definition of done: a one-page "campaign design" note plus screenshots showing (1) a user whose access is group-derived and (2) a target-resource list (app + group) that closes the indirect-access gap. If you have an OIG-enabled org, launch it and capture a revoke β application.user_membership.remove System Log event instead. Turn it into a LinkedIn post β "The access-review gap 9 of 10 orgs miss: group-derived access" β with the screenshots; it reads as governance depth, not admin trivia.
QLoRA + DPO: Fine-Tuning on One GPU Without RLHF's Headaches
π‘ Key Concept
Full fine-tuning updates billions of parameters β you must hold weights, gradients, and optimizer states (roughly 16 bytes/param with Adam), which pushes an 8B model well past a single accessible GPU. Parameter-efficient fine-tuning (PEFT) sidesteps this: LoRA freezes the base weights and injects small trainable low-rank matrices (AΒ·B, rank r) into the attention/MLP projections, so you train ~0.1β1% of parameters. QLoRA goes further β it quantizes the frozen base to 4-bit (NF4) and trains the LoRA adapters on top, cutting memory enough that an 8B model fine-tunes on a single 24β80GB GPU at near-full-fine-tune quality.
That covers teaching a task or style (supervised fine-tuning, SFT). But SFT can't cleanly teach preferences β "prefer the concise answer," "refuse this class of request." That used to mean RLHF (reward model + PPO): powerful but unstable and operationally heavy. Direct Preference Optimization (DPO) replaced it for most teams: given pairs of (chosen, rejected) responses, DPO optimizes the policy directly against a frozen reference model with a simple classification-style loss β no separate reward model, no RL loop. The 2026 default recipe is QLoRA SFT first, layer DPO if you have preference pairs, and reserve GRPO for when the reward is programmatically verifiable.
The architect's real question isn't "how do I fine-tune" β it's "should I." Fine-tuning changes behavior and format reliably but bakes in knowledge that goes stale; RAG changes knowledge without retraining; prompting changes nothing durably. Fine-tune for form and behavior, RAG for facts, prompt for iteration β and never fine-tune to inject knowledge you'll need to update next quarter.
π¬ Deep Dive
Rank and alpha are a ratio, not two independent knobs. LoRA's effective update scales as alpha / r. Raise r for capacity (a 2026 rule of thumb: 16 for style, 32 for general SFT, 64 for complex/coding) but keep alpha tracking it (start near alpha = 2r). Bump r alone and you silently shrink the update magnitude β a model whose loss dropped nicely but whose behavior barely moved.
Practitioner trap β QLoRA merge-to-fp16 reintroduces the error you quantized away. You train with a 4-bit (NF4) frozen base, but you often merge the adapter into fp16 to serve. The merge runs against the dequantized base, so the served model isn't bit-identical to what you trained and evaluated β subtle drift, worst on edge cases. Either evaluate the exact merged artifact you'll ship, or serve the adapter unmerged (hot-swapped) so training and serving see the same weights.
Practitioner trap β DPO's beta silently trades alignment for collapse.beta controls how far the policy may drift from the reference. Too low and preferences don't stick; too high and the model over-optimizes the signal, degrading fluency and diversity (mode collapse) while your reward-margin metric looks great. Watch generation diversity and a held-out quality eval, not just DPO loss β and make sure your (chosen, rejected) pairs differ on the axis you care about, or DPO learns a proxy (length, formatting) instead.
Staff-level framing β one base + many adapters is a serving architecture, not just a training trick. Because adapters are tiny (MBs), multi-adapter serving (vLLM / S-LoRA-style) keeps one quantized base resident and hot-swaps per-tenant/per-task adapters at request time β far cheaper than one full model per use case. That reframes fine-tuning as fleet economics: an adapter registry, versioning, an eval gate before prod, and rollback. And since data quality dominates (1,000 curated pairs routinely beat 100,000 noisy ones), the durable investment is the labeling/eval pipeline, not GPU hours.
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig
from trl import DPOTrainer, DPOConfig
import torch
# 1) 4-bit NF4 base (QLoRA): base frozen + quantized, adapters trainable
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True)
base = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct", quantization_config=bnb, device_map="auto")
# 2) LoRA: keep alpha tracking rank (effective update is proportional to alpha/r)
lora = LoraConfig(r=32, lora_alpha=64, lora_dropout=0.05, bias="none",
target_modules=["q_proj","k_proj","v_proj","o_proj",
"gate_proj","up_proj","down_proj"],
task_type="CAUSAL_LM")
# 3) DPO: align to (chosen, rejected) pairs β no reward model, no PPO loop.
# beta too high -> mode collapse; watch a held-out eval, not just the margin.
cfg = DPOConfig(beta=0.1, learning_rate=5e-6, per_device_train_batch_size=2,
gradient_accumulation_steps=8, bf16=True, max_length=1024)
trainer = DPOTrainer(model=base, ref_model=None, # ref = a frozen copy of base
args=cfg, peft_config=lora, train_dataset=pref_pairs)
trainer.train()
# Serve the adapter UNMERGED (hot-swap) OR eval the exact fp16-merged artifact you deploy.
Cross-pollination: an adapter in prod with no eval gate or version owner is the model-layer version of today's IAM pill's un-recertified entitlement β ungoverned standing state that drifts until it fails an audit, or a user.
π§ Recall
A week ago you looked at GraphRAG. Given today's fine-tune-vs-RAG rule, if a support bot must answer from a knowledge base that changes weekly, which approach do you reach for β and why not QLoRA?
Show answer
Reach for RAG (GraphRAG if the questions need multi-hop/global synthesis over the corpus). Fine-tuning bakes knowledge into weights at training time, so weekly changes would force a re-train and re-deploy of the adapter every week β slow and costly, and the model still can't cite sources or reflect an edit made an hour ago. RAG updates by re-indexing the store; the model stays fixed. Fine-tune for how the bot answers (tone, format, refusals); RAG for what facts it answers with.
πΌ Market Signal
Per the KORE1 AI Engineer Salary Guide (2026), AI-engineer base pay runs $145Kβ$310K, and per the same 2026 analysis, engineers who can do LLM fine-tuning + RAG architecture command a $20Kβ$50K+ premium over generalists, with dedicated LLM specialists at $220Kβ$280K and demand up ~136% year-over-year. Levels.fyi's 2026 data puts big-tech MLE median total comp at ~$264K. Fine-tuning sits squarely in that premium band β and the highest-paid combination cited is LLM serving infrastructure (vLLM / TensorRT-LLM / Triton) layered on top, which is exactly the multi-adapter serving skill in this pill's Staff framing.
β‘ Action This Week
Take a 500β1,000-example instruction dataset (curate your own for signal) and QLoRA-fine-tune an 8B model on a single GPU (Colab/A100 or a local 24GB card), then run the same prompt set against base vs. adapter. Definition of done: a side-by-side eval table (β₯20 prompts) with a win/lose/tie judgment per row, plus the peak-VRAM number proving it fit on one GPU. Bonus: log generation diversity so you can catch over-fitting. Post the win-rate table + VRAM figure on LinkedIn ("fine-tuned Llama-3 8B on one GPU β here's the before/after and the memory receipts"); a measured before/after beats another "I trained a model" post.
Okta for AI Agents: Giving Your Non-Human Workforce a Governed Identity
π‘ Key Concept
The moment an AI agent calls an API on its own behalf it becomes an identity your IdP has never seen β no owner, no lifecycle, no policy, no audit trail. Teams paper over this with a long-lived API key baked into the agent's environment: a "god token" that never expires, belongs to no one, and β if the agent is prompt-injected β hands an attacker everything that token can reach. Okta for AI Agents (Early Access now; GA April 30, 2026) makes the agent a first-class principal in the same control plane that governs humans: it gets registered, owned, scoped, monitored, and de-provisioned like any other identity.
Three primitives carry the model. A registry: every agent is an identity with an accountable human/team owner, a stated purpose, and a TTL. Credential vaulting: Okta holds the durable upstream secret and mints the agent only short-lived, audience-scoped tokens, so a compromised agent has no reusable secret to exfiltrate. Connection governance: every downstream hop β MCP server, API, or agent-to-agent β is evaluated against policy with dynamic, context-and-risk-aware decisions, and identity follows the agent across apps via Cross App Access (XAA), now an official MCP authorization extension. The through-line: an agent's authority becomes governed data you audit, not a secret you hope nobody leaks.
π¬ Deep Dive
Registration is the whole game β an unregistered agent is shadow IT with an API key. Model each agent as an identity with an accountable human owner, a purpose, and an expiry. The owner is your audit anchor: "which human is responsible for what agent:invoice-bot did at 03:14?" must have an answer before the agent ships, not after the incident.
Practitioner trap β vaulting does not mean the agent never holds a token. The agent still receives short-lived, audience-scoped access tokens to call tools; what is vaulted is the long-lived upstream credential (client secret, refresh token, downstream API key). Conflate the two and inject the raw client secret into the agent's env, and you have re-created the god-token you were trying to kill. Vault the durable secret; mint ephemeral, resource-indicated tokens per call.
Practitioner trap β the confused deputy across MCP. An agent acting on behalf of user A must not be able to reach a tool as if it were the whole system. Without the user's context propagated into the token (subject + audience), the MCP server can't distinguish "agent doing A's bidding" from "agent with god rights." XAA / token-exchange carries the user identity through the hop; a shared service token silently erases it β this is how over-broad agents turn a low-priv request into privileged action.
Staff-level framing β govern the connection graph, not just the agent. The blast radius of an agentic enterprise is the union of every tool every agent can reach. Treat it as a graph you can query and attest: enumerate agent β MCP-server β API edges, gate "no new high-privilege connection without review," and stream agent auth events to your SIEM so an agent suddenly reaching a new resource fires a detection. That org-wide control is what an architect owns; app-by-app key rotation is not.
# Agent authenticates with an ASYMMETRIC key (private_key_jwt) β no shared secret to leak.
# The durable key is VAULTED by Okta; the agent only ever holds a short-lived token.
curl -X POST https://your-org.okta.com/oauth2/aus1a.../v1/token \
-d grant_type=client_credentials \
-d client_assertion_type=urn:ietf:params:oauth:client-assertion-type:jwt-bearer \
-d client_assertion=$SIGNED_JWT \
-d scope='tools.read tickets.write' \
-d resource='https://mcp.internal/servers/helpdesk' # RFC 8707 β token bound to ONE resource
# -> access_token: aud=helpdesk-mcp, exp=300s, scope='tools.read tickets.write'
# A prompt-injected agent CANNOT replay this token against billing-mcp: wrong audience -> 401.
# ---- access policy for the agent (least privilege + owner + TTL, as governed data) ----
agent: helpdesk-bot
owner: team-support # human-accountable β the audit anchor
expires: 2026-11-18 # agents get a TTL, like a contractor badge
allowed_connections:
- mcp: helpdesk scopes: [tools.read, tickets.write]
- api: kb-search scopes: [read]
deny_by_default: true # any tool not listed is unreachable
step_up:
- action: tickets.delete require: human_approval # CIBA-style async sign-off
Cross-pollination: as agent systems fan out into supervisor/worker topologies (see today's AI pill), each worker should carry its own vaulted, least-privilege token β the difference between "one worker compromised" and "the whole system compromised."
π§ Recall
One week ago you looked at CIBA. When an agent must get a human's sign-off before a high-risk action (like the tickets.delete step-up above), which property of CIBA lets it trigger that approval without ever holding the user's interactive session?
Show answer
CIBA is a decoupled / backchannel flow: the agent is the initiator and pushes the approval request to the user's separate authenticator device out-of-band, then polls (or is pinged) for the result. The agent never sees the user's credentials or session β it only receives a token after the human approves on their own device, so the sensitive action stays gated behind a real person.
πΌ Market Signal
Per Okta's Q2 FY2026 results (reported Aug 26, 2025), total revenue was $728M (+13% YoY) with RPO of $4.15B (+18% YoY) and a record 28% non-GAAP operating margin; CEO Todd McKinnon explicitly framed managing the identities of AI agents as a new customer-expansion driver. That thesis rests on demand: per Gartner (press release, Aug 26, 2025), 40% of enterprise apps will embed task-specific AI agents by end of 2026, up from <5% in 2025. Every one of those agents needs an identity, a credential home, and a policy β the exact skill set that separates "SSO admin" from "agentic-identity architect."
β‘ Action This Week
In a free Okta developer org, register a service app, switch it from a client secret to private_key_jwt (asymmetric auth), and mint a token with a resource indicator scoped to one intended audience. Then replay that token against a different audience and capture the rejection. Definition of done: two terminal captures β (1) the decoded access-token JWT showing aud = your intended resource and a β€5-minute exp, and (2) a 401 proving audience binding blocks cross-resource replay. Wrap it in a LinkedIn post β "Why your AI agent should never hold a client secret" β with both screenshots; it signals hands-on NHI depth, not just theory.
The Supervisor Pattern: Multi-Agent Orchestration Without Losing Control
π‘ Key Concept
A single agent wired to 30 tools degrades in predictable ways: the model burns context reasoning about tools irrelevant to the current step, tool-selection accuracy drops as the menu grows, and one runaway loop poisons the entire trajectory. The multi-agent answer is decomposition. A supervisor (orchestrator) receives the request, breaks it into subtasks, routes each to a specialized worker agent with a small, focused toolset, and reduces their results into one answer. Workers operate on non-overlapping subtasks and don't see each other's outputs during execution β which keeps each worker's context small and its behavior legible in a trace.
Two topologies dominate. Supervisor (orchestrator-worker): one routing node owns global state and every handoff β clear control flow, every decision visible, easy to reason about. Swarm: agents hand off directly to each other with no intermediary β fewer LLM calls and faster, but control flow is emergent and harder to audit. In 2026 the supervisor pattern has the most native framework support: LangGraph v1.0's create_agent, the OpenAI Agents SDK's handoffs, and the Claude Agent SDK's subagents (one level deep, no nesting) are all supervisor primitives. The hard part isn't wiring the graph β it's what crosses a handoff (context) and with whose authority (identity).
π¬ Deep Dive
Handoffs are control transfer, not function calls. In LangGraph a worker returns a Command(goto=..., graph=Command.PARENT) to jump to another node in the parent graph; in the OpenAI Agents SDK a handoff is a tool the model can call. Either way conversation ownership moves β name your handoff tools explicitly (who, when-to-use, what payload) or the supervisor routes to the wrong specialist.
Practitioner trap β forwarding full history to every worker is the #1 cost/latency killer. Naive implementations pass the entire running transcript into each sub-agent, so a 6-worker run re-processes the same 40K tokens six times and worker prompts drift as they inherit irrelevant context. Pass a distilled task brief (the subtask plus only the facts it needs); instrument tokens-per-subtask, not just end-to-end latency.
Practitioner trap β no termination guard means infinite ping-pong. Two workers that can hand off to each other (or back to the supervisor) will loop forever if no node is authoritative about "done." Make the supervisor the only node that can emit the final answer, cap total hops, and give the orchestrator a step budget it decrements each turn.
Staff-level framing β topology is an org & security decision, not just latency. Supervisor gives you a single auditable chokepoint (every action attributable, every handoff traced) β exactly what you need when workers touch production systems; swarm trades that auditability for speed. Attach the per-worker scoped identity at that chokepoint so "which agent did what, with whose authority" is answerable. That governance story β not the framework β is what makes a multi-agent system shippable in a regulated enterprise.
π οΈ Artifact β a supervisor graph that distills briefs and scopes worker auth
from langgraph.graph import StateGraph, END
from langgraph.types import Command
# handoff = control transfer + a DISTILLED brief (not the whole transcript)
def to_worker(worker, brief):
return Command(goto=worker, update={"task_brief": brief})
def supervisor(state):
if state["phase"] == "research":
return to_worker("researcher", brief=state["question"]) # pass ONLY what it needs
if state["phase"] == "answer":
return Command(goto=END, update={"final": synthesize(state["findings"])})
def researcher(state):
# small, focused toolset + its OWN scoped token (never the supervisor's god-token)
hits = web_tool.run(state["task_brief"], auth=state["worker_token"])
# hand control back with a SUMMARY, not the raw pages
return Command(goto="supervisor",
update={"findings": summarize(hits), "phase": "answer"})
g = StateGraph(dict)
g.add_node("supervisor", supervisor)
g.add_node("researcher", researcher)
g.set_entry_point("supervisor")
graph = g.compile() # every hop is one traced node β your audit + cost chokepoint
Cross-pollination: mint each worker its own least-privilege token via Okta for AI Agents (see today's IAM pill) instead of sharing one key β so a prompt-injected worker can act only within its slice, never with the whole system's rights.
π§ Recall
One week ago you looked at LangGraph's interrupt() for human-in-the-loop. In a supervisor graph where a human might take hours to approve, what makes interrupt() safe β i.e., why doesn't the whole multi-agent run restart from scratch on resume?
Show answer
The checkpointer (durable state persistence). interrupt() pauses execution at a specific node and the graph's full state is checkpointed; on resume, execution continues from that exact node with prior state intact β it does not replay from the entry point. Durable execution is what lets a multi-agent graph survive long human waits (or process crashes) without re-running completed workers or re-billing their tokens.
πΌ Market Signal
Per Gartner (press release, Aug 26, 2025), 40% of enterprise apps will embed task-specific AI agents by end of 2026, up from <5% in 2025 β and 2026 adoption analyses aggregating Forrester/Gartner data report the agentic-AI market surpassing $9B with ~22% of production deployments now coordinating 3+ agents (i.e., multi-agent is going mainstream). Comp follows the skill: per Levels.fyi (2026), US AI-engineer average total compensation is ~$242,500, with senior LLM/agent specialists reaching $340Kβ$550K TC. The multiplier is orchestration + governance β Gartner also warns >40% of agentic initiatives may be scrapped by 2027 for failing precisely the control/ROI fundamentals this pill is about.
β‘ Action This Week
Build a 2-worker supervisor graph (supervisor + researcher + writer) in LangGraph or the OpenAI Agents SDK, and instrument token usage per worker. Run it twice: once forwarding the full transcript to each worker, once passing a distilled brief. Definition of done: a run log proving each worker received a distilled brief (not the full history) plus a two-row table of total tokens β naive baseline vs. brief-passing β showing the reduction. Post the before/after token table on LinkedIn ("how context isolation cut my multi-agent token bill by X%"); it turns an architecture principle into a measured, shareable result.
Okta FGA: Relationship-Based Authorization When RBAC Runs Out of Road
π‘ Key Concept
RBAC answers "what role does this user have?" β but real products ask "can this user view this specific document, because they're a member of the team that owns the folder it lives in?". That relational question is where role explosion begins: every new sharing rule spawns another role, and authorization logic metastasizes across every service. Okta FGA (the productized, hosted form of Auth0 FGA, built on the CNCF project OpenFGA, itself a re-implementation of Google's Zanzibar paper) answers it with Relationship-Based Access Control (ReBAC): you externalize authorization into a central graph of relationship tuples and query it with a millisecond Check call.
Two artifacts define the system. An authorization model (a typed DSL) declares object types and the relations between them β including computed relations like "a viewer is anyone who is an owner, or a viewer of the parent folder". Tuples are the data: user:anne is member of group:eng. Your app stops encoding permissions in if statements and instead asks FGA Check(anne, viewer, doc:roadmap) β allowed:true. Okta pitches ~100B relationships and 1M+ checks/sec at low latency β the point is that authz becomes data you audit and evolve, not code you redeploy.
π¬ Deep Dive
The model is the contract. Computed relations (or, and, but not, and X from Y) let one line express inheritance that RBAC needs dozens of roles to fake. Model it wrong once and every downstream check inherits the flaw β treat the DSL like a schema migration, versioned and reviewed.
Practitioner trap β eventual consistency bites read-after-write. FGA defaults to MINIMIZE_LATENCY consistency: write a tuple, immediately Check, and you can get a stale allowed:false because the change hasn't propagated. In share-then-open UX this reads as "I granted access but they still can't see it." Fix: pass consistency: HIGHER_CONSISTENCY on the critical check, or supply the just-written relationship as a contextual tuple so the check evaluates it without waiting for the write to settle.
Practitioner trap β ListObjects is not free. "Show every doc anne can see" fans out across the graph and is far heavier than a point Check. Teams that back a listing page with naive ListObjects per-render hit latency cliffs at scale β cache results, or invert the query.
Staff-level framing β centralization is leverage and blast radius at once. Moving authz to FGA makes it auditable, testable, and consistent across services β but the model and its uptime become a Tier-0 dependency: a bad model deploy or an FGA outage denies everything. Budget for a fail-safe cache, model-change canaries with the FGA test framework (.fga.yaml assertions in CI), and an explicit policy-as-data vs. policy-as-code (OPA/Rego) decision β FGA owns "who relates to what", OPA still fits stateless request-context rules.
π οΈ Artifact β model + tuples + check
# authorization model (OpenFGA DSL, schema 1.1)
model
schema 1.1
type user
type group
relations
define member: [user]
type folder
relations
define viewer: [user, group#member]
type document
relations
define parent: [folder]
define owner: [user]
# computed: owner OR direct viewer OR inherited from parent folder
define viewer: [user, group#member] or owner or viewer from parent
# seed relationships (fga CLI)
fga tuple write user:anne member group:eng
fga tuple write group:eng#member viewer folder:plans
fga tuple write folder:plans parent document:roadmap
# the question your app asks on every request
fga query check user:anne viewer document:roadmap
# -> { "allowed": true } (anne never had a role on the doc)
# read-after-write on a share flow β force consistency
fga query check user:anne viewer document:roadmap --consistency HIGHER_CONSISTENCY
Cross-pollination: this same relationship graph is exactly what an identity-aware GraphRAG index needs to gate which knowledge-graph nodes an AI agent may traverse per user β see today's AI pill.
π§ Recall
One week ago you looked at Okta Realms. In a realm-based delegated-admin design, what is the single property that most limits a compromised partner-admin's blast radius?
Show answer
Realm isolation of the administrative scope: a realm admin can only see and act on users/groups assigned to their realm, so a compromised partner admin cannot enumerate or modify identities outside it β the realm boundary caps lateral reach regardless of the admin role's permissions.
πΌ Market Signal
Per Glassdoor salary data (accessed Aug 2026), the average total pay for an "Okta IAM Engineer" in the US is ~$170,900/yr, with the range spanning ~$132K (25th pctl) to ~$223K (75th pctl). ZipRecruiter (Aug 2026) lists active "Okta IAM" roles at $95Kβ$170K. The differentiator is depth: authorization-modeling skill (FGA/ReBAC, OPA) sits above generic SSO admin work β it is the part that scales into Staff/Architect scope where you own the org-wide authz contract, not just app assignments.
β‘ Action This Week
Run OpenFGA locally: docker run -p 8080:8080 openfga/openfga run, then use the fga CLI to load the doc-sharing model above and prove inheritance works. Definition of done: a terminal screenshot showing fga query check user:anne viewer document:roadmap returning {"allowed": true} where anne's access comes only via group:eng β folder:plans β parent β no direct tuple on the document. That screenshot plus a 4-line "RBAC role explosion vs. ReBAC" caption is a crisp LinkedIn post that signals authorization-architect depth.
GraphRAG: When Vector RAG Can't Answer "What Are the Big Themes?"
π‘ Key Concept
Vector RAG retrieves the top-k chunks most similar to a query β which is perfect for "what's our PTO policy?" and useless for "what are the recurring themes across all 400 incident postmortems?". The second question is a global sensemaking query: no single chunk contains the answer, and semantic similarity to the question retrieves nothing, because the answer is a property of the whole corpus. GraphRAG (Microsoft Research, arXiv 2404.16130) restructures the corpus into a knowledge graph before any query arrives, so the structure itself carries meaning that flat chunks throw away.
Indexing is an LLM pipeline: extract entities and relationships from every chunk into a graph, then run Leiden community detection to cluster densely-connected entities into a hierarchy, and pre-generate an LLM summary of each community. Query time offers two modes. Local search anchors on specific entities and walks their neighborhood β best for "tell me about X". Global search does map-reduce over the community summaries β each summary partially answers the question, then a reduce step synthesizes β which is how it answers themes-across-everything questions vector RAG structurally cannot.
π¬ Deep Dive
Local β global β pick per query, not per system. Route point-lookups to local (or plain vector) search and only sensemaking questions to global. A production router that classifies query intent first is what keeps GraphRAG affordable; running everything through global is the naive mistake.
Practitioner trap β indexing cost is a step function, not a slope. Extraction fires an LLM call per chunk; a modest corpus is thousands of calls, and re-indexing on new docs repeats much of it. Teams pilot on 20 documents, love the demo, then hit a five-figure indexing bill at real scale. Microsoft's LazyGraphRAG (announced 2024, maturing through 2026) defers the expensive summarization to query time and reportedly cut indexing cost to ~0.1% of vanilla GraphRAG for comparable quality β evaluate it before committing to full pre-summarization.
Practitioner trap β global search is token-hungry per query. Map-reduce touches many community summaries every time, so a single global question can cost 50β100Γ a vector lookup. Constrain the community hierarchy level you query, and cache reduce outputs for repeated question shapes.
Staff-level framing β GraphRAG is a build-vs-buy and cost-governance decision, not a default. Its own paper reports strong win rates on comprehensiveness and diversity for global questions β but that value only lands on corpora people ask holistic questions about (research, incident/postmortem bodies, competitive intel). Own the decision explicitly: budget indexing as a recurring cost line, set an incremental-reindex SLA, and treat "is this corpus even asked global questions?" as the gate before any pipeline work.
π οΈ Artifact β index once, query two ways
pip install graphrag
# scaffold + drop .txt files into ./ragtest/input, set your model key in settings.yaml
graphrag init --root ./ragtest
graphrag index --root ./ragtest # LLM extraction + Leiden + community summaries (the $$ step)
# GLOBAL: corpus-wide sensemaking β no single chunk holds this answer
graphrag query --root ./ragtest --method global \
--query "What are the recurring root-cause themes across all postmortems?"
# LOCAL: entity-anchored neighborhood walk
graphrag query --root ./ragtest --method local \
--query "What did the payments outage on 2026-07-02 depend on?"
Cross-pollination: an identity-aware GraphRAG must prune the graph to nodes the caller may see β exactly a per-user Check(user, viewer, node) against the relationship graph in today's Okta FGA pill. Filter during traversal, not after retrieval, or you leak entity names in summaries.
π§ Recall
Ten days ago: in a two-stage retriever, why does adding a cross-encoder reranker beat simply retrieving more chunks with the bi-encoder?
Show answer
A bi-encoder embeds query and doc independently, so it can't model token-level interaction; a cross-encoder jointly encodes the (query, doc) pair and scores true relevance. You retrieve broadly and cheaply with the bi-encoder, then rerank a small candidate set with the accurate-but-expensive cross-encoder β better precision at the top without paying cross-encoder cost over the whole corpus.
πΌ Market Signal
Per Levels.fyi data (through Q1 2026), AI engineers report ~$244K+ total comp overall, and the Recruiting from Scratch analysis of 1.9M postings (2026) puts the remote AI-engineer median at ~$194K ($155Kβ$225K interquartile). LLM/GenAI specialists cited at $175Kβ$260K base, outpacing generalist ML by $30Kβ$60K+. The recurring phrase across guides β engineers who can "build production-grade RAG pipelines and evaluate outputs at scale" are supply-constrained β is precisely the GraphRAG-vs-vector routing and cost-governance judgment above, not just wiring an embedding call.
β‘ Action This Week
Take ~20 documents you actually know (past postmortems, a set of meeting notes), run graphrag index, then ask the same holistic question via --method global and via a plain vector-RAG baseline. Definition of done: a side-by-side of the two answers plus the indexing token/cost number from the run logs. The comparison β "global surfaced 3 cross-document themes the vector baseline missed, at $X indexing cost" β is a portfolio artifact that demonstrates the judgment (when GraphRAG earns its cost), which is what Staff-level AI roles actually screen for.
When the Agent Must Ask Permission: Auth0 CIBA for Async Human-in-the-Loop Approval β and Why Your Approval Binds the Human, Not the Transaction
π‘ Key Concept
An autonomous agent authenticated as its own principal is fine for reading a calendar β but the moment it wants to wire funds, delete a production database, or send an email to your CFO, you want a human in the loop. The problem: the human isn't sitting in the chat. CIBA (Client-Initiated Backchannel Authentication, OpenID spec) solves exactly this β it decouples the device that initiates authorization (the agent, the "Consumption Device") from the device that approves it (the user's phone, the "Authentication Device"). Auth0's Async Authorization for AI Agents, which reached GA in the May 2026 Auth0 Platform release alongside Auth for MCP and Agent-as-Principal, packages CIBA as the human-approval primitive for agentic workflows.
The flow is poll-based and backchannel: the agent calls /bc-authorize with the target user, a binding_message, and (critically) the scopes it wants. Okta pushes a rich approval to the user's enrolled authenticator; the agent receives an auth_req_id and polls the token endpoint with the CIBA grant until the human approves, denies, or the request expires. Only then does the agent get an access token β narrowly scoped to the approved action β to actually call the downstream API. This is the identity-layer twin of a framework pausing an agent graph for human oversight: the same async pause boundary, enforced by tokens instead of application code.
π¬ Deep Dive
The backchannel request and poll are two distinct grants. You initiate with /bc-authorize and redeem with the CIBA grant type on the token endpoint. Respect the returned interval β poll faster and you get slow_down, which increments your backoff:
Practitioner trap β the approval binds a message string, not the transaction.binding_message is display text shown to the user; nothing cryptographically ties the token you receive to the exact $4,200 / Acme Corp the human saw. An agent (or a compromised prompt) can display one amount, get approval, then reuse that same access token for a different transfer within the token's lifetime. Fix: carry the transaction as RAR authorization_details so the approved parameters are embedded in the token and enforced by the resource server β and keep CIBA token TTLs to seconds, one action per token. The binding_message is for the human's eyes; authorization_details is for the machine's enforcement.
Staff-level framing β design the timeout as a policy, not an accident. CIBA requests expire (expires_in), and a human who never responds must resolve to deny by default, not to a silently retried or dropped action β otherwise your "human oversight" is fail-open under time pressure. At org scale, wire every bc-authorize, approval, denial, and expiry into your System Log stream so each async agent action has an immutable, human-attributed audit trail; then set an approval-SLA dashboard so a flood of pending requests (a runaway agent) is visible, not buried. The governance question a Staff engineer owns isn't "can the agent ask?" β it's "who approved what, how fast, and what happens when nobody does."
π§ Recall
From ~a week ago: a CIBA access token is a bearer token β anyone who steals it in transit can replay it. What Okta mechanism (covered Aug 3) would cryptographically bind that token to the agent's own key so a stolen copy is useless?
Show answer
DPoP (sender-constrained tokens). The agent proves possession of a private key via a DPoP proof JWT bound to the token; the resource server rejects the token if presented without a matching proof, defeating bearer-token replay. Pair CIBA (who approved) with DPoP (who may present the token) for agent actions.
Spin up a free Auth0 tenant, enable Async Authorization (CIBA) with Guardian push, and script the two-call flow above against a dummy "transfer:funds" scope. Definition of done = a terminal recording (or screenshots) showing (1) the bc-authorize response with auth_req_id, (2) the push landing on your phone with the binding_message, and (3) the token endpoint flipping from authorization_pending to a scoped access_token after you approve. Turn the recording into a 60-second LinkedIn demo captioned "human-in-the-loop for AI agents, enforced at the identity layer" β a visible artifact that signals agentic-identity fluency to hiring managers.
Pausing an Agent Without Losing Its Mind: LangGraph interrupt() + Durable Checkpointers β and Why InMemorySaver Silently Eats Your Approvals
π‘ Key Concept
A production agent that wants to send an email or execute a trade shouldn't just do it β it should pause, surface the proposed action, and wait for a human. In LangGraph the primitive for this is interrupt(): called inside a node, it halts graph execution, persists the entire graph state to a checkpointer, and returns control to the caller. Later you resume by invoking the graph with Command(resume=value) β the human's decision flows back into the exact spot execution stopped. Unlike static breakpoints, interrupts are dynamic: they can fire anywhere and conditionally, based on the agent's own logic (e.g., only pause when a tool call's dollar amount exceeds a threshold).
The load-bearing word is durable. The pause is only as reliable as the checkpointer behind it. LangGraph writes a state snapshot at every superstep, organized by thread_id, into two tables (checkpoints and writes). A human might approve in three seconds or three days; across that gap your process can restart, your container can be rescheduled, your pod can OOM. If state lives only in memory, the pause β and the pending approval β vanishes. This is the application-layer counterpart to an identity-layer async-authorization flow: the graph interrupt is precisely where you'd trigger a CIBA push and block on the human's answer.
π¬ Deep Dive
The pattern is: interrupt to pause, Command to resume β and it needs a thread_id. The interrupt() payload is what your UI shows the reviewer; the value passed to Command(resume=...) becomes the return value of that same interrupt() call on the next run:
from langgraph.types import interrupt, Command
from langgraph.checkpoint.postgres import PostgresSaver # durable, NOT InMemorySaver
def approval_node(state):
decision = interrupt({ # pauses here; state persisted
"action": "send_email",
"to": state["recipient"], "body": state["draft"],
})
if decision != "approve":
return {"status": "rejected"}
return {"status": "sent"} # only runs after resume
graph = builder.compile(checkpointer=PostgresSaver.from_conn_string(DB_URI))
cfg = {"configurable": {"thread_id": "run-42"}} # required to locate the paused state
graph.invoke({"recipient": "cfo@corp.com", "draft": "..."}, cfg) # β stops at interrupt
# ...minutes or days later, from a fresh process...
graph.invoke(Command(resume="approve"), cfg) # resumes exactly where it paused
Practitioner trap β resume re-runs the whole node from the top, so pre-interrupt side effects fire twice.interrupt() doesn't freeze a Python stack frame; on resume LangGraph replays the node from its beginning and fast-forwards previously completed work, but any non-checkpointed side effect (an API POST, a row insert, an LLM call) placed before the interrupt() executes again. Rules that save you: put interrupt() as early as possible in the node, make anything before it idempotent, and never place two interrupts in one node without accounting for replay ordering. Teams discover this when their agent sends the email twice β once before approval, once after.
Staff-level framing β durability is an infra decision, not a demo detail.InMemorySaver is perfect for notebooks and evals and catastrophic in production: it is not restart-durable, so a redeploy or crash during a pending human review silently drops the run. For real workloads use a Postgres/Redis checkpointer, treat thread_id as a first-class business key (so approvals map to auditable runs), and budget for replay cost β every resume re-executes node logic, so a graph with expensive pre-interrupt LLM calls pays that token cost twice. The architecture question you own: what's your SLA on a paused run surviving a deploy, and does your checkpoint store meet it?
π§ Recall
From ~5 days ago: if your approval node re-runs and pays for an expensive LLM call twice on resume, one inference-side technique (covered Aug 6) cuts that per-call latency by drafting several tokens cheaply and verifying them in one pass. Name it.
Show answer
Speculative decoding (e.g., EAGLE-3). A small draft model proposes k tokens; the target model verifies them in a single forward pass, accepting the longest correct prefix β same output distribution, lower wall-clock latency per generation, which softens the cost of replayed nodes.
πΌ Market Signal
Agentic engineering is a seller's market: US Agentic AI Engineers average ~$190K with top earners past $300K (Glassdoor, 2026), and job postings mentioning LangGraph or MCP pay 15β25% more than LangChain-only listings β LangGraph being the highest-paying framework-specific skill (KORE1 Agentic AI Hiring Survey, 2026). Agentic postings grew ~280% YoY to roughly 90,000 US listings in 2026 (jobsbyculture, 2026). Human-in-the-loop with durable checkpointing is exactly the "reliable agents in production" competency these listings screen for β not toy demos, but pauseable, resumable, auditable workflows. (Cross-pollination: today's IAM pill shows the identity half β the interrupt() boundary here is where a CIBA async-authorization push enforces the human approval at the token layer.)
β‘ Action This Week
Build a two-node LangGraph agent (draft β approval) with a PostgresSaver checkpointer, hit the interrupt(), then kill the Python process entirely and resume from a fresh process with Command(resume="approve") using the same thread_id. Definition of done = the run completes correctly after the restart, proving the pause survived process death β plus a screenshot of the checkpoints table row holding the paused state. Write it up as a short "durable human-in-the-loop in 40 lines" post; the process-kill demo is the detail that separates it from every InMemorySaver tutorial and makes it portfolio-worthy.
One Org, Many Blast Radii: Okta Realms for Subsidiary & Partner Isolation β and Why "Realm β Policy Boundary" Is the Migration That Bites
π‘ Key Concept
When an M&A closes or a partner ecosystem grows, the reflex is "spin up a second Okta org and org2org it back to the hub." That works, but it fragments licensing, splits your global threat-signal surface, and forces cross-org federation plumbing for every app. Okta Realms offers the other lever: partition a single org into isolated user populations β a subsidiary, an acquired company, an external partner β where a delegated admin for Realm A cannot see, edit, or reset users in Realm B, yet everyone still shares one org's apps, SSO domain, and licensing pool.
The mental model a Staff engineer must hold: a realm is an administrative and directory isolation boundary, not a security-policy boundary. Realms scope who an admin can manage (users, groups, and profile sources) and which population a user belongs to. They do not automatically scope authentication policies, sign-on rules, or app assignments β those still live at org level and span realms unless you deliberately scope them by group. Treat realms as delegated-admin blast-radius control first; layer group-scoped policy on top to get the tenant-like behavior people assume realms give them for free.
π¬ Deep Dive
Assignment is priority-ordered and source-driven. A user can match multiple realm assignments; Okta walks them by ascending priority and applies the first whose conditions all pass. Build the rule against a profile-source attribute (country, employee type, source app) β the API object shape:
# Create a realm (Realm Assignments API β gated by OIG / Secure Partner Access)
POST /api/v1/realms
{ "profile": { "name": "EU-Subsidiary" } }
# Route new HR/AD-sourced users into it, by priority
POST /api/v1/realm-assignments
{ "name": "route-eu-staff",
"priority": 10,
"status": "ACTIVE",
"profileSourceId": "0oa1hrsource",
"conditions": { "expression":
{ "value": "user.countryCode == \"DE\"",
"type": "urn:okta:expression:1.0" } },
"actions": { "assignUserToRealm": { "realm": { "id": "rlm1a2b3c" } } } }
Practitioner trap β realm assignments do NOT re-home existing users. Rules evaluate only when a user is created or imported from a profile source. Add a new realm rule after go-live and every current user stays in the Default Realm with a green, healthy-looking status β you discover the misalignment months later when a delegated subsidiary admin quietly can't see "their" people. Backfill deliberately: move existing users with the API (POST /api/v1/realms/{realmId}/users) or an Okta Workflows flow, and reconcile realm membership as a scheduled job, not a one-time import.
Staff-level framing β realms vs. a second org is a blast-radius/cost decision. Separate orgs give the hardest isolation (independent policies, tenants, rate-limit buckets) but fragment SSO, licensing, and your global ITP/threat-signal view, and add org2org federation to maintain. Realms keep one org's shared SSO, licensing, and threat surface while cutting delegated-admin blast radius β at the price of shared global sign-on policies and shared org rate limits across every realm. Decide per driver: hostile-divestiture or regulatory data-residency β separate org; subsidiary/partner delegation with common apps β realms.
π§ Recall
From ~9 days ago: a realm is delegated to a subsidiary admin by binding it to which two Okta primitives β and what stops that admin from touching users in a different realm?
Show answer
A custom admin role (the permission set) bound to a resource set (the scope) that lists only the target realm. The resource set is the fence: because it enumerates just Realm A's users/groups, the granted role can never resolve identities in Realm B β same least-privilege mechanic as the custom-admin-roles pill, now scoped to a realm instead of a group.
πΌ Market Signal
The governance surface realms live inside is where budget is flowing: the Identity Governance & Administration market is ~$9.57B in 2026, growing at a 13.62% CAGR to ~$18.12B by 2031 (per Mordor Intelligence, "IGA Market" report, 2026), with North America ~45% of spend. Okta was again named a Gartner Peer Insights Customers' Choice for IGA (Okta analyst-research page, 2026).
In a dev/preview org, create one realm, add a priority realm-assignment rule keyed on a profile attribute, import a test user, and confirm placement. Then create a custom admin role + resource set scoped to that realm and sign in as the delegated admin. Definition of done = a screenshot of the delegated admin's People view showing only the realm's users (and the Default Realm user absent), plus a one-line note proving a pre-existing user did not auto-move when you added the rule. That contrast β "isolation works, but only for new users" β is a strong LinkedIn post on realms as blast-radius control, not tenancy.
Speculative Decoding with EAGLE-3: Lossless 2β5Γ Latency β and the Batch-Size Trap Where It Quietly Destroys Your Throughput
π‘ Key Concept
Autoregressive decoding is memory-bandwidth bound: each token needs a full forward pass over billions of weights just to emit one token, and the GPU's compute sits mostly idle waiting on memory. Speculative decoding breaks that one-token-per-pass ceiling. A cheap draft proposes k tokens ahead; the expensive target model then verifies all k in a single forward pass. Accepted tokens are kept, the first rejected one is resampled β so multiple tokens can clear per target pass instead of one.
The property that makes this production-safe is losslessness: the accept/reject step (rejection sampling) is constructed so the output distribution is provably identical to the target model decoding alone β no quality regression, ever. EAGLE-3, the current default in vLLM, doesn't use a separate small LLM as the drafter; it trains a lightweight autoregressive head on the target's own hidden-state features, which pushes acceptance rates to ~70β85% on Llama-3-70B and delivers up to ~2.5Γ end-to-end on general workloads, 3.5β5Γ on code (AWS's P-EAGLE, merged to mainline vLLM in early 2026, reports 4β5Γ on coding benchmarks). Because it's lossless, you can A/B it in production with zero eval-quality risk β the only thing you're trading is GPU compute.
π¬ Deep Dive
Enabling it is one config block β reading the metric is the real skill. In vLLM you attach an EAGLE-3 drafter and set the lookahead depth; then you watch the acceptance length (mean accepted tokens per step), not just latency:
from vllm import LLM
llm = LLM(
model="meta-llama/Llama-3.1-70B-Instruct",
tensor_parallel_size=4,
speculative_config={
"method": "eagle3",
"model": "yuhuili/EAGLE3-LLaMA3.1-Instruct-70B",
"num_speculative_tokens": 5, # k: bigger k helps only if acceptance stays high
},
)
# Watch the server log: "Speculative metrics: ... acceptance rate / draft accept length"
# accept_len ~4/5 -> great; ~1.5/5 -> drafter is mismatched to your workload
Practitioner trap β it's a latency win, not a throughput win, and the crossover is batch size. At low concurrency the GPU is memory-bound and idle compute is free, so verifying speculative tokens costs nothing you weren't wasting anyway. At high batch sizes the server becomes compute-bound, and every drafted-then-rejected token now burns FLOPs you needed for real requests β enabling spec decode globally can reduce total tokens/sec under load. Never benchmark it at batch=1 and ship it; measure throughput across your real concurrency curve and gate it by QPS.
Staff-level framing β pick the regime, then set the policy. Speculative decoding is right for latency-SLA interactive traffic (chat, coding copilots, agent loops with tight per-step budgets) and wrong for offline/batch throughput jobs where cost-per-token dominates. Because it's output-lossless there's no accuracy blast radius, so the safe rollout is a per-endpoint flag: enable on the low-latency tier, keep it off on the batch tier, and add an autoscaler rule that disables it when sustained batch size crosses your measured crossover point. Also pin the drafter to the workload β a draft tuned on general text underperforms on code/JSON, and the drafter must share the target's tokenizer/vocab or acceptance collapses.
π§ Recall
From ~8 days ago: with Matryoshka (MRL) embeddings, why can you truncate a 1024-dim vector to 256 dims and still retrieve well β and how is that "cheap first pass, expensive verify" shape the same idea as speculative decoding?
Show answer
MRL training nests coarse-to-fine information so the leading dimensions carry most of the semantic signal β truncated vectors give a fast, approximate coarse filter, and you rerank survivors with the full dimensions. Same two-stage economics as spec decoding: a cheap proposer narrows the field, an expensive verifier confirms β you only pay full cost on what clears the first stage.
πΌ Market Signal
Inference-cost engineering is where the AI premium concentrates. The remote AI Engineer median sits at ~$194K in 2026 (25thβ75th percentile ~$155Kβ$225K), per Recruiting from Scratch's analysis of ~1.9M job postings (2026); the same market commentary notes ~$250K packages becoming standard for senior engineers who can demonstrably cut inference cost and run RAG at scale.
Speculative decoding is exactly that credential: it's a concrete, measurable latency/cost lever (2β5Γ with a config change) that separates "I call an LLM API" from "I operate an inference stack." A benchmark you ran β acceptance rate, latency, and the throughput crossover β is a portfolio artifact that reads as production judgment, and job posts increasingly name vLLM / inference optimization as a differentiator rather than a nice-to-have.
β‘ Action This Week
Stand up vLLM with a 7β8B target model and its EAGLE-3 drafter, and run 50 coding prompts twice β spec decode on and off β sweeping batch sizes 1, 8, and 32. Log mean acceptance length, P50/P99 latency, and total tokens/sec at each batch. Definition of done = a small table (or one line chart) showing latency improving at batch=1 and throughput crossing over β the batch where spec decode stops helping. Post it as a "speculative decoding is a latency win, not a throughput win β here's the crossover" LinkedIn chart; a real curve beats every blog restating the paper's headline number.
The Account That Never Died: Okta Lifecycle Management, SCIM Deprovisioning & Why "Deactivate in Okta" Is Not "Access Revoked"
π‘ Key Concept
Provisioning is the glamorous half of Lifecycle Management β a joiner lights up on day one, HR is happy, the demo works. Deprovisioning is where the breaches live. When a leaver is terminated, Okta drives an HR-as-master source through Universal Directory, group rules, and Workflows to flip that user off across every SCIM-connected app. The trap is believing that flip is atomic and complete. It rarely is: SCIM apps disagree on what "deprovision" even means, and a large fraction of your estate has no SCIM at all.
The core distinction a Staff engineer must internalize: SCIM DELETE /Users/{id} vs. PATCH active:false. The spec (RFC 7644) treats a DELETE as "remove the resource"; most well-behaved SaaS apps implement it as a soft delete (retain the record, set inactive) β but some hard-delete and lose the audit trail, and others ignore DELETE entirely and only honor a PATCH to active:false. Okta's "Deactivate" button abstracts this per-connector, so two apps behind the same termination event can end up in opposite states. And none of it revokes the app's existing OAuth refresh tokens or browser sessions β the account can be "inactive" in Okta while a long-lived session keeps working for hours.
π¬ Deep Dive
Kill the session, not just the account. Deprovisioning removes future authentication, but existing tokens survive. Pair deactivation with a session/token clear so a leaver's open tab dies immediately:
# 1) Deactivate in Okta (halts new sign-ins)
POST /api/v1/users/{userId}/lifecycle/deactivate
# 2) Revoke ALL active Okta sessions + OAuth grants (kills live access)
DELETE /api/v1/users/{userId}/sessions?oauthTokens=true
# 3) SCIM downstream: what the connector actually sends
PATCH /scim/v2/Users/{id}
{ "schemas":["urn:ietf:params:scim:api:messages:2.0:PatchOp"],
"Operations":[{ "op":"replace","path":"active","value":false }] }
Practitioner trap β the DELETE/PATCH disagreement. A downstream app that only implements SCIM DELETE will silently ignore Okta's deactivate-as-PATCH, leaving the account fully live with a green checkmark in the Okta admin console. Always test each connector with a canary user and read the real HTTP call in the System Log (application.provision.user.deactivate) β do not trust the UI status. The inverse also bites: apps that hard-delete on DELETE destroy the record you needed for the offboarding audit.
Staff-level framing β govern the un-SCIM'd tail. The residual risk is not the automated 80%; it's the SWA/bookmark apps and shadow SaaS with no SCIM connector. Inventory every app, and for each one without SCIM, name a deprovisioning runbook owner and document the manual steps so the process survives owner turnover. Extend the same discipline to non-human identities: a leaver who owned a service account or an agent's client-credentials app leaves a token that no HR event will ever deactivate β link ownership so departure triggers rotation (this is where yesterday's DPoP/NHI work pays off).
π§ Recall
From ~5 days ago: which Okta System Log event stream would you tap to detect a deprovisioning that silently failed downstream, and where would you route it?
Show answer
Okta Log Streaming (System Log β EventBridge/Splunk) β stream application.provision.user.deactivate and its failure variants into your SIEM, then build a detection that alerts when a user.lifecycle.deactivate is not followed by matching downstream deactivations within an SLA window.
πΌ Market Signal
Orphaned access is a measured, not theoretical, risk: identity-lifecycle guidance in 2026 cites that ~83% of former employees retain access to a previous employer's systems after departure (per SSOJet / identity-lifecycle best-practice roundups, 2026) β the exact failure deprovisioning governance exists to close.
The role that owns this pays for it: US Senior IAM Engineer base sits around $150Kβ$200K in 2026, with senior/lead identity engineers commonly in the $150Kβ$200K band (per PayScale & Salary.com, July 2026) β and joiner-mover-leaver automation + SCIM connector depth is exactly the "identity management skills" premium those surveys flag.
β‘ Action This Week
Pick one SCIM-connected app in a dev/preview Okta org, create a canary user, then deactivate it and capture the actual downstream call from the System Log. Definition of done = a screenshot showing the application.provision.user.deactivate event and the resolved HTTP verb (DELETE vs PATCH active:false) for that connector, plus a one-line note on whether existing sessions survived. This becomes a sharp LinkedIn post β "The difference between 'deactivated in Okta' and 'access actually revoked'" β with a concrete artifact behind it.
Teacher in the Loop: On-Policy Distillation, or How an 8B Student Learns to Reason at a Fraction of RL's Cost
π‘ Key Concept
Classic off-policy distillation (SFT on a teacher's transcripts) has a structural flaw: the student only ever sees the teacher's perfect trajectories, never its own mistakes. At inference the student walks paths the teacher never demonstrated, errors compound, and reasoning collapses β the exposure-bias problem. RL fixes the on-policy part (learn from your own rollouts) but its reward is sparse: one scalar per episode, brutally sample-inefficient.
On-policy distillation (OPD) takes the best of both. The student generates the rollout, then the teacher grades it densely β a per-token target distribution (typically reverse-KL against the teacher's logits) instead of a single end-of-episode reward. You get RL's relevance (learning on the distribution you actually deploy on) with distillation's dense signal (feedback on every token). In 2026 this went from research trick to production default: Qwen3, MiMo, and GLM-5 all use OPD in post-training, and Thinking Machines Lab reproduced Qwen3's result β training reasoning into Qwen3-8B-Base using Qwen3-32B as teacher β at a fraction of the compute a pure-RL run would cost (Thinking Machines Lab, "On-Policy Distillation," 2026).
π¬ Deep Dive
The loss is dense, per-token KL β not a scalar reward. Sample from the student, then minimize divergence from the teacher on the student's own tokens:
# student generates its OWN rollout (on-policy)
with torch.no_grad():
seq = student.generate(prompt, do_sample=True, max_new_tokens=T)
t_logits = teacher(seq).logits # dense supervision
s_logits = student(seq).logits # requires grad
# reverse KL: KL(student || teacher) over the student's trajectory
logp_s = F.log_softmax(s_logits, dim=-1)
p_t = F.softmax(t_logits, dim=-1)
loss = F.kl_div(logp_s, p_t, reduction="batchmean") # per-token, not per-episode
loss.backward()
Practitioner trap β shared tokenizer is a hard prerequisite. Per-token KL is only defined if student and teacher emit over the same vocabulary. Distilling across model families with different tokenizers (e.g., Llama teacher β Qwen student) silently misaligns positions and produces garbage gradients unless you add a token-alignment/logit-projection step. Pick a teacher in the student's own family whenever you can.
Staff-level framing β distillation is a build-vs-buy lever, and a repair tool. The org decision isn't "distill or not"; it's when a distilled small model beats paying frontier-API rates per call β high-QPS, latency-sensitive, or data-residency-bound workloads flip the math toward a self-served student. OPD also fixes catastrophic forgetting: after domain fine-tuning erases prior behavior, distill from the pre-fine-tune checkpoint as its own teacher to restore it while keeping the new knowledge. Treat the distillation pipeline as a reusable platform capability, not a one-off. (The teacher's outputs are training data β govern who can reach the teacher endpoint with the same identity-scoped access controls you'd put on any model API.)
π§ Recall
From ~6 days ago: if you distill a student and then need one embedding model to serve both a cheap first-pass filter and a high-fidelity rerank without storing two indexes, which technique gives you that "one vector, many budgets" property?
Show answer
Matryoshka Representation Learning (MRL) β a single embedding whose leading dimensions are independently usable, so you truncate (e.g., 1536 β 256) for coarse recall and use the full vector for fine-grained rerank, from one stored representation.
πΌ Market Signal
The skill that OPD sits inside is the top-paying AI band: LLM fine-tuning & inference specialists command $220Kβ$350K total comp, with senior LLM engineers running ~$195Kβ$265K base and north of $300K all-in where equity is real (per KORE1 & jobsbyculture 2026 salary guides). Surveys explicitly flag a $20Kβ$50K+ premium for specialists over generalists β and "knowing when a distilled 7B beats a frontier model for the task" is called out by name as a premium skill.
Adoption is the proof: Qwen3, MiMo, and GLM-5 all ship OPD in their 2026 post-training pipelines (per HuggingFace "Distillation in 2026" roundup), so this is table stakes for anyone claiming production post-training experience.
β‘ Action This Week
Run a minimal OPD sanity loop: take a small open student (e.g., Qwen3-1.7B) and a same-family teacher (Qwen3-8B), and on 20 math prompts compare off-policy SFT-on-teacher-transcripts vs. on-policy (student samples, teacher scores per-token) for 50 steps each. Definition of done = a plot or table showing the on-policy run's per-token KL dropping faster and matching/beating the off-policy run's held-out accuracy. Write it up as a "why on-policy beats SFT for reasoning, in one chart" post β a concrete, reproducible artifact that signals real post-training depth.