Frameworks & SDKs
LangChain
LangGraph
Vercel AI SDK
Mastra
LlamaIndex
Pydantic AI
We review LLM applications, agents, RAG workflows, and tool integrations to establish what the system can do on its own, and what it should have to ask permission for.
Technology coverage
We review how your application uses these frameworks to retrieve data, retain memory, and call tools.
This list is not exhaustive.
Describe the tools your agents can call, data they can access, and actions that need approval.
What we review
System instructions, prompt boundaries, injection paths, and control assumptions.
Sources, access, poisoning, sensitive data, provenance, and isolation.
Persistence, cross-user leakage, manipulation, retention, and deletion.
APIs, MCP servers, credentials, scope, delegation, and dangerous actions.
User, agent, service, and human-in-the-loop authorization.
Logs, anomaly signals, rollback, containment, and investigation.
Timing
An agent can read sensitive data or call business-critical tools.
Automation can change records, move value, deploy code, or trigger workflows.
RAG, memory, MCP, or multi-agent architecture changes authority boundaries.
Leadership needs a bounded view of AI risk before launch or expansion.
Our approach
Map the tools, data sources, and actions available to the agent. Document who can influence its inputs and where human approval is required.
Follow prompts and retrieved content through data access, tool calls, and approvals. Record the permission and data boundaries each flow must preserve.
Challenge the agent with prompt injection, unauthorized access, and tool misuse scenarios. Add security tests to your repo and record which controls hold or fail.
We always verify fixes through retesting, including changes to permissions, approvals, monitoring, and recovery. Tie the principal’s signed opinion to the reviewed commit, permissions, and evidence.
What you receive
Technical and workflow findings, each with the way it gets abused or fails on its own.
Agent access, human approval boundaries, and the results of tested attack scenarios.
Prompt, retrieval, tool call, approval, and the high-impact action at the end, each with the property it has to hold.
Each security property, its test result, and supporting evidence.
Regression tests for agent permissions, data access, and tool use, ready to rerun after fixes and future changes.
The principal’s conclusions, tied to the reviewed commit, permissions, and evidence.
Your review team
A principal leads each engagement. Your proposal names the specialists assigned to the scope.
Specialist Advisor
Security engineer and trainer with around 10 years of experience across application and protocol security. At ZKsync, reviewed Solidity and Rust code, including account abstraction, then built AI-assisted vulnerability-analysis workflows.
Specialist Advisor
Security specialist with 60+ Web3 reviews across DeFi, L1s, bridges, oracles, and other critical infrastructure. At Kraken, worked on the security of funding services, custody, APIs, and on-chain systems.
Founder & Partner
Led a security engineering practice for Rust and non-EVM systems across Substrate and NEAR. Earlier, built vulnerability-detection engines at Invicti used by Fortune 50 and public-sector organizations.
FAQ
Yes. We test how untrusted instructions can affect retrieval, memory, tool calls, and approvals. The review checks whether those inputs can expose data or trigger actions outside the intended permissions.
Yes. The review can cover your application’s instructions, data access, tool integrations, and approval controls. Access to the model provider’s internal systems is a separate scope question; we state any limits on what we can verify.
We need to understand the agent’s tools, data sources, identities, permissions, and expected behavior. For a review that includes code, we need access to that code during scoping, alongside architecture documentation, configuration, your main concerns, and your deadline. We also agree access to logs and test accounts for the workflows in scope.
We agree the test environment and permitted actions before testing. Test data, restricted credentials, and simulated tools can help isolate sensitive actions. Any production testing needs explicit boundaries and stop conditions.
Yes. We always provide remediation guidance and verify fixes through retesting. We rerun relevant scenarios and check changes to permissions, approvals, and data access. Regression tests help your engineers repeat those checks as the application changes.
Describe the agent’s tools, data access, and actions. We will define the review or design scope.
Discuss your scope