Book an Appointment
AI Products · Embodied AI · AIoT

Penetration testing for AI products: robots, AIoT devices and edge AI

Once AI is built into a device, a prompt injection turns into a movement, an opened lock or a camera feed on the wrong network. We model threats with STRIDE and MAESTRO, test everything from the debug interface to the LLM, and run cyber-physical attacks as a red team – along the OWASP LLM Top 10, AISVS, ISTG and IoT Top 10.

  • Threat modeling with STRIDE × MAESTRO
  • Hardware, firmware, wireless, cloud, model & LLM
  • Evidence for CRA, EN 18031 & AI Act
  • 01Sensors
  • 02Edge model
  • 03Actuators
  • 04Cloud & LLM
  • 05Wireless & app
  • 06Firmware
Six attack surfaces – one product
21.1 bn
connected IoT devices at the end of 2025 (forecast) – 39 bn expected by 2030
IoT Analytics, 10/2025
5 m
industrial robots in operation, more than 600,000 newly installed in 2025 alone
IFR World Robotics, 09/2026
up to 100%
jailbreak success against LLM-controlled robots in a research test, incl. Unitree Go2
RoboPAIR, UPenn 2024
> 20%
of the breached organizations studied reported a breach targeting AI models or applications
IBM Cost of a Data Breach, 07/2026
In brief

What is AI product penetration testing?

AI product penetration testing examines devices with embedded artificial intelligence – robots, cobots, humanoid systems, AI cameras, voice assistants, AI toys and edge AI in machinery – across all layers: hardware, firmware, wireless, cloud API, app, on-device model and LLM integration. The goal is to establish whether an attacker can take over the product digitally or make it perform unwanted actions in the physical world.

Pure LLM and agent applications without a device are covered by our Agentic AI Pentesting service, connected devices without AI features by our IoT Pentesting service.

Last updated: · Responsible: Valeri Milke, VamiSec GmbH

Key takeaways

  • 1AI products combine three risk levels: classic IoT weaknesses, AI-specific attacks and physical impact via actuators and sensors.
  • 2Threat modeling combines STRIDE (what can go wrong?) with the seven MAESTRO layers (where in the AI architecture?) – extended to cover the physical world.
  • 3Testing is performed against recognized standards: OWASP Top 10 for LLM Applications 2026, OWASP AISVS 1.01, OWASP ISTG, OWASP IoT Top 10, MITRE ATLAS and ETSI EN 303 645.
  • 4The reporting obligations of the Cyber Resilience Act have applied since 11 Sep 2026; all requirements apply from 11 Dec 2027 – including regular security testing.
  • 5Tests on robots with actuators only take place under an agreed safety concept: a defined test zone, emergency stop and reduced speeds.
Why AI products need dedicated testing

When AI gets a body, the risks add up.

An AI product is an IoT device, an AI system and – with a motor, lock or valve – a cyber-physical system all at once. Testing only one of these layers misses the attack chains in between.

Layer 1

The IoT legacy

Hardware, firmware and wireless bring along the well-known weaknesses of connected devices.

  • Hardcoded keys and default passwords
  • Open debug ports (UART, JTAG) and unsigned updates
  • Insecure cloud APIs and companion apps
Layer 2

The AI layer

On-device models and cloud LLMs open up new classes of attack.

  • Prompt injection via voice, text and camera images
  • Model extraction, adversarial examples, data poisoning
  • LLM outputs that turn into commands without validation
Layer 3

The physical impact

Actuators and sensors link digital attacks to real-world consequences.

  • Movements beyond safety limits
  • Camera and microphone as a bug in the room
  • Outages in production, care or security systems

Attack chain from research: from a calendar invitation to an opened window

  1. 1The attacker sends a calendar invitation with hidden instructions.
  2. 2The user asks the AI assistant to summarize her appointments.
  3. 3The LLM reads the invitation – the instructions end up in its context (indirect prompt injection).
  4. 4In response to a harmless “thanks”, the assistant calls smart home functions.
  5. 5Windows open, the boiler heats up: digital manipulation with physical impact.

Demonstrated by researchers from Tel Aviv University, the Technion and SafeBreach against Gemini and Google Home (Black Hat USA, August 2025). Google had rolled out countermeasures before publication, such as confirmations for risky actions. Paper “Invitation Is All You Need”

Which AI products we test

Six product classes, six attack profiles.

Each product class has its own focus areas. Select a category to see typical attack surfaces, the main threats, our testing focus and the relevant standards.

Robots, cobots, humanoids and robot dogs

Robots combine actuators, cameras, wireless connectivity and, increasingly, LLM-based planners. A compromised robot is not just a data leak but a safety risk for the people around it. From 20 Jan 2027, the Machinery Regulation requires control systems to withstand foreseeable malicious attacks and self-evolving systems not to leave their defined task and movement space.

ROS 2 / DDSBLE provisioningControl system & safety PLCTeleoperationLLM/VLA plannerFleet cloud

Typical threats

  • Command injection via wireless and provisioning interfaces
  • A planner jailbreak leads to dangerous movements
  • Covert telemetry to manufacturer servers
  • Bypassing speed, force and zone limits

What we test

  • Wireless, BLE and network services incl. key management
  • ROS 2/DDS configuration, SROS 2 and node authentication
  • Jailbreak-to-action scenarios against the planner in the test area
  • Separation of AI planning and the deterministic safety layer
Attack surface of an AI product

Ten layers – from the environment to the supply chain.

This is how we break down an AI product during testing. Each layer has its own threats, test methods and references. Click on a layer.

03 / 10

Edge model & NPU

The on-device model – image recognition, language model or vision-language-action model – is both an attack target and valuable know-how worth protecting.

Threats

  • Extraction of model files from flash or app
  • Side-channel attacks on accelerators (NPU, TPU)
  • Backdoors in pre-trained weights

Test methods

  • Memory and file system analysis for unprotected models
  • Review of encryption, signature and loading process
  • Robustness tests with adversarial examples
ReferencesML05:2023AISVS C11ISTG-PROC-SIDEC
Threat modeling

STRIDE × MAESTRO: finding threats systematically before attackers do.

Before we test, we model. STRIDE provides the six questions about what can go wrong. MAESTRO, the Cloud Security Alliance framework for agentic AI, tells you where in the AI architecture. For products with sensors and actuators, we add a layer for the physical world.

STRIDE

What can go wrong?

Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service and Elevation of Privilege – applied to every data flow and every trust boundary of the product.

MAESTRO

Where in the AI architecture?

Seven layers from Foundation Models to Agent Ecosystem, including cross-layer threats. For AI products, we extend this to the physical world with sensors, actuators and safety.

SSpoofingAuthentication
TTamperingIntegrity
RRepudiationNon-repudiation
IInformation DisclosureConfidentiality
DDenial of ServiceAvailability
EElevation of PrivilegeAuthorization
L1

Foundation Models

On-device models, language and vision-language-action models, cloud LLMs

S
Substituted model

A manipulated model poses as the manufacturer’s model because signature and hash are not verified at load time.

T
Backdoor in the weights

Trojanized weights or fine-tunes respond to a trigger – such as a pattern in the camera image – with a deliberately wrong decision.

R
Unclear model version

Without model ID, prompt and confidence in the log, there is no way to prove which model triggered an action.

I
Model extraction

Unencrypted model files or side channels on the NPU reveal architecture and know-how.

D
Resource attacks

Crafted inputs drive up compute load, latency and battery drain until real-time functions fail.

E
Jailbreak

Prompt injection or jailbreaks make the model ignore rules and roles (LLM01:2026).

Threat modeling results

  • Data flow diagram with trust boundaries incl. physical interfaces
  • Threat register based on STRIDE per MAESTRO layer
  • Rating by CVSS 4.0 and safety impact
  • Attack trees for the most critical scenarios
  • Test plan that feeds directly into pentesting and red teaming
  • Mapping to OWASP, MITRE ATLAS, CRA and AI Act
OWASP Top 10 for LLM Applications 2026

The LLM Top 10, translated into the physical world.

The edition of the OWASP list published in August 2026 describes the risks of LLM applications – and now explicitly counts attacks via images and audio as prompt injection. Inside a device, these risks take on a new quality:

LLM01:2026

Prompt Injection

A voice command from the TV, a sign in the camera image or a calendar invitation controls the device – also cross-modally via image and sound.

LLM02:2026

Sensitive Information Disclosure

The assistant reveals Wi-Fi keys, floor plans, faces or conversation content.

LLM03:2026

Excessive Agency

The model operates door locks, stove tops or robot arms without confirmation – the most consequential climber on the list.

LLM04:2026

Supply Chain

Pre-trained models, adapters and model hubs bring backdoors or malicious code onto the device.

LLM05:2026

Data and Model Poisoning

Poisoned fleet, training or fine-tuning data makes the robot overlook obstacles.

LLM06:2026

Unbounded Consumption

Endless requests drain the battery and API budget or overheat the NPU.

LLM07:2026

Misinformation

Hallucinated maintenance instructions or incorrect object detection lead to dangerous actions.

LLM08:2026

Hidden Context Exposure

The device’s system prompt, tool schemas and security rules can be extracted and used as a blueprint for attacks.

LLM09:2026

Vector and Embedding Weaknesses

Manipulated entries in the local knowledge store – such as maintenance manuals – steer responses.

LLM10:2026

Improper Output Handling

Unvalidated model outputs become motor commands, shell commands or ROS messages.

For agentic functions, we additionally test against the OWASP Top 10 for Agentic Applications 2026 (ASI01–ASI10). On request, we also map findings to the previous 2025 edition.

Standards & methodologies

What we measure against – and what we use it for.

We combine risk lists, verification standards, testing guides and formal standards. This makes findings comparable, traceable and directly usable for your technical documentation. All version information as of 10 Oct 2026.

Risk list2026 edition · 08/2026

OWASP Top 10 for LLM Applications 2026

Guardrail for all LLM functions – from Prompt Injection (LLM01) through Excessive Agency (LLM03) to Improper Output Handling (LLM10).

Verification standardv1.01 · 10/2026

OWASP AISVS

Testable requirements in 12 chapters (C1–C12) and three levels (L1–L3) – our basis for test plans and acceptance criteria for the AI layer.

Testing guidev1 · 11/2025

OWASP AI Testing Guide

32 test cases in four areas: AI application, model, infrastructure and data – from prompt injection to evasion attacks.

Red teaming methodologyv1.0 · 01/2025

OWASP GenAI Red Teaming Guide

Four phases from model evaluation to runtime and agent evaluation – the framework for our red teaming scenarios.

Knowledge basev2026.09

MITRE ATLAS

16 tactics and 120 techniques of real-world attacks on AI – including Physical Environment Access (AML.T0041) and physical deception tools such as adversarial stickers (AML.T0008.003).

Testing guidev1.0.1 · 06/2024

OWASP ISTG

Test cases per device component – processor, memory, firmware, interfaces, wireless, user interface – with attacker models for physical access and privileges.

Verification standard1.0.0-RC2 · draft

OWASP ISVS

Requirements for secure IoT ecosystems in five chapters, from ecosystem to hardware platform – the basis for hardening criteria.

Risk list2018 edition

OWASP IoT Top 10

The ten most common vulnerability classes of connected devices – still the common language in IoT pentesting.

Methodology9 phases

OWASP FSTM

Approach to firmware analysis: from information gathering through extraction and emulation to runtime analysis and binary exploitation.

SpecificationDDS Security v1.2 · 02/2026

OMG DDS Security & SROS 2

Security mechanisms of DDS, the middleware underneath ROS 2: authentication, access control and encryption between nodes, implemented with SROS 2.

StandardV3.1.3 · 09/2024

ETSI EN 303 645

13 groups of baseline requirements for consumer IoT – from “no universal default passwords” to input data validation.

Harmonised standardlisted with restrictions

EN 18031-1/-2/-3

Demonstrates compliance with the cybersecurity requirements of the Radio Equipment Directive. The restrictions concern, among other things, passwords, parental controls and update criteria.

StandardV2.1.1 · 12/2025

ETSI EN 304 223

13 principles across five lifecycle phases for AI models and systems – including mandatory security testing before release (Principle 9).

Standard series4-1:2018 · 4-2:2019

IEC 62443-4-1/-4-2

Secure development process (4-1) and technical requirements for components (4-2) for industrial products.

TaxonomyE2025 · 03/2025

NIST AI 100-2 E2025

Common terminology for evasion, poisoning, privacy attacks and misuse – including indirect prompt injection and agents.

OWASP IoT Top 10 – the foundation of every device test

  1. I1Weak, Guessable, or Hardcoded Passwords
  2. I2Insecure Network Services
  3. I3Insecure Ecosystem Interfaces
  4. I4Lack of Secure Update Mechanism
  5. I5Use of Insecure or Outdated Components
  6. I6Insufficient Privacy Protection
  7. I7Insecure Data Transfer and Storage
  8. I8Lack of Device Management
  9. I9Insecure Default Settings
  10. I10Lack of Physical Hardening

Official titles of the OWASP Internet of Things Top 10 (2018) – still the current edition today. For AI products, these classes remain relevant: they are often the entry point for attacks on the AI layer.

Services

Four modules – individually or as a complete package.

From the first architecture workshop to the evidence package for conformity assessment. The modules build on each other but can also be commissioned individually.

01
Module 1 · Design

Threat modeling for AI products

We model your product with STRIDE and MAESTRO – including the physical layer – and derive prioritized risks and a test plan.

  • Workshops with engineering, safety and product management
  • Data flow diagram with trust boundaries
  • Threat register with CVSS 4.0 and safety rating
  • Measures and test plan for the next steps

Supported by VamiThreat: STRIDE analysis, MAESTRO and automatic mapping to MITRE ATT&CK and ATLAS.

More on AI threat modeling
02
Module 2 · Test

Penetration test of the AI product

Grey-box or white-box testing across all layers: hardware, firmware, wireless, cloud API, app, on-device model and LLM integration.

  • Hardware and debug interfaces, memory readout
  • Firmware analysis and reverse engineering
  • Wireless, pairing and network services
  • Cloud API, app and LLM following OWASP methodology

AI-assisted firmware and binary analysis with VamiReverse.

Request a pentest
03
Module 3 · Attack

AI red teaming & cyber-physical scenarios

Goal-oriented attack simulation across layer boundaries: can an attacker make your product perform an unwanted action in the real world?

  • Jailbreak-to-action chains against planners and assistants
  • Physical prompt injection, adversarial patches, audio attacks
  • Fleet scenarios via cloud, app and remote maintenance
  • TTPs along MITRE ATLAS, safely in the test area

Scaled with the AI Pentesting Agent from VamiRedteam (AI-specific test suites based on MITRE ATLAS).

Discover VamiRedteam
04
Module 4 · Evidence

Hardening & compliance evidence

We translate findings into evidence for your technical documentation and support remediation through to a successful retest.

  • Mapping to CRA Annex I, EN 18031, ETSI EN 303 645
  • Reference to AI Act Art. 15 and the Machinery Regulation
  • SBOM and AI-BOM, vulnerability and reporting process
  • Retest and retest report

SBOM per release and vulnerability tracking with VamiAppSec.

Explore the Cyber Resilience Act
Approach

In six steps from scoping to evidence.

The threat model drives the test: we first examine what is truly critical for your product.

  1. 1

    Scoping & safety briefing

    Product variants, test scope, test environment, access and rules of engagement – including a safety concept where actuators are involved.

    Output: Test agreement and safety approval

  2. 2

    Threat modeling

    STRIDE per data flow and trust boundary, structured by MAESTRO layers plus the physical world.

    Output: Threat register and attack trees

  3. 3

    Test plan

    Test cases derived from the threat model, mapped to AISVS, ISTG, LLM Top 10 and IoT Top 10.

    Output: Prioritized test plan

  4. 4

    Component pentest

    Hardware, firmware, wireless, cloud API, app, model and LLM – each layer with suitable methods.

    Output: Validated findings with proof of concept

  5. 5

    End-to-end red teaming

    Chained scenarios across layer boundaries: from initial access to impact in the physical world.

    Output: Attack chains with storyboard

  6. 6

    Report, retest & evidence

    Management summary, technical findings, measures and standards mapping – after remediation, we test again.

    Output: Report, retest report, evidence package

Safety first: we only perform tests on robots and machinery with actuators under an agreed safety concept – a defined test zone, emergency stop within reach, reduced speeds and approval by your safety officers. Where possible, we test in simulation or on a digital twin first.

How it differs

IoT pentest, LLM pentest – or both together?

An AI product needs the depth of an IoT pentest and the methods of an LLM pentest – plus a view of the physical consequences.

Comparison of IoT pentesting, LLM/agentic pentesting and AI product pentesting
Test areaClassic IoT pentestLLM/agentic pentestAI product pentest
Hardware, debug interfaces & firmwarecoverednot coveredcovered
Wireless & pairing (BLE, Wi-Fi, Matter, Zigbee)coverednot coveredcovered
Cloud backend, APIs & companion appcoveredpartially coveredcovered
Prompt injection & jailbreaks (text, voice, image)not coveredcoveredcovered
Model extraction, adversarial examples & poisoningnot coveredpartially coveredcovered
Physical consequences & safety limits of actuatorsnot coverednot coveredcovered
Threat model based on STRIDE × MAESTRO incl. physical layerpartially coveredpartially coveredcovered
Evidence for CRA, EN 18031, AI Act & Machinery Regulationpartially coveredpartially coveredcovered

covered partially covered not covered

Documented cases

Not theory: attacks on AI products from research and practice.

A selection of documented vulnerabilities and research results from the past two years – each with its source.

2025Humanoids & robot dogs

UniPwn: root via Bluetooth on Unitree robots

Hardcoded AES keys in BLE provisioning allowed command injection with root privileges on the G1, H1, Go2 and B2. The researchers showed that an infected robot can spread to others within radio range. Four CVEs were assigned (CVE-2025-35027, -60017, -60250, -60251).

What we test as a result: Provisioning, wireless and key management in every robot test.

IEEE Spectrum, 09/2025
2025Humanoids

Covert telemetry in the Unitree G1

According to Alias Robotics, the G1 humanoid robot sends sensor and status data via MQTT to external servers every 300 seconds – without notifying the operator. In addition, the BLE interface uses a static AES key that is identical on all devices.

What we test as a result: Data flow analysis and review of hidden connections.

arXiv 2509.14139
2025Smart home & LLM

Calendar invitation controls Google Home

Hidden instructions in calendar invitations got Gemini to operate smart home devices. The researchers rated 73% of the threats demonstrated as high or critical; Google rolled out countermeasures.

What we test as a result: Indirect prompt injection via all connected data sources.

arXiv 2508.12175
2026Drones & vehicles

Physical prompt injection via signs in the camera view

Optimized text on signs in the camera view hijacked vision-language agents: 95.5% success in drone object tracking in simulation, up to 92.5% on a real robotic vehicle with GPT-4o.

What we test as a result: The environment as an attack vector in red teaming.

arXiv 2510.00181 (IEEE SaTML 2026)
2025Robot AI (VLA)

Adversarial patches against vision-language-action models

A small colorful patch in the field of view caused robots running OpenVLA to fail their tasks – in simulation up to 100% in the best attack scenario, in the physical experiment in more than 43% of cases.

What we test as a result: Model robustness tests under real-world conditions.

ICCV 2025, arXiv 2411.13587
2026AI toys

Around 50,000 children’s chats from an AI plush toy exposed

The parent portal of the AI plush toy Bondu only checked whether a Google account was present. As a result, children’s conversation logs, names and dates of birth were visible to third parties. The manufacturer closed the gap immediately.

What we test as a result: Authorization of portals, APIs and logs.

Malwarebytes, 02/2026
2025AI cameras

AI cameras accessible without login

At least 60 AI-powered cameras from Flock Safety exposed live streams, archives and admin panels to the internet without authentication. The manufacturer described it as a misconfiguration that has been fixed.

What we test as a result: External attack surface and authentication of every interface.

404 Media, 12/2025
2026Cobots

CVSS 9.8 in a cobot controller

CVE-2026-8153 in PolyScope 5 from Universal Robots allowed unauthenticated operating system commands via the Dashboard Server – according to the researcher who discovered it (Claroty), with consequences extending to entire cobot fleets, provided the Dashboard Server is enabled and reachable.

What we test as a result: Network services and segmentation of robot controllers.

SecurityWeek, 05/2026
2024Edge AI

Model architecture extracted via side channel

Using electromagnetic emanations, researchers at NC State University reconstructed the hyperparameters of neural networks on a Google Edge TPU with 99.91% accuracy – assuming physical access.

What we test as a result: Protecting the model as intellectual property.

IACR TCHES 2025
Regulation

Deadlines that affect AI products.

Product safety, the Radio Equipment Directive, the Cyber Resilience Act, product liability, the Machinery Regulation and the AI Act interlock. Security testing is becoming a mandatory building block rather than optional evidence.

  1. GPSR

    Product safety includes cybersecurity

    The General Product Safety Regulation (EU) 2023/988 requires cybersecurity features as well as learning and predictive functionalities to be included in the safety assessment.

  2. RED

    Cybersecurity for radio equipment

    Delegated Regulation (EU) 2022/30 applies – compliance can be demonstrated e.g. via EN 18031-1/-2/-3, which are only listed with restrictions. From 11 Dec 2027, the CRA takes over.

  3. AI Act

    Transparency obligations under Art. 50

    Anyone interacting with an AI system – such as a voice assistant – must be informed of this, unless it is obvious. For marking AI-generated content (para. 2), legacy systems have a transition period until 2 Dec 2026.

  4. CRA

    Reporting obligations under the Cyber Resilience Act

    Actively exploited vulnerabilities and severe incidents must be reported via the ENISA reporting platform: early warning within 24 hours, notification within 72 hours – also for products already on the market.

  5. Product Liability

    New Product Liability Directive

    For products placed on the market after 9 Dec 2026: software – including AI – counts as a product, and safety-relevant cybersecurity requirements are taken into account. A lack of security updates does not relieve the manufacturer of liability where they are within its control (Art. 7, 11 Directive (EU) 2024/2853).

  6. Machinery

    EU Machinery Regulation applies

    Protection against corruption (Annex III 1.1.9) and attack-resistant control systems (1.2.1). Safety functions with ML-based self-evolving behavior require third-party conformity assessment.

  7. AI Act

    High-risk AI under Annex III

    Among other things, remote biometric identification and emotion recognition – relevant for AI cameras with facial recognition.

  8. CRA

    Cyber Resilience Act fully applicable

    All requirements of Annex I apply – including effective and regular security testing. At the same time, the Delegated Regulation under the Radio Equipment Directive ceases to apply.

  9. AI Act

    High-risk AI in products under Annex I

    For AI as a safety component in products such as toys, radio equipment and medical devices, the high-risk obligations including Art. 15 apply where a third-party conformity assessment is required. For machinery, the AI requirements come via delegated acts under the Machinery Regulation, which apply by 2 Aug 2028 at the latest.

As of 10 Oct 2026. AI Act dates reflect the Digital Omnibus on AI (Regulation (EU) 2026/1744), which moved machinery to Annex I Section B. Not legal advice – we provide technical support in demonstrating compliance.

Deliverables

What you walk away with.

Management summary

Risk posture, critical attack chains and recommendations on two pages – for executive management and product owners.

Threat model

Data flow diagram, trust boundaries and threat register based on STRIDE × MAESTRO – as a living document for engineering.

Findings report with proof of concept

Every finding reproducible, rated by CVSS 4.0 and – where actuators are involved – by safety impact.

Attack chains as a storyboard

Red teaming scenarios step by step: initial access, escalation, physical impact.

Standards and regulatory mapping

Mapping to OWASP, MITRE ATLAS, ETSI EN 303 645, EN 18031, CRA and AI Act.

Prioritized action plan

Concrete fixes for hardware, firmware, cloud and model – sorted by risk and effort.

Retest and retest report

After remediation, we re-test the findings and document their status.

Evidence package

Test reports and mappings prepared for technical documentation and conformity assessment.

Glossary

Key terms, briefly explained.

Embodied AI
AI that interacts with the world through a physical body – such as robots, drones or autonomous vehicles with sensors and actuators.
AIoT (Artificial Intelligence of Things)
Connected devices that use AI functions on the device, at the edge or in the cloud.
Edge AI
AI models that run directly on the device or a nearby gateway instead of in the cloud – often on specialized accelerators (NPU).
Vision-Language-Action model (VLA)
A model that processes camera images and language instructions and directly generates control commands for a robot from them.
Prompt injection
An attack in which inputs – direct or hidden in data such as emails, images or calendar entries – cause a language model to behave in unwanted ways.
Physical prompt injection
Prompt injection via the physical environment, such as text on signs or objects in a camera’s field of view.
Adversarial example / patch
A deliberately altered input – such as a printed pattern – that tricks a model into a wrong detection or decision.
Model extraction
Theft of a model or its architecture, for example by reading out files, systematic queries or side channels.
STRIDE
Threat modeling method from Microsoft with six threat categories: Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege.
MAESTRO
“Multi-Agent Environment, Security, Threat, Risk, and Outcome” – the Cloud Security Alliance’s threat modeling framework (2025) with seven layers for agentic AI.
FAQ

Frequently asked questions about AI product pentesting

What is AI product penetration testing?

AI product penetration testing is a security test for devices with embedded AI – such as robots, AI cameras, voice assistants or edge AI in machinery. It covers hardware, firmware, wireless, cloud API, app, the on-device model and the LLM integration. The central question is whether an attacker can take over the product or make it perform actions in the physical world.

How does it differ from a classic IoT pentest?

An IoT pentest examines hardware, firmware, wireless and cloud. With an AI product, attacks on the model and the LLM integration are added – prompt injection, jailbreaks, model extraction, adversarial examples – as well as the question of what physical consequences an attack via actuators and sensors can have. For devices without AI functionality, we recommend our IoT pentesting.

Can AI robots and humanoid robots be hacked?

Yes. Documented cases include root access via Bluetooth to Unitree robots (UniPwn, 2025), jailbreaks of LLM-controlled robots with success rates of up to 100% in a research test (RoboPAIR, 2024) and a critical vulnerability in the cobot software from Universal Robots (CVE-2026-8153, CVSS 9.8). Typical entry points are wireless interfaces, network services, cloud APIs and the AI planner itself.

Can a prompt injection make a robot perform physical actions?

Yes, if model outputs become commands without validation or the model has overly broad permissions – in the 2026 OWASP edition, these are Improper Output Handling (LLM10) and Excessive Agency (LLM03). Researchers have shown that voice commands, text in the camera image or calendar invitations can control robots, drones and smart home devices. That is why we specifically test whether a deterministic safety layer is in place between AI planning and actuators.

Which standards do you use?

For the AI layer: the OWASP Top 10 for LLM Applications 2026, the OWASP Top 10 for Agentic Applications 2026, OWASP AISVS 1.01, the OWASP AI Testing Guide, the OWASP GenAI Red Teaming Guide and MITRE ATLAS. For device and wireless: OWASP ISTG, OWASP ISVS, the OWASP IoT Top 10, the OWASP FSTM firmware methodology and, for robots, DDS Security with SROS 2. As formal standards: ETSI EN 303 645, EN 18031, ETSI EN 304 223 and IEC 62443-4-1/-4-2.

What is MAESTRO and why do you combine it with STRIDE?

MAESTRO is a threat modeling framework from the Cloud Security Alliance for agentic AI, with seven architecture layers from Foundation Models to Agent Ecosystem. STRIDE answers what type of threat is involved, MAESTRO where in the AI architecture it arises. For products with sensors and actuators, we add a layer for the physical world – Microsoft already lists adversarial examples in the physical domain as a separate threat class in its threat modeling for AI/ML systems.

Does the Cyber Resilience Act require a penetration test?

The CRA does not prescribe a procedure called a pentest, but Annex I Part II does require effective and regular tests and reviews of the security of the product. A pentest is the common way to demonstrate this. The reporting obligations have applied since 11 Sep 2026, all other requirements apply from 11 Dec 2027.

Are smart speakers, security cameras and AI toys important products under the CRA?

Yes: in Class I, Annex III of the CRA lists smart home general-purpose virtual assistants (No. 16), smart home products with security functionalities such as security cameras, baby monitors and alarm systems (No. 17), connected toys with social interactive or location tracking features (No. 18) and certain wearables (No. 19). Implementing Regulation (EU) 2025/2392 explicitly mentions smart speakers with voice assistants. Self-assessment (Module A) is only possible for Class I if harmonised standards, common specifications or a European cybersecurity certification scheme at assurance level “substantial” or higher are applied in full (Art. 32(2) CRA) – as long as no standards have been published in the Official Journal, the route usually leads through a notified body.

When does the AI Act apply to AI in robots, machinery and devices?

That depends on the applicable product legislation. For AI as a safety component in products under Annex I Section A – such as toys, radio equipment or medical devices – the high-risk obligations including Art. 15 apply from 2 Aug 2028 under the Digital Omnibus, where a third-party conformity assessment is required. The Omnibus moved machinery to Section B: for AI in robots and machinery, the requirements come via delegated acts under the Machinery Regulation, which itself applies from 20 Jan 2027. The transparency obligations under Art. 50 have already applied since 2 Aug 2026.

How do you test robots without endangering people?

With an agreed safety concept: a defined test zone, emergency stop within reach, reduced speeds and approval by the manufacturer’s safety officers. We first test critical scenarios in simulation or on a digital twin before reproducing them on the real system.

Do you also test the on-device models?

Yes. We check whether model files are stored unprotected on the device or in the app, whether the loading process and signature are secure, and how robust the model is against adversarial examples, manipulated sensor data and poisoning. Where relevant, we also assess side-channel risks on accelerators (ISTG-PROC-SIDEC).

Do you find hidden telemetry and backdoors?

Yes. We analyze firmware, services and network traffic for undocumented connections, remote maintenance tunnels and hidden functions. Examples from research include the covert telemetry of the Unitree G1 and the remote access tunnel in the Unitree Go1 (CVE-2025-2894).

What do you need from us – black, grey or white box?

A grey-box or white-box approach is most effective: two to three test devices, access to the app and cloud test environment, architecture documentation and – if possible – firmware images and source code excerpts. A black-box test is possible, but covers less attack surface in the same amount of time.

How much does an AI product pentest cost and how long does it take?

That depends on the product class, the number of interfaces, testing depth and the modules you need. After a free scoping call, you receive a fixed-price quote. A focused test of a single device including the report usually takes a few weeks.

How often should an AI product be retested?

At a minimum before every major product release and after significant changes to firmware, model or cloud functions. Because models and attack techniques evolve quickly, we also recommend an annual retest – the CRA requires regular testing throughout the entire support period.

Sources

Sources & primary literature

  1. OWASP Top 10 for LLM Applications 2026 — OWASP GenAI Security Project, 08/2026
  2. OWASP Top 10 for Agentic Applications 2026 — OWASP GenAI Security Project, 12/2025
  3. OWASP AISVS – Artificial Intelligence Security Verification Standard — OWASP, v1.01
  4. OWASP AI Testing Guide — OWASP, v1 (11/2025)
  5. OWASP ISTG – IoT Security Testing Guide — OWASP, v1.0.1
  6. OWASP Internet of Things Top 10 (2018) — OWASP
  7. MAESTRO: Agentic AI Threat Modeling Framework — Cloud Security Alliance, 02/2025
  8. Threat Modeling AI/ML Systems and Dependencies — Microsoft
  9. MITRE ATLAS — MITRE, v2026.09
  10. Regulation (EU) 2024/2847 – Cyber Resilience Act — EUR-Lex
  11. Regulation (EU) 2026/1744 – Digital Omnibus on AI — EUR-Lex
  12. Regulation (EU) 2023/1230 – Machinery Regulation — EUR-Lex
  13. Directive (EU) 2024/2853 – Product Liability — EUR-Lex
  14. Implementing Decision (EU) 2025/138 – EN 18031 — EUR-Lex
  15. ETSI EN 304 223: Baseline Cyber Security Requirements for AI — ETSI, V2.1.1 (12/2025)
  16. NIST AI 100-2 E2025: Adversarial Machine Learning — NIST, 03/2025
  17. Robey et al.: Jailbreaking LLM-Controlled Robots (RoboPAIR) — arXiv 2410.13691, 2024
  18. Nassi et al.: Invitation Is All You Need — arXiv 2508.12175, 2025
  19. Mayoral-Vilches et al.: Cybersecurity AI: Humanoid Robots as Attack Vectors — arXiv 2509.14139, 2025
  20. UniPwn: Unitree robots can be taken over via Bluetooth — IEEE Spectrum, 09/2025
  21. NVD: CVE-2025-35027 (UniPwn, Unitree) — NIST National Vulnerability Database
  22. CHAI: Command Hijacking against Embodied AI — arXiv 2510.00181
  23. Wang et al.: Adversarial Vulnerabilities of VLA Models in Robotics — ICCV 2025
  24. TPUXtract: hyperparameter extraction on Google Edge TPU — IACR TCHES 2025
  25. State of IoT 2025: 21.1 bn connected devices — IoT Analytics, 10/2025
  26. World Robotics 2026: five million industrial robots — IFR, 09/2026
  27. Cost of a Data Breach Report 2026 — IBM, 07/2026
Valeri Milke – Founder & CEO of VamiSec GmbH
Valeri MilkeFounder & CEO · VamiSec GmbH
Your contact

“With AI products, the question is no longer just whether an attacker can get in – but what they can make the device do in the real world.”

We combine hardware and IoT pentesting, AI red teaming and regulatory expertise in the CRA, AI Act and Machinery Regulation – for robust security evidence before market launch and throughout the entire support period.