Skip to content

x86 support

picaro packs one architecture at a time. build.output.arch decides which:

archStub target (Windows / cross)Payloadwindows import lib
x64 (default)x86_64-pc-windows-msvc / -gnuPE32+windows_x86_64_msvc / _gnu
x86i686-pc-windows-msvc / -gnuPE32windows_i686_msvc / _gnu

The rule: the stub and the payload share an architecture

Section titled “The rule: the stub and the payload share an architecture”

Reflective loading runs the payload in the stub’s own process, so the two must match: 32-bit code cannot run in a 64-bit process, and vice versa. A plan whose payload does not match output.arch is rejected at build time, before anything is encrypted or compiled:

build.output.arch is 'x64' but the payload is an x86 PE (machine 0x014c):
the stub and the payload must share an architecture

An arch: x86 artifact is a 32-bit PE. On 64-bit Windows it runs under WOW64 and sees the WOW64 filesystem redirection (C:\Windows\System32\… resolves to C:\Windows\SysWOW64\…, which is usually what you want — the 32-bit svchost.exe and the 32-bit ntdll.dll). A 64-bit stub cannot run on 32-bit-only Windows.

MilestoneStateWhat it covers
M0 — foundationsdonePE32 parsing, schema/validation coupling, capability gating, i686 compiler, PE32 imports, x86 args, fixtures
M1 — reflective_loadingdonePE32 mapper, 32-bit IMAGE_TLS_DIRECTORY, x86 TLS provisioning, x86 entry call
M2 — process_hollowingdonePE32 target parsing, CONTEXT32, Eip hand-off, fs:[0x2C] TLS slot, WOW64 target
M3 — preparationsdoneanti_debug (x86 debug registers), patch_amsi, patch_etw, bouncer
M4 — x86 indirect syscallsparkedneeds a Heaven’s Gate thunk and 64-bit SSN resolution — a redesign, not a flag (see below)
M5 — .NET hostingdoneformat: dotnet + dotnet_hosting: the stub hosts the CLR and loads the assembly from memory (x86 and x64)

Verified end to end on Windows (32-bit payloads, under WOW64 on an x64 host):

  • packs_and_runs_an_x86_reflective_payload — a PE32 payload mapped into the 32-bit stub prints its marker and reads its own thread_local;
  • packs_and_runs_an_x86_hollowed_payload — a PE32 payload hollowed into the WOW64 svchost.exe writes its marker, reads tls=42 and its exit code (42) comes back through the stub;
  • x86_preparations_run_before_the_payload — anti_debug, patch_amsi, patch_etw and bouncer run in a 32-bit stub before the payload;
  • packs_and_runs_a_managed_payload — a 32-bit managed assembly runs on the CLR hosted by the 32-bit stub, receives the payload args, reports x64=False and its exit code (42) comes back through the stub;
  • packs_and_runs_sharphound — the real-world x86 assembly in the repository root packs and its output is identical to the unpacked one.

The foundations (M0) are unchanged from before: the schema accepts arch: x86, the compiler builds a valid i686 stub, and the shared blocks are architecture-aware:

  • engine::pe parses PE32 and PE32+, and build enforces that the payload’s arch matches output.arch.
  • The compiler is arch-parameterized: target triple, dependency set and windows_i686_* / windows_x86_64_* import library.
  • The shared import resolver reads the image’s optional-header magic: 4-byte thunks and IMAGE_ORDINAL_FLAG32 for PE32, 8-byte and IMAGE_ORDINAL_FLAG64 for PE32+, with the matching IAT slot width.
  • The payload-argument block uses the right PEB offsets (fs:[0x30], +0x10, +0x40 on x86; gs:[0x60], +0x20, +0x70 on x64).
  • Resources and audit handle both layouts; the CI and the Docker image install and build the i686 target.

reflective_loading and process_hollowing declare WindowsX64 | WindowsX86; so do anti_debug, patch_amsi, patch_etw and bouncer. sleep_obfuscation stays x64-only, and an arch: x86 pipeline that uses it is rejected at validation with a message naming the technique:

technique 'sleep_obfuscation' does not support build.output.arch 'x86' yet:
it requires x64; the packer can compile an x86 stub, but this technique has no
x86 fragment
  • reflective_loading. IMAGE_NT_HEADERS32, the hand-defined 32-bit IMAGE_TLS_DIRECTORY, 4-byte IAT slots and IMAGE_REL_BASED_HIGHLOW relocations come from the shared mapper; the entry point is called through the x86 ABI (arguments on the stack — there is no register hand-off).
  • process_hollowing. PE32 offsets in the payload parser (ImageBase is a u32 at a different offset), the native CONTEXT32, the PEB’s ImageBaseAddress at +0x08 read from Ebx, the entry written to Eip, and the thread’s TLS slot at fs:[0x2C] with 4-byte entries. The default target (System32\svchost.exe) is redirected by WOW64 to SysWOW64\svchost.exe.
  • TLS provisioning. Both techniques give a manually mapped image a real TLS index, block and slot; the arrays differ per architecture (see the technique READMEs and docs/measurements.md).
  • dotnet_hosting. The COM interfaces and vtable slots are the same on both architectures; what differs is the VARIANT layout (16 bytes on x86, 24 on x64) and the Invoke_3 ABI — the variant is passed by value on x86 and by reference on x64, and the fragment asserts the size it expects at build time.
  • Architecture must match. See the rule above; it is a hard build-time check.
  • syscalls.mode: indirect is x64-only, and M4 parked it deliberately. The x64 path uses a syscall; ret gadget plus an x64 global_asm! trampoline. A 32-bit process cannot execute syscall at all: it runs in compatibility mode, where the instruction is invalid, and the only way in is Heaven’s Gate — a far transfer to a 64-bit code segment (0x33) around a hand-written 64-bit thunk. On top of that the SSNs a 32-bit stub can read from the live ntdll stubs are the WOW64 service numbers, which the kernel maps through a separate table; they are not the numbers the 64-bit syscall path expects. A real x86 implementation therefore needs the gate thunk and 64-bit SSN resolution (the 64-bit ntdll is mapped above 4 GB and unreachable from 32-bit pointers), so it is a redesign of the syscall layer, not a per-architecture flag. mode: indirect is rejected on arch: x86 instead of silently degrading. Because api_resolution.mode: proxy requires indirect, it is x64-only too.
  • .NET assemblies need CLR hosting (done, M5). A managed assembly’s entry point calls mscoree.dll!_CorExeMain, which starts the CLR for the process’s main module — the stub, not the mapped payload — so reflectively mapping a .NET assembly does not run it (that is now a build-time error). A managed payload declares format: dotnet and the dotnet_hosting execution technique: the stub hosts the CLR (CLRCreateInstance → ICorRuntimeHost → _AppDomain.Load_3 → GetEntryPoint → Invoke_3) and loads the assembly from memory. The assembly runs on the CLR of the stub’s bitness, so a 32BITREQUIRED assembly needs arch: x86 and an AnyCPU/IL-only one runs on either — checked at build time (SharpHound.exe, 32-bit .NET, is the acceptance case).
  • A mapped image’s TLS index can collide with a loaded module’s. TlsAlloc returns the lowest free index and the index has to be reserved that way (see docs/measurements.md); when it coincides with a module’s _tls_index, that module’s thread-locals in the affected thread are the payload’s template bytes. The stub logs the index under build.stub.debug.
  • sleep_obfuscation (ekko) is x64-only. Its ROP chain is built on the x64 ABI (Rip/Rsp, Rcx–R9); an x86 port needs the cdecl/stack convention.
  • process_hollowing needs an arch-matching target. Under WOW64 the System32\svchost.exe literal already resolves to the 32-bit one, so the default happens to work for an x86 stub — but it is implicit, and a target given as $arg.N$ cannot be checked at build time.
  • shellcode has no architecture in its bytes, so it follows output.arch.
  • x86 SEH is stack-based (fs:[0] chain), while x64 is table-based (.pdata/.xdata). x86 needs no unwind-table registration for a mapped image, but the x64 side does not register one either today, so payload exception unwinding inside a mapped image is unsupported on x64 and is a separate feature. IMAGE_DIRECTORY_ENTRY_EXCEPTION is not meaningful in PE32.
  • TLS callbacks under process hollowing are not run (a never-started thread has no TLS array to point at) — unchanged by architecture.
  • dotnet_hosting is reflective_loading-only in practice. The CLR is hosted in the stub’s own process, and a pipeline has exactly one execution technique, so process_hollowing and dotnet_hosting cannot be combined; format: dotnet and dotnet_hosting imply each other at validation.