mirror of
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
synced 2026-09-18 23:09:29 +02:00
Document how pre-opened interpreters are accounted. Link: https://patch.msgid.link/20260803-work-binfmt_misc-interplimit-v1-3-4a2435500bd9@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
392 lines
20 KiB
ReStructuredText
392 lines
20 KiB
ReStructuredText
Kernel Support for miscellaneous Binary Formats (binfmt_misc)
|
|
=============================================================
|
|
|
|
This Kernel feature allows you to invoke almost (for restrictions see below)
|
|
every program by simply typing its name in the shell.
|
|
This includes for example compiled Java(TM), Python or Emacs programs.
|
|
|
|
To achieve this you must tell binfmt_misc which interpreter has to be invoked
|
|
with which binary. Binfmt_misc recognises the binary-type by matching some bytes
|
|
at the beginning of the file with a magic byte sequence (masking out specified
|
|
bits) you have supplied. Binfmt_misc can also recognise a filename extension
|
|
aka ``.com`` or ``.exe``.
|
|
|
|
First you must mount binfmt_misc::
|
|
|
|
mount binfmt_misc -t binfmt_misc /proc/sys/fs/binfmt_misc
|
|
|
|
To actually register a new binary type, you have to set up a string looking like
|
|
``:name:type:offset:magic:mask:interpreter:flags`` (where you can choose the
|
|
``:`` upon your needs) and echo it to ``/proc/sys/fs/binfmt_misc/register``.
|
|
|
|
Here is what the fields mean:
|
|
|
|
- ``name``
|
|
is an identifier string. A new /proc file will be created with this
|
|
name below ``/proc/sys/fs/binfmt_misc``; cannot contain slashes ``/`` for
|
|
obvious reasons.
|
|
- ``type``
|
|
is the type of recognition. Give ``M`` for magic, ``E`` for extension and
|
|
``B`` for a bpf-backed handler (see below).
|
|
- ``offset``
|
|
is the offset of the magic/mask in the file, counted in bytes. This
|
|
defaults to 0 if you omit it (i.e. you write ``:name:type::magic...``).
|
|
Ignored when using filename extension matching.
|
|
- ``magic``
|
|
is the byte sequence binfmt_misc is matching for. The magic string
|
|
may contain hex-encoded characters like ``\x0a`` or ``\xA4``. Note that you
|
|
must escape any NUL bytes; parsing halts at the first one. In a shell
|
|
environment you might have to write ``\\x0a`` to prevent the shell from
|
|
eating your ``\``.
|
|
If you chose filename extension matching, this is the extension to be
|
|
recognised (without the ``.``, the ``\x0a`` specials are not allowed).
|
|
Extension matching is case sensitive, and slashes ``/`` are not allowed!
|
|
- ``mask``
|
|
is an (optional, defaults to all 0xff) mask. You can mask out some
|
|
bits from matching by supplying a string like magic and as long as magic.
|
|
The mask is anded with the byte sequence of the file. Note that you must
|
|
escape any NUL bytes; parsing halts at the first one. Ignored when using
|
|
filename extension matching.
|
|
- ``interpreter``
|
|
is the program that should be invoked with the binary as first
|
|
argument (specify the full path). For ``B`` entries this field
|
|
carries the name of the bpf handler instead (see below).
|
|
- ``flags``
|
|
is an optional field that controls several aspects of the invocation
|
|
of the interpreter. It is a string of capital letters, each controls a
|
|
certain aspect. The following flags are supported:
|
|
|
|
``P`` - preserve-argv[0]
|
|
Legacy behavior of binfmt_misc is to overwrite
|
|
the original argv[0] with the full path to the binary. When this
|
|
flag is included, binfmt_misc will add an argument to the argument
|
|
vector for this purpose, thus preserving the original ``argv[0]``.
|
|
e.g. If your interp is set to ``/bin/foo`` and you run ``blah``
|
|
(which is in ``/usr/local/bin``), then the kernel will execute
|
|
``/bin/foo`` with ``argv[]`` set to ``["/bin/foo", "/usr/local/bin/blah", "blah"]``. The interp has to be aware of this so it can
|
|
execute ``/usr/local/bin/blah``
|
|
with ``argv[]`` set to ``["blah"]``.
|
|
``O`` - open-binary
|
|
Legacy behavior of binfmt_misc is to pass the full path
|
|
of the binary to the interpreter as an argument. When this flag is
|
|
included, binfmt_misc will open the file for reading and pass its
|
|
descriptor into the auxilary vector with the key "AT_EXECFD", thus
|
|
allowing the interpreter to execute non-readable binaries. This
|
|
feature should be used with care - the interpreter has to be trusted
|
|
not to emit the contents of the non-readable binary.
|
|
``C`` - credentials
|
|
Currently, the behavior of binfmt_misc is to calculate
|
|
the credentials and security token of the new process according to
|
|
the interpreter. When this flag is included, these attributes are
|
|
calculated according to the binary. It also implies the ``O`` flag.
|
|
This feature should be used with care as the interpreter
|
|
will run with root permissions when a setuid binary owned by root
|
|
is run with binfmt_misc.
|
|
``F`` - fix binary
|
|
The usual behaviour of binfmt_misc is to spawn the
|
|
binary lazily when the misc format file is invoked. However,
|
|
this doesn't work very well in the face of mount namespaces and
|
|
changeroots, so the ``F`` mode opens the binary as soon as the
|
|
emulation is installed and uses the opened image to spawn the
|
|
emulator, meaning it is always available once installed,
|
|
regardless of how the environment changes.
|
|
``T`` - transparent
|
|
Run the interpreter transparently. The binary is handed to
|
|
the interpreter through ``AT_EXECFD`` (``T`` implies ``O``),
|
|
the argument vector is left exactly as the caller built it
|
|
and the kernel labels ``/proc/pid/exe`` with the binary
|
|
instead of the interpreter. The interpreter has to load the
|
|
binary from ``AT_EXECFD`` and follow the
|
|
``AT_FLAGS_TRANSPARENT_INTERP`` contract. Combining ``T``
|
|
with ``P`` is rejected: transparency preserves the whole
|
|
argument vector, argv[0] included.
|
|
``L`` - loader substitution
|
|
Do not run the interpreter on the binary at all: load the
|
|
binary itself as a fully native exec and substitute the
|
|
interpreter for the loader named in the binary's
|
|
``PT_INTERP``. See the "Loader substitution" section
|
|
below. ``L`` rejects ``T``, ``P``, ``O`` and ``C``;
|
|
``F`` composes.
|
|
``D`` - registered disabled
|
|
The entry is created disabled instead of being matchable at
|
|
once, and has to be enabled by writing ``1`` to its file
|
|
before it dispatches anything. This splits a registration
|
|
into creating the entry and activating it, leaving room to
|
|
configure it in between - which is what a ``B`` entry that
|
|
binds interpreters needs; see the bpf section below. The flag
|
|
is spent on the registration and is not read back: what an
|
|
entry file reports afterwards is whether it is enabled.
|
|
|
|
|
|
There are some restrictions:
|
|
|
|
- the whole register string may not exceed 1920 characters
|
|
- the magic must reside in the first 128 bytes of the file, i.e.
|
|
offset+size(magic) has to be less than 128
|
|
- the interpreter string may not exceed 127 characters
|
|
- an interpreter used with ``C`` or ``L`` but without ``F`` has to be
|
|
named by an absolute path. It is opened when the binary is executed, so
|
|
a relative one would be resolved against the working directory of
|
|
whoever runs the binary
|
|
- the amount of pre-opened interpreters by ``F``, or bound to a ``B`` entry
|
|
is limited by the ``/proc/sys/user/max_binfmt_misc_interpreters`` sysctl. A
|
|
registration past the limit is refused with ``-ENOSPC``. This limits an
|
|
unprivileged namespace pinning files. A nested namespace can raise only its
|
|
own limit and every ancestor is charged too
|
|
|
|
|
|
To use binfmt_misc you have to mount it first. You can mount it with
|
|
``mount -t binfmt_misc none /proc/sys/fs/binfmt_misc`` command, or you can add
|
|
a line ``none /proc/sys/fs/binfmt_misc binfmt_misc defaults 0 0`` to your
|
|
``/etc/fstab`` so it auto mounts on boot.
|
|
|
|
You may want to add the binary formats in one of your ``/etc/rc`` scripts during
|
|
boot-up. Read the manual of your init program to figure out how to do this
|
|
right.
|
|
|
|
Think about the order of adding entries! Later added entries are matched first!
|
|
|
|
|
|
A few examples (assumed you are in ``/proc/sys/fs/binfmt_misc``):
|
|
|
|
- enable support for em86 (like binfmt_em86, for Alpha AXP only)::
|
|
|
|
echo ':i386:M::\x7fELF\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x03:\xff\xff\xff\xff\xff\xfe\xfe\xff\xff\xff\xff\xff\xff\xff\xff\xff\xfb\xff\xff:/bin/em86:' > register
|
|
echo ':i486:M::\x7fELF\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x06:\xff\xff\xff\xff\xff\xfe\xfe\xff\xff\xff\xff\xff\xff\xff\xff\xff\xfb\xff\xff:/bin/em86:' > register
|
|
|
|
- enable support for packed DOS applications (pre-configured dosemu hdimages)::
|
|
|
|
echo ':DEXE:M::\x0eDEX::/usr/bin/dosexec:' > register
|
|
|
|
- enable support for Windows executables using wine::
|
|
|
|
echo ':DOSWin:M::MZ::/usr/local/bin/wine:' > register
|
|
|
|
For java support see Documentation/admin-guide/java.rst
|
|
|
|
|
|
You can enable/disable binfmt_misc or one binary type by echoing 0 (to disable)
|
|
or 1 (to enable) to ``/proc/sys/fs/binfmt_misc/status`` or
|
|
``/proc/.../the_name``.
|
|
Catting the file tells you the current status of ``binfmt_misc/the_entry``.
|
|
|
|
You can remove one entry or all entries by echoing -1 to ``/proc/.../the_name``
|
|
or ``/proc/sys/fs/binfmt_misc/status``. A single entry can also be removed
|
|
by simply unlinking (``rm``) ``/proc/.../the_name``.
|
|
|
|
|
|
bpf-backed handlers
|
|
-------------------
|
|
|
|
With ``CONFIG_BINFMT_MISC_BPF`` both the matching and the interpreter
|
|
selection can be delegated to bpf programs. A handler is an instance of the
|
|
``binfmt_misc_ops`` struct_ops with a ``match`` and a ``load`` program and a
|
|
``name``. Once the struct_ops map is registered the handler can be activated
|
|
with a ``B`` entry that references it by name in the ``interpreter`` field
|
|
and carries neither offset, magic, nor mask::
|
|
|
|
echo ':qemu:B::::my_handler:' > register
|
|
|
|
Both programs receive the ``linux_binprm`` of the binary and both can
|
|
sleep. The ``match`` program decides whether the handler applies: it is
|
|
consulted during the entry walk exactly like magic and extension matching,
|
|
in the same registration order with the same first-match-wins semantics.
|
|
Unlike static matching it is not limited to the prefetched first bytes of
|
|
the file in ``bprm->buf``: it can read the file, e.g. to parse ELF program
|
|
headers whose data sits at arbitrary offsets. It only decides, though: the
|
|
selection kfuncs below are rejected in it. The ``load`` program of the
|
|
matched handler then selects the interpreter: it can equally read the file
|
|
and derive the interpreter from the binary's location. It selects the
|
|
interpreter by calling the ``bpf_binprm_set_interp()`` kfunc with an
|
|
absolute path and returning ``0``. A match is committed: a failing
|
|
``load`` fails the exec with its error instead of falling through to later
|
|
entries; ``-ENOEXEC`` lets the remaining binary formats have a go. A path
|
|
selected this way is opened with the credentials of the task doing the
|
|
exec, exactly as a statically registered interpreter without ``F`` would
|
|
be.
|
|
|
|
An entry can instead bind the interpreters its handler may use, so that no
|
|
path is resolved at exec time at all. An entry registered with ``D`` is not
|
|
matchable yet, which is what leaves it open to being given them, one
|
|
``+name path`` write at a time::
|
|
|
|
echo ':qemu:B::::my_handler:D' > register
|
|
echo '+aarch64 /usr/bin/qemu-aarch64' > qemu
|
|
echo '+arm /usr/bin/qemu-arm' > qemu
|
|
echo 1 > qemu
|
|
|
|
Each path is opened during its write, in the writing process's context and
|
|
with the credentials the entry file was opened with, exactly the way ``F``
|
|
pre-opens a static entry's interpreter; the paths must be absolute. The
|
|
path is everything past the first space, so there is nothing it cannot
|
|
express, and no interpreter has to fit in a register string. An entry
|
|
binds at most 100 interpreters, and each one is charged against
|
|
``max_binfmt_misc_interpreters`` like any other binding. A write past either
|
|
limit is refused with ``-ENOSPC``.
|
|
|
|
The ``load`` program then selects one per exec by name with the
|
|
``bpf_binprm_select_interp()`` kfunc, and every exec runs a clone of the
|
|
file that was opened. The path decides which file is bound and nothing
|
|
else: it is not resolved again, in any namespace, so what it holds later -
|
|
or what it holds in the namespace of whoever runs the binary - no longer
|
|
decides anything.
|
|
|
|
Enabling the entry ends this. Its interpreters are read at exec time with
|
|
nothing but a reference held on the entry, so an entry that has ever been
|
|
matchable can never have its set changed again: the first ``1`` seals it,
|
|
from then on ``+`` is refused with ``-EBUSY``, and an entry registered
|
|
without ``D`` is sealed from the start. Binding a name twice is refused
|
|
with ``-EEXIST``.
|
|
|
|
Selection is by name so that the configuration and the program need not
|
|
agree on an order, and so that a handler is not tied to where a distribution
|
|
puts its interpreters. A name is a single word of printable ASCII, at most
|
|
32 characters; a name the entry did not bind gives the program ``-ENOENT``,
|
|
which it can act on or return. The interpreter runs under the path it was
|
|
registered under, and the entry reports what it bound::
|
|
|
|
$ cat /proc/sys/fs/binfmt_misc/qemu
|
|
enabled
|
|
bpf my_handler
|
|
bpf-interpreter aarch64 /usr/bin/qemu-aarch64
|
|
bpf-interpreter arm /usr/bin/qemu-arm
|
|
flags:
|
|
|
|
The path reported is the one the interpreter was bound under, which named
|
|
the file at that moment; it is not re-resolved, so it is a record of what
|
|
was bound rather than a promise about what that path holds now.
|
|
|
|
The ``load`` program can also pass a single argument to the interpreter with
|
|
the ``bpf_binprm_set_interp_arg()`` kfunc. It is inserted between the
|
|
interpreter and the binary, exactly like the optional argument of a ``#!``
|
|
interpreter line, e.g. for a handler that resolves ``$ORIGIN`` in a script's
|
|
``#!`` path and needs to preserve the argument that followed it.
|
|
|
|
The invocation flags a static entry fixes at registration - ``P``, ``C``,
|
|
``O``, ``T`` and ``L`` - are per-exec choices for a bpf handler, made by the
|
|
``load`` program with the ``bpf_binprm_set_flags()`` kfunc, so a single
|
|
handler can decide them differently for each binary it handles:
|
|
|
|
- ``BPF_BINPRM_PRESERVE_ARGV0`` keeps the caller's ``argv[0]`` (the ``P``
|
|
flag).
|
|
- ``BPF_BINPRM_CREDENTIALS`` computes credentials from the binary (the ``C``
|
|
flag), bounded to user namespaces that map the binary's owner just like
|
|
any other setuid exec.
|
|
- ``BPF_BINPRM_EXECFD`` opens the binary on the interpreter's behalf and
|
|
passes it through the ``AT_EXECFD`` aux vector entry (the ``O`` flag), so
|
|
the interpreter can run binaries it could not open by path.
|
|
- ``BPF_BINPRM_TRANSPARENT`` runs the interpreter transparently (the ``T``
|
|
flag): the binary is handed over through ``AT_EXECFD`` as
|
|
with ``BPF_BINPRM_EXECFD``, but the argument vector is also left as the
|
|
caller passed it. An interpreter that loads the binary from ``AT_EXECFD``
|
|
then appears in ``argv[0]`` and ``/proc/pid/cmdline`` as a direct
|
|
execution of the binary. ``BPF_BINPRM_PRESERVE_ARGV0`` and a staged
|
|
interpreter argument are rejected in combination with it, just as ``P``
|
|
is with ``T``. It also lets a handler
|
|
run a binary passed as an inaccessible ``O_CLOEXEC`` file descriptor to
|
|
``execveat()``, which a path-splicing dispatch cannot: the interpreter
|
|
has no path by which to open it.
|
|
- ``BPF_BINPRM_LOADER`` substitutes the interpreter for the binary's
|
|
``PT_INTERP`` and runs the binary as a fully native exec (the ``L``
|
|
flag). It excludes the other flags and a staged interpreter argument.
|
|
|
|
Because these are program choices, a ``B`` entry carries no invocation
|
|
flags in the register string; ``F`` has none to spell for it either, since
|
|
the interpreters it binds already pre-open what ``F`` would. The
|
|
registration directive ``D`` is the exception: it decides how the entry
|
|
starts out, not how the interpreter is invoked.
|
|
|
|
A handler is looked up only in the user namespace the struct_ops map was
|
|
registered in. Handlers are not inherited, so an entry can only reference a
|
|
handler registered in the same user namespace as its binfmt_misc instance.
|
|
The entry keeps the handler alive; deleting the struct_ops map only prevents
|
|
new activations.
|
|
|
|
|
|
Transparent interpreters
|
|
------------------------
|
|
|
|
With the ``T`` flag or ``BPF_BINPRM_TRANSPARENT`` the dispatch is invisible
|
|
to the resulting process. The argument vector is left exactly as the caller
|
|
built it. The binary is passed through ``AT_EXECFD``. The kernel also labels
|
|
``/proc/pid/exe`` correctly. The binary's file is write-denied while the
|
|
process runs and the interpreter's is not, exactly as if the binary had been
|
|
executed directly. A transparent entry does not change how credentials are
|
|
derived. As
|
|
with any other entry, set*id bits of the binary are only honored with ``C`` (or
|
|
``BPF_BINPRM_CREDENTIALS``).
|
|
|
|
The interpreter has to be built for this contract. The kernel announces it
|
|
with ``AT_FLAGS_TRANSPARENT_INTERP`` in the ``AT_FLAGS`` aux vector entry
|
|
next to ``AT_EXECFD``. The argument vector belongs entirely to the program,
|
|
nothing was spliced in, so the interpreter doesn't consume arguments and
|
|
simply loads the program from the descriptor. The bit is also the loader's
|
|
license to finish the identity. After mapping the program it may retarget the
|
|
``AT_PHDR``/``AT_ENTRY``/``AT_BASE`` entries of ``/proc/pid/auxv`` and the
|
|
code/data statistics markers via one ``PR_SET_MM_MAP`` which completes
|
|
what attaching debuggers observe. What remains visibly different from a
|
|
direct execution is the address space layout. The interpreter occupies
|
|
the main-image position and the program lives in the mmap region.
|
|
|
|
|
|
Loader substitution
|
|
-------------------
|
|
|
|
The ``L`` flag turns the execution model around. Instead of running the
|
|
registered interpreter with the binary as its payload the kernel loads
|
|
the matched binary itself as the main image and substitutes the registered
|
|
interpreter for the loader named in the binary's ``PT_INTERP``.
|
|
|
|
Because the exec is native, there is no dispatch identity to
|
|
reconstruct and no contract the substitute has to implement. A stock
|
|
dynamic loader works unchanged. The argument vector is untouched,
|
|
credentials and ``AT_SECURE`` derive from the binary, there is no
|
|
``AT_EXECFD`` and no marker in the aux vector, the binary sits in the
|
|
main-image slot with the native brk placement so ``/proc/pid/maps``,
|
|
core dumps and perf mmap records have the native shape, and the
|
|
identity is already complete when ``PTRACE_EVENT_EXEC`` stops the
|
|
tracee. So launching under a debugger works, not just attaching. ``L``
|
|
entries are for ELF binaries of a native architecture. Foreign-arch
|
|
emulation and non-ELF payloads remain the domain of the classic and
|
|
transparent modes.
|
|
|
|
The override applies when the format that finally claims the file is
|
|
ELF with a ``PT_INTERP``. A matched binary without one or an
|
|
interpreter-less ``ET_DYN`` drops the override and runs natively. A file
|
|
claimed by another format - a ``#!`` script, say - is handled by that
|
|
format as if the entry had not matched. ``L`` is therefore not an
|
|
enforcement mechanism: it decides how a binary that asks for a loader is
|
|
run, it does not guarantee that everything matching the entry runs under
|
|
the substitute. A format that cannot consume the override at all instead
|
|
refuses the exec with ``ENOEXEC`` before the point of no return.
|
|
|
|
A wrong-architecture ELF fails the whole exec with ``ENOEXEC`` exactly
|
|
as if no entry had matched. A substitute that is not ELF of the right
|
|
architecture fails with ``ELIBBAD``. The usual ``PT_INTERP`` sanity
|
|
checks on the binary still apply. But the segment's content is otherwise
|
|
irrelevant.
|
|
|
|
``L`` rejects the classic-dispatch flags ``T``, ``P``, ``O`` and ``C``
|
|
at registration. ``F`` composes and is valuable: with it the substitute
|
|
is opened at registration time, so later mount namespace or path changes
|
|
cannot redirect it. Without it the substitute is opened when the binary
|
|
is executed, and the path is resolved in the mount namespace and root of
|
|
whoever runs the binary, which is why it has to be absolute. As with
|
|
``C``, register only trusted interpreters. The substituted loader runs
|
|
with credentials derived from the binary.
|
|
|
|
|
|
Hints
|
|
-----
|
|
|
|
If you want to pass special arguments to your interpreter, you can
|
|
write a wrapper script for it.
|
|
See :doc:`Documentation/admin-guide/java.rst <./java>` for an example.
|
|
|
|
Your interpreter should NOT look in the PATH for the filename; the kernel
|
|
passes it the full filename (or the file descriptor) to use. Using ``$PATH`` can
|
|
cause unexpected behaviour and can be a security hazard.
|
|
|
|
|
|
Richard Günther <rguenth@tat.physik.uni-tuebingen.de>
|