요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
===========================================
Seccomp BPF (SECure COMPuting with filters)
===========================================
Introduction
============
A large number of system calls are exposed to every userland process
with many of them going unused for the entire lifetime of the process.
As system calls change and mature, bugs are found and eradicated. A
certain subset of userland applications benefit by having a reduced set
of available system calls. The resulting set reduces the total kernel
surface exposed to the application. System call filtering is meant for
use with those applications.
Seccomp filtering provides a means for a process to specify a filter for
incoming system calls. The filter is expressed as a Berkeley Packet
Filter (BPF) program, as with socket filters, except that the data
operated on is related to the system call being made: system call
number and the system call arguments. This allows for expressive
filtering of system calls using a filter program language with a long
history of being exposed to userland and a straightforward data set.
Additionally, BPF makes it impossible for users of seccomp to fall prey
to time-of-check-time-of-use (TOCTOU) attacks that are common in system
call interposition frameworks. BPF programs may not dereference
pointers which constrains all filters to solely evaluating the system
call arguments directly.
What it isn't
=============
System call filtering isn't a sandbox. It provides a clearly defined
mechanism for minimizing the exposed kernel surface. It is meant to be
a tool for sandbox developers to use. Beyond that, policy for logical
behavior and information flow should be managed with a combination of
other system hardening techniques and, potentially, an LSM of your
choosing. Expressive, dynamic filters provide further options down this
path (avoiding pathological sizes or selecting which of the multiplexed
system calls in socketcall() is allowed, for instance) which could be
construed, incorrectly, as a more complete sandboxing solution.
Usage
=====
An additional seccomp mode is added and is enabled using the same
prctl(2) call as the strict seccomp. If the architecture has
``CONFIG_HAVE_ARCH_SECCOMP_FILTER``, then filters may be added as below:
``PR_SET_SECCOMP``:
Now takes an additional argument which specifies a new filter
using a BPF program.
The BPF program will be executed over struct seccomp_data
reflecting the system call number, arguments, and other
metadata. The BPF program must then return one of the
acceptable values to inform the kernel which action should be
taken.
Usage::
prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);
The 'prog' argument is a pointer to a struct sock_fprog which
will contain the filter program. If the program is invalid, the
call will return -1 and set errno to ``EINVAL``.
If ``fork``/``clone`` and ``execve`` are allowed by @prog, any child
processes will be constrained to the same filters and system
call ABI as the parent.
Prior to use, the task must call ``prctl(PR_SET_NO_NEW_PRIVS, 1)`` or
run with ``CAP_SYS_ADMIN`` privileges in its namespace. If these are not
true, ``-EACCES`` will be returned. This requirement ensures that filter
programs cannot be applied to child processes with greater privileges
than the task that installed them.
Additionally, if ``prctl(2)`` is allowed by the attached filter,
additional filters may be layered on which will increase evaluation
time, but allow for further decreasing the attack surface during
execution of a process.
The above call returns 0 on success and non-zero on error.
Return values
=============
A seccomp filter may return any of the following values. If multiple
filters exist, the return value for the evaluation of a given system
call will always use the highest precedent value. (For example,
``SECCOMP_RET_KILL_PROCESS`` will always take precedence.)
In precedence order, they are:
``SECCOMP_RET_KILL_PROCESS``:
Results in the entire process exiting immediately without executing
the system call. The exit status of the task (``status & 0x7f``)
will be ``SIGSYS``, not ``SIGKILL``.
``SECCOMP_RET_KILL_THREAD``:
Results in the task exiting immediately without executing the
system call. The exit status of the task (``status & 0x7f``) will
be ``SIGSYS``, not ``SIGKILL``.
``SECCOMP_RET_TRAP``:
Results in the kernel sending a ``SIGSYS`` signal to the triggering
task without executing the system call. ``siginfo->si_call_addr``
will show the address of the system call instruction, and
``siginfo->si_syscall`` and ``siginfo->si_arch`` will indicate which
syscall was attempted. The program counter will be as though
the syscall happened (i.e. it will not point to the syscall
instruction). The return value register will contain an arch-
dependent value -- if resuming execution, set it to something
sensible. (The architecture dependency is because replacing
it with ``-ENOSYS`` could overwrite some useful information.)
The ``SECCOMP_RET_DATA`` portion of the return value will be passed
as ``si_errno``.
``SIGSYS`` triggered by seccomp will have a si_code of ``SYS_SECCOMP``.
``SECCOMP_RET_ERRNO``:
Results in the lower 16-bits of the return value being passed
to userland as the errno without executing the system call.
``SECCOMP_RET_USER_NOTIF``:
Results in a ``struct seccomp_notif`` message sent on the userspace
notification fd, if it is attached, or ``-ENOSYS`` if it is not. See
below on discussion of how to handle user notifications.
``SECCOMP_RET_TRACE``:
When returned, this value will cause the kernel to attempt to
notify a ``ptrace()``-based tracer prior to executing the system
call. If there is no tracer present, ``-ENOSYS`` is returned to
userland and the system call is not executed.
A tracer will be notified if it requests ``PTRACE_O_TRACESECCOMP``
using ``ptrace(PTRACE_SETOPTIONS)``. The tracer will be notified
of a ``PTRACE_EVENT_SECCOMP`` and the ``SECCOMP_RET_DATA`` portion of
the BPF program return value will be available to the tracer
via ``PTRACE_GETEVENTMSG``.
The tracer can skip the system call by changing the syscall number
to -1. Alternatively, the tracer can change the system call
requested by changing the system call to a valid syscall number. If
the tracer asks to skip the system call, then the system call will
appear to return the value that the tracer puts in the return value
register.
The seccomp check will not be run again after the tracer is
notified. (This means that seccomp-based sandboxes MUST NOT
allow use of ptrace, even of other sandboxed processes, without
extreme care; ptracers can use this mechanism to escape.)
``SECCOMP_RET_LOG``:
Results in the system call being executed after it is logged. This
should be used by application developers to learn which syscalls their
application needs without having to iterate through multiple test and
development cycles to build the list.
This action will only be logged if "log" is present in the
actions_logged sysctl string.
``SECCOMP_RET_ALLOW``:
Results in the system call being executed.
If multiple filters exist, the return value for the evaluation of a
given system call will always use the highest precedent value.
Precedence is only determined using the ``SECCOMP_RET_ACTION`` mask. When
multiple filters return values of the same precedence, only the
``SECCOMP_RET_DATA`` from the most recently installed filter will be
returned.
Pitfalls
========
The biggest pitfall to avoid during use is filtering on system call
number without checking the architecture value. Why? On any
architecture that supports multiple system call invocation conventions,
the system call numbers may vary based on the specific invocation. If
the numbers in the different calling conventions overlap, then checks in
the filters may be abused. Always check the arch value!
Example
=======
The ``samples/seccomp/`` directory contains both an x86-specific example
and a more generic example of a higher level macro interface for BPF
program generation.
Userspace Notification
======================
The ``SECCOMP_RET_USER_NOTIF`` return code lets seccomp filters pass a
particular syscall to userspace to be handled. This may be useful for
applications like container managers, which wish to intercept particular
syscalls (``mount()``, ``finit_module()``, etc.) and change their behavior.
To acquire a notification FD, use the ``SECCOMP_FILTER_FLAG_NEW_LISTENER``
argument to the ``seccomp()`` syscall:
.. code-block:: c
fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);
which (on success) will return a listener fd for the filter, which can then be
passed around via ``SCM_RIGHTS`` or similar. Note that filter fds correspond to
a particular filter, and not a particular task. So if this task then forks,
notifications from both tasks will appear on the same filter fd. Reads and
writes to/from a filter fd are also synchronized, so a filter fd can safely
have many readers.
The interface for a seccomp notification fd consists of two structures:
.. code-block:: c
struct seccomp_notif_sizes {
__u16 seccomp_notif;
__u16 seccomp_notif_resp;
__u16 seccomp_data;
};
struct seccomp_notif {
__u64 id;
__u32 pid;
__u32 flags;
struct seccomp_data data;
};
struct seccomp_notif_resp {
__u64 id;
__s64 val;
__s32 error;
__u32 flags;
};
The ``struct seccomp_notif_sizes`` structure can be used to determine the size
of the various structures used in seccomp notifications. The size of ``struct
seccomp_data`` may change in the future, so code should use:
.. code-block:: c
struct seccomp_notif_sizes sizes;
seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);
to determine the size of the various structures to allocate. See
samples/seccomp/user-trap.c for an example.
Users can read via ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)`` (or ``poll()``) on a
seccomp notification fd to receive a ``struct seccomp_notif``, which contains
five members: the input length of the structure, a unique-per-filter ``id``,
the ``pid`` of the task which triggered this request (which may be 0 if the
task is in a pid ns not visible from the listener's pid namespace). The
notification also contains the ``data`` passed to seccomp, and a filters flag.
The structure should be zeroed out prior to calling the ioctl.
Userspace can then make a decision based on this information about what to do,
and ``ioctl(SECCOMP_IOCTL_NOTIF_SEND)`` a response, indicating what should be
returned to userspace. The ``id`` member of ``struct seccomp_notif_resp`` should
be the same ``id`` as in ``struct seccomp_notif``.
Userspace can also add file descriptors to the notifying process via
``ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)``. The ``id`` member of
``struct seccomp_notif_addfd`` should be the same ``id`` as in
``struct seccomp_notif``. The ``newfd_flags`` flag may be used to set flags
like O_CLOEXEC on the file descriptor in the notifying process. If the supervisor
wants to inject the file descriptor with a specific number, the
``SECCOMP_ADDFD_FLAG_SETFD`` flag can be used, and set the ``newfd`` member to
the specific number to use. If that file descriptor is already open in the
notifying process it will be replaced. The supervisor can also add an FD, and
respond atomically by using the ``SECCOMP_ADDFD_FLAG_SEND`` flag and the return
value will be the injected file descriptor number.
The notifying process can be preempted, resulting in the notification being
aborted. This can be problematic when trying to take actions on behalf of the
notifying process that are long-running and typically retryable (mounting a
filesystem). Alternatively, at filter installation time, the
``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`` flag can be set. This flag makes it
such that when a user notification is received by the supervisor, the notifying
process will ignore non-fatal signals until the response is sent. Signals that
are sent prior to the notification being received by userspace are handled
normally.
It is worth noting that ``struct seccomp_data`` contains the values of register
arguments to the syscall, but does not contain pointers to memory. The task's
memory is accessible to suitably privileged traces via ``ptrace()`` or
``/proc/pid/mem``. However, care should be taken to avoid the TOCTOU mentioned
above in this document: all arguments being read from the tracee's memory
should be read into the tracer's memory before any policy decisions are made.
This allows for an atomic decision on syscall arguments.
Sysctls
=======
Seccomp's sysctl files can be found in the ``/proc/sys/kernel/seccomp/``
directory. Here's a description of each file in that directory:
``actions_avail``:
A read-only ordered list of seccomp return values (refer to the
``SECCOMP_RET_*`` macros above) in string form. The ordering, from
left-to-right, is the least permissive return value to the most
permissive return value.
The list represents the set of seccomp return values supported
by the kernel. A userspace program may use this list to
determine if the actions found in the ``seccomp.h``, when the
program was built, differs from the set of actions actually
supported in the current running kernel.
``actions_logged``:
A read-write ordered list of seccomp return values (refer to the
``SECCOMP_RET_*`` macros above) that are allowed to be logged. Writes
to the file do not need to be in ordered form but reads from the file
will be ordered in the same way as the actions_avail sysctl.
The ``allow`` string is not accepted in the ``actions_logged`` sysctl
as it is not possible to log ``SECCOMP_RET_ALLOW`` actions. Attempting
to write ``allow`` to the sysctl will result in an EINVAL being
returned.
Adding architecture support
===========================
See ``arch/Kconfig`` for the authoritative requirements. In general, if an
architecture supports both ptrace_event and seccomp, it will be able to
support seccomp filter with minor fixup: ``SIGSYS`` support and seccomp return
value checking. Then it must just add ``CONFIG_HAVE_ARCH_SECCOMP_FILTER``
to its arch-specific Kconfig.
Caveats
=======
The vDSO can cause some system calls to run entirely in userspace,
leading to surprises when you run programs on different machines that
fall back to real syscalls. To minimize these surprises on x86, make
sure you test with
``/sys/devices/system/clocksource/clocksource0/current_clocksource`` set to
something like ``acpi_pm``.
On x86-64, vsyscall emulation is enabled by default. (vsyscalls are
legacy variants on vDSO calls.) Currently, emulated vsyscalls will
honor seccomp, with a few oddities:
- A return value of ``SECCOMP_RET_TRAP`` will set a ``si_call_addr`` pointing to
the vsyscall entry for the given call and not the address after the
'syscall' instruction. Any code which wants to restart the call
should be aware that (a) a ret instruction has been emulated and (b)
trying to resume the syscall will again trigger the standard vsyscall
emulation security checks, making resuming the syscall mostly
pointless.
- A return value of ``SECCOMP_RET_TRACE`` will signal the tracer as usual,
but the syscall may not be changed to another system call using the
orig_rax register. It may only be changed to -1 order to skip the
currently emulated call. Any other change MAY terminate the process.
The rip value seen by the tracer will be the syscall entry address;
this is different from normal behavior. The tracer MUST NOT modify
rip or rsp. (Do not rely on other changes terminating the process.
They might work. For example, on some kernels, choosing a syscall
that only exists in future kernels will be correctly emulated (by
returning ``-ENOSYS``).
To detect this quirky behavior, check for ``addr & ~0x0C00 ==
0xFFFFFFFFFF600000``. (For ``SECCOMP_RET_TRACE``, use rip. For
``SECCOMP_RET_TRAP``, use ``siginfo->si_call_addr``.) Do not check any other
condition: future kernels may improve vsyscall emulation and current
kernels in vsyscall=native mode will behave differently, but the
instructions at ``0xF...F600{0,4,8,C}00`` will not be system calls in these
cases.
Note that modern systems are unlikely to use vsyscalls at all -- they
are a legacy feature and they are considerably slower than standard
syscalls. New code will use the vDSO, and vDSO-issued system calls
are indistinguishable from normal system calls.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
소개
1-29모든 사용자 공간 프로세스에는 많은 system call이 노출되지만, 상당수는 프로세스 수명 동안 한 번도 쓰이지 않습니다. system call이 변하고 성숙하는 과정에서 버그가 발견되고 제거되므로 일부 애플리케이션은 사용할 수 있는 system call 집합을 줄여 커널 공격 표면을 축소하는 것이 유리합니다. system call filtering은 이런 애플리케이션을 위한 기능입니다.
seccomp filtering을 사용하면 프로세스가 들어오는 system call에 적용할 필터를 지정할 수 있습니다. 필터는 socket filter와 같은 Berkeley Packet Filter, 즉 BPF 프로그램이지만 입력 데이터는 호출 번호, 인자와 관련 메타데이터입니다. 오래 검증된 사용자 공간 필터 언어와 단순한 입력 집합으로 표현력 있는 system call 정책을 작성할 수 있습니다.
BPF 프로그램은 포인터를 역참조할 수 없고 system call 인자 값 자체만 평가합니다. 이 제약 덕분에 seccomp 사용자는 일반적인 system call interposition 프레임워크의 time-of-check-time-of-use, 즉 TOCTOU 공격을 피할 수 있습니다.
===========================================
Seccomp BPF (SECure COMPuting with filters)
===========================================
Introduction
============
A large number of system calls are exposed to every userland process
with many of them going unused for the entire lifetime of the process.
As system calls change and mature, bugs are found and eradicated. A
certain subset of userland applications benefit by having a reduced set
of available system calls. The resulting set reduces the total kernel
surface exposed to the application. System call filtering is meant for
use with those applications.
Seccomp filtering provides a means for a process to specify a filter for
incoming system calls. The filter is expressed as a Berkeley Packet
Filter (BPF) program, as with socket filters, except that the data
operated on is related to the system call being made: system call
number and the system call arguments. This allows for expressive
filtering of system calls using a filter program language with a long
history of being exposed to userland and a straightforward data set.
Additionally, BPF makes it impossible for users of seccomp to fall prey
to time-of-check-time-of-use (TOCTOU) attacks that are common in system
call interposition frameworks. BPF programs may not dereference
pointers which constrains all filters to solely evaluating the system
call arguments directly.
seccomp가 아닌 것
30-42system call filtering 자체는 sandbox가 아닙니다. 노출된 커널 표면을 최소화하는 명확한 메커니즘이며 sandbox 개발자가 사용하는 구성 요소입니다. 논리적 동작과 정보 흐름 정책은 다른 시스템 강화 기법, 필요하다면 선택한 LSM과 함께 관리해야 합니다.
동적이고 표현력 있는 필터로 비정상적으로 큰 정책을 피하거나 `socketcall()`이 multiplex하는 호출 중 일부만 허용할 수 있지만, 이를 완전한 sandbox 해법으로 오해해서는 안 됩니다.
What it isn't
=============
System call filtering isn't a sandbox. It provides a clearly defined
mechanism for minimizing the exposed kernel surface. It is meant to be
a tool for sandbox developers to use. Beyond that, policy for logical
behavior and information flow should be managed with a combination of
other system hardening techniques and, potentially, an LSM of your
choosing. Expressive, dynamic filters provide further options down this
path (avoiding pathological sizes or selecting which of the multiplexed
system calls in socketcall() is allowed, for instance) which could be
construed, incorrectly, as a more complete sandboxing solution.
필터 설치와 상속
43-83아키텍처가 `CONFIG_HAVE_ARCH_SECCOMP_FILTER`를 제공하면 strict seccomp와 같은 `prctl(2)` 호출에 추가 모드를 지정해 필터를 설치할 수 있습니다.
`PR_SET_SECCOMP`는 BPF 프로그램으로 새 필터를 지정하는 인자를 추가로 받습니다. 프로그램은 system call 번호, 인자와 메타데이터를 담은 `struct seccomp_data`를 평가하고 커널이 수행할 action 값을 반환해야 합니다.
prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);
`prog`는 필터 프로그램을 담은 `struct sock_fprog` 포인터입니다. 프로그램이 잘못되면 호출은 -1을 반환하고 `errno`를 `EINVAL`로 설정합니다.
필터가 `fork`/`clone`과 `execve`를 허용하면 자식 프로세스도 부모와 같은 필터 및 system call ABI 제약을 받습니다.
설치 전에 태스크가 `prctl(PR_SET_NO_NEW_PRIVS, 1)`을 호출했거나 자신의 namespace에서 `CAP_SYS_ADMIN` 권한을 가져야 합니다. 아니면 `-EACCES`를 반환합니다. 이 조건은 설치자보다 권한이 큰 자식에게 필터 프로그램을 적용해 권한을 악용하는 일을 막습니다.
연결된 필터가 `prctl(2)`을 허용하면 필터를 추가로 겹쳐 적용할 수 있습니다. 평가 시간은 늘지만 실행 중 공격 표면을 더 줄일 수 있습니다. 설치 호출은 성공 시 0, 오류 시 0이 아닌 값을 반환합니다.
Usage
=====
An additional seccomp mode is added and is enabled using the same
prctl(2) call as the strict seccomp. If the architecture has
``CONFIG_HAVE_ARCH_SECCOMP_FILTER``, then filters may be added as below:
``PR_SET_SECCOMP``:
Now takes an additional argument which specifies a new filter
using a BPF program.
The BPF program will be executed over struct seccomp_data
reflecting the system call number, arguments, and other
metadata. The BPF program must then return one of the
acceptable values to inform the kernel which action should be
taken.
Usage::
prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);
The 'prog' argument is a pointer to a struct sock_fprog which
will contain the filter program. If the program is invalid, the
call will return -1 and set errno to ``EINVAL``.
If ``fork``/``clone`` and ``execve`` are allowed by @prog, any child
processes will be constrained to the same filters and system
call ABI as the parent.
Prior to use, the task must call ``prctl(PR_SET_NO_NEW_PRIVS, 1)`` or
run with ``CAP_SYS_ADMIN`` privileges in its namespace. If these are not
true, ``-EACCES`` will be returned. This requirement ensures that filter
programs cannot be applied to child processes with greater privileges
than the task that installed them.
Additionally, if ``prctl(2)`` is allowed by the attached filter,
additional filters may be layered on which will increase evaluation
time, but allow for further decreasing the attack surface during
execution of a process.
The above call returns 0 on success and non-zero on error.
반환값: 종료와 trap
84-120여러 필터가 있으면 주어진 system call 평가에서 가장 우선순위가 높은 action이 항상 선택됩니다. 예를 들어 `SECCOMP_RET_KILL_PROCESS`가 언제나 우선합니다.
위에서 아래 순서로 우선순위가 낮아집니다.
TRAP의 `siginfo->si_call_addr`는 system call 명령 주소를, `si_syscall`과 `si_arch`는 시도한 호출을 나타냅니다. program counter는 호출이 일어난 뒤처럼 보이며 syscall 명령 자체를 가리키지 않습니다. 반환값 register는 아키텍처 의존 값이므로 실행을 재개하려면 의미 있는 값으로 설정해야 합니다.
TRAP 반환값의 `SECCOMP_RET_DATA` 부분은 `si_errno`로 전달되고, seccomp가 발생시킨 `SIGSYS`의 `si_code`는 `SYS_SECCOMP`입니다.
Return values
=============
A seccomp filter may return any of the following values. If multiple
filters exist, the return value for the evaluation of a given system
call will always use the highest precedent value. (For example,
``SECCOMP_RET_KILL_PROCESS`` will always take precedence.)
In precedence order, they are:
``SECCOMP_RET_KILL_PROCESS``:
Results in the entire process exiting immediately without executing
the system call. The exit status of the task (``status & 0x7f``)
will be ``SIGSYS``, not ``SIGKILL``.
``SECCOMP_RET_KILL_THREAD``:
Results in the task exiting immediately without executing the
system call. The exit status of the task (``status & 0x7f``) will
be ``SIGSYS``, not ``SIGKILL``.
``SECCOMP_RET_TRAP``:
Results in the kernel sending a ``SIGSYS`` signal to the triggering
task without executing the system call. ``siginfo->si_call_addr``
will show the address of the system call instruction, and
``siginfo->si_syscall`` and ``siginfo->si_arch`` will indicate which
syscall was attempted. The program counter will be as though
the syscall happened (i.e. it will not point to the syscall
instruction). The return value register will contain an arch-
dependent value -- if resuming execution, set it to something
sensible. (The architecture dependency is because replacing
it with ``-ENOSYS`` could overwrite some useful information.)
The ``SECCOMP_RET_DATA`` portion of the return value will be passed
as ``si_errno``.
``SIGSYS`` triggered by seccomp will have a si_code of ``SYS_SECCOMP``.
반환값: errno, 알림, trace, log, allow
121-173각 action은 호출 실행 여부와 관찰 주체가 다릅니다.
`SECCOMP_RET_LOG`는 애플리케이션 개발자가 여러 차례의 시험·개발 주기를 반복하지 않고 실제로 필요한 syscall 목록을 학습하도록 마련된 action입니다. 호출은 차단되지 않으며 기록 여부는 `actions_logged` 설정에도 좌우됩니다.
TRACE에서 `PTRACE_O_TRACESECCOMP`를 `ptrace(PTRACE_SETOPTIONS)`로 요청한 tracer는 `PTRACE_EVENT_SECCOMP` 통지와 `PTRACE_GETEVENTMSG`를 통한 `SECCOMP_RET_DATA`를 받습니다.
tracer는 syscall 번호를 -1로 바꿔 호출을 건너뛰거나 유효한 번호로 바꿀 수 있습니다. 건너뛰면 tracer가 반환값 register에 넣은 값을 system call 결과로 보게 됩니다.
tracer 통지 뒤 seccomp 검사를 다시 하지 않습니다. 따라서 seccomp sandbox가 ptrace를 허용하면 tracer가 이 경로로 탈출할 수 있으므로 극도로 주의해야 합니다.
여러 필터 결과의 우선순위는 `SECCOMP_RET_ACTION` mask만으로 정합니다. 같은 우선순위라면 가장 최근에 설치한 필터의 `SECCOMP_RET_DATA`만 반환합니다.
``SECCOMP_RET_ERRNO``:
Results in the lower 16-bits of the return value being passed
to userland as the errno without executing the system call.
``SECCOMP_RET_USER_NOTIF``:
Results in a ``struct seccomp_notif`` message sent on the userspace
notification fd, if it is attached, or ``-ENOSYS`` if it is not. See
below on discussion of how to handle user notifications.
``SECCOMP_RET_TRACE``:
When returned, this value will cause the kernel to attempt to
notify a ``ptrace()``-based tracer prior to executing the system
call. If there is no tracer present, ``-ENOSYS`` is returned to
userland and the system call is not executed.
A tracer will be notified if it requests ``PTRACE_O_TRACESECCOMP``
using ``ptrace(PTRACE_SETOPTIONS)``. The tracer will be notified
of a ``PTRACE_EVENT_SECCOMP`` and the ``SECCOMP_RET_DATA`` portion of
the BPF program return value will be available to the tracer
via ``PTRACE_GETEVENTMSG``.
The tracer can skip the system call by changing the syscall number
to -1. Alternatively, the tracer can change the system call
requested by changing the system call to a valid syscall number. If
the tracer asks to skip the system call, then the system call will
appear to return the value that the tracer puts in the return value
register.
The seccomp check will not be run again after the tracer is
notified. (This means that seccomp-based sandboxes MUST NOT
allow use of ptrace, even of other sandboxed processes, without
extreme care; ptracers can use this mechanism to escape.)
``SECCOMP_RET_LOG``:
Results in the system call being executed after it is logged. This
should be used by application developers to learn which syscalls their
application needs without having to iterate through multiple test and
development cycles to build the list.
This action will only be logged if "log" is present in the
actions_logged sysctl string.
``SECCOMP_RET_ALLOW``:
Results in the system call being executed.
If multiple filters exist, the return value for the evaluation of a
given system call will always use the highest precedent value.
Precedence is only determined using the ``SECCOMP_RET_ACTION`` mask. When
multiple filters return values of the same precedence, only the
``SECCOMP_RET_DATA`` from the most recently installed filter will be
returned.
아키텍처 확인과 예제
174-190가장 큰 함정은 architecture 값을 확인하지 않고 system call 번호만 필터링하는 것입니다. 여러 호출 규약을 지원하는 아키텍처에서는 규약에 따라 번호가 다르고 서로 겹칠 수 있어 필터 검사가 악용될 수 있습니다. 항상 arch 값을 검사해야 합니다.
`samples/seccomp/` 디렉터리에는 x86 전용 예제와 BPF 프로그램 생성을 위한 상위 수준 macro 인터페이스의 범용 예제가 있습니다.
Pitfalls
========
The biggest pitfall to avoid during use is filtering on system call
number without checking the architecture value. Why? On any
architecture that supports multiple system call invocation conventions,
the system call numbers may vary based on the specific invocation. If
the numbers in the different calling conventions overlap, then checks in
the filters may be abused. Always check the arch value!
Example
=======
The ``samples/seccomp/`` directory contains both an x86-specific example
and a more generic example of a higher level macro interface for BPF
program generation.
사용자 공간 알림 설정과 구조체
191-248`SECCOMP_RET_USER_NOTIF`는 특정 system call을 사용자 공간으로 넘겨 처리하게 합니다. container manager가 `mount()`나 `finit_module()` 같은 호출을 가로채 동작을 바꾸는 경우에 유용합니다.
notification FD를 얻으려면 `seccomp()` 호출에 `SECCOMP_FILTER_FLAG_NEW_LISTENER`를 지정합니다.
fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);
성공하면 listener FD를 반환하며 `SCM_RIGHTS` 등으로 전달할 수 있습니다. filter FD는 특정 태스크가 아니라 특정 필터에 대응하므로 설치 태스크가 fork하면 두 태스크의 알림이 같은 FD에 나타납니다. FD read/write는 동기화되어 여러 reader가 안전하게 공유할 수 있습니다.
seccomp notification FD 인터페이스는 크기 정보와 요청·응답 구조체를 사용합니다.
struct seccomp_notif_sizes {
__u16 seccomp_notif;
__u16 seccomp_notif_resp;
__u16 seccomp_data;
};
struct seccomp_notif {
__u64 id;
__u32 pid;
__u32 flags;
struct seccomp_data data;
};
struct seccomp_notif_resp {
__u64 id;
__s64 val;
__s32 error;
__u32 flags;
};
`struct seccomp_data` 크기는 미래에 바뀔 수 있으므로 `SECCOMP_GET_NOTIF_SIZES`로 각 구조체의 실제 크기를 조회해 할당해야 합니다.
struct seccomp_notif_sizes sizes;
seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);
구현 예제는 `samples/seccomp/user-trap.c`에서 볼 수 있습니다.
Userspace Notification
======================
The ``SECCOMP_RET_USER_NOTIF`` return code lets seccomp filters pass a
particular syscall to userspace to be handled. This may be useful for
applications like container managers, which wish to intercept particular
syscalls (``mount()``, ``finit_module()``, etc.) and change their behavior.
To acquire a notification FD, use the ``SECCOMP_FILTER_FLAG_NEW_LISTENER``
argument to the ``seccomp()`` syscall:
.. code-block:: c
fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);
which (on success) will return a listener fd for the filter, which can then be
passed around via ``SCM_RIGHTS`` or similar. Note that filter fds correspond to
a particular filter, and not a particular task. So if this task then forks,
notifications from both tasks will appear on the same filter fd. Reads and
writes to/from a filter fd are also synchronized, so a filter fd can safely
have many readers.
The interface for a seccomp notification fd consists of two structures:
.. code-block:: c
struct seccomp_notif_sizes {
__u16 seccomp_notif;
__u16 seccomp_notif_resp;
__u16 seccomp_data;
};
struct seccomp_notif {
__u64 id;
__u32 pid;
__u32 flags;
struct seccomp_data data;
};
struct seccomp_notif_resp {
__u64 id;
__s64 val;
__s32 error;
__u32 flags;
};
The ``struct seccomp_notif_sizes`` structure can be used to determine the size
of the various structures used in seccomp notifications. The size of ``struct
seccomp_data`` may change in the future, so code should use:
.. code-block:: c
struct seccomp_notif_sizes sizes;
seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);
to determine the size of the various structures to allocate. See
samples/seccomp/user-trap.c for an example.
알림 수신, 응답과 FD 주입
249-291사용자 공간은 notification FD에서 `ioctl(SECCOMP_IOCTL_NOTIF_RECV)` 또는 `poll()`로 `struct seccomp_notif`를 받습니다. 알림에는 입력 구조체 길이, 필터별 고유 `id`, 요청을 발생시킨 태스크의 `pid`, seccomp `data`, filter flag가 들어갑니다. listener의 PID namespace에서 태스크가 보이지 않으면 pid는 0일 수 있습니다. ioctl 전 구조체를 0으로 초기화해야 합니다.
supervisor는 정보를 바탕으로 결정을 내린 뒤 `ioctl(SECCOMP_IOCTL_NOTIF_SEND)`로 사용자 공간에 반환할 결과를 보냅니다. `struct seccomp_notif_resp.id`는 요청의 `struct seccomp_notif.id`와 같아야 합니다.
`ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)`로 알림을 발생시킨 프로세스에 file descriptor를 추가할 수도 있습니다. `struct seccomp_notif_addfd.id`도 요청 id와 같아야 하고, `newfd_flags`로 O_CLOEXEC 같은 flag를 설정합니다.
특정 번호로 FD를 주입하려면 `SECCOMP_ADDFD_FLAG_SETFD`와 `newfd`를 사용합니다. 대상 번호가 열려 있으면 교체됩니다. `SECCOMP_ADDFD_FLAG_SEND`를 사용하면 FD 추가와 응답을 원자적으로 수행하고 주입된 FD 번호를 반환받습니다.
알림 프로세스가 preempt되면 notification이 중단될 수 있어 filesystem mount처럼 오래 걸리고 보통 재시도 가능한 작업에 문제가 됩니다. 필터 설치 때 `SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`를 설정하면 supervisor가 알림을 받은 뒤 응답을 보낼 때까지 알림 프로세스가 fatal이 아닌 signal을 무시합니다. 사용자 공간 수신 전의 signal은 정상 처리됩니다.
`struct seccomp_data`에는 syscall register 인자 값이 있지만 메모리 포인터가 가리키는 내용은 없습니다. 충분한 권한의 tracer는 `ptrace()`나 `/proc/pid/mem`으로 태스크 메모리를 읽을 수 있지만 TOCTOU를 피하려면 정책 결정을 내리기 전에 tracee 메모리의 모든 인자를 tracer 메모리로 먼저 복사해야 합니다. 그래야 syscall 인자에 대한 원자적 결정을 내릴 수 있습니다.
Users can read via ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)`` (or ``poll()``) on a
seccomp notification fd to receive a ``struct seccomp_notif``, which contains
five members: the input length of the structure, a unique-per-filter ``id``,
the ``pid`` of the task which triggered this request (which may be 0 if the
task is in a pid ns not visible from the listener's pid namespace). The
notification also contains the ``data`` passed to seccomp, and a filters flag.
The structure should be zeroed out prior to calling the ioctl.
Userspace can then make a decision based on this information about what to do,
and ``ioctl(SECCOMP_IOCTL_NOTIF_SEND)`` a response, indicating what should be
returned to userspace. The ``id`` member of ``struct seccomp_notif_resp`` should
be the same ``id`` as in ``struct seccomp_notif``.
Userspace can also add file descriptors to the notifying process via
``ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)``. The ``id`` member of
``struct seccomp_notif_addfd`` should be the same ``id`` as in
``struct seccomp_notif``. The ``newfd_flags`` flag may be used to set flags
like O_CLOEXEC on the file descriptor in the notifying process. If the supervisor
wants to inject the file descriptor with a specific number, the
``SECCOMP_ADDFD_FLAG_SETFD`` flag can be used, and set the ``newfd`` member to
the specific number to use. If that file descriptor is already open in the
notifying process it will be replaced. The supervisor can also add an FD, and
respond atomically by using the ``SECCOMP_ADDFD_FLAG_SEND`` flag and the return
value will be the injected file descriptor number.
The notifying process can be preempted, resulting in the notification being
aborted. This can be problematic when trying to take actions on behalf of the
notifying process that are long-running and typically retryable (mounting a
filesystem). Alternatively, at filter installation time, the
``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`` flag can be set. This flag makes it
such that when a user notification is received by the supervisor, the notifying
process will ignore non-fatal signals until the response is sent. Signals that
are sent prior to the notification being received by userspace are handled
normally.
It is worth noting that ``struct seccomp_data`` contains the values of register
arguments to the syscall, but does not contain pointers to memory. The task's
memory is accessible to suitably privileged traces via ``ptrace()`` or
``/proc/pid/mem``. However, care should be taken to avoid the TOCTOU mentioned
above in this document: all arguments being read from the tracee's memory
should be read into the tracer's memory before any policy decisions are made.
This allows for an atomic decision on syscall arguments.
Seccomp sysctl
292-320seccomp sysctl은 `/proc/sys/kernel/seccomp/`에 있습니다.
현재 커널이 지원하고 기록할 수 있는 action 집합을 제공합니다.
`SECCOMP_RET_ALLOW` action은 기록할 수 없으므로 `actions_logged`는 `allow` 문자열을 받지 않으며 쓰려고 하면 `EINVAL`을 반환합니다.
Sysctls
=======
Seccomp's sysctl files can be found in the ``/proc/sys/kernel/seccomp/``
directory. Here's a description of each file in that directory:
``actions_avail``:
A read-only ordered list of seccomp return values (refer to the
``SECCOMP_RET_*`` macros above) in string form. The ordering, from
left-to-right, is the least permissive return value to the most
permissive return value.
The list represents the set of seccomp return values supported
by the kernel. A userspace program may use this list to
determine if the actions found in the ``seccomp.h``, when the
program was built, differs from the set of actions actually
supported in the current running kernel.
``actions_logged``:
A read-write ordered list of seccomp return values (refer to the
``SECCOMP_RET_*`` macros above) that are allowed to be logged. Writes
to the file do not need to be in ordered form but reads from the file
will be ordered in the same way as the actions_avail sysctl.
The ``allow`` string is not accepted in the ``actions_logged`` sysctl
as it is not possible to log ``SECCOMP_RET_ALLOW`` actions. Attempting
to write ``allow`` to the sysctl will result in an EINVAL being
returned.
아키텍처 지원 추가
321-331권위 있는 요구사항은 `arch/Kconfig`에 있습니다. 일반적으로 ptrace_event와 seccomp를 모두 지원하는 아키텍처는 `SIGSYS` 지원과 seccomp 반환값 검사를 약간 보완하면 필터를 지원할 수 있습니다. 그 뒤 아키텍처별 Kconfig에 `CONFIG_HAVE_ARCH_SECCOMP_FILTER`를 추가합니다.
Adding architecture support
===========================
See ``arch/Kconfig`` for the authoritative requirements. In general, if an
architecture supports both ptrace_event and seccomp, it will be able to
support seccomp filter with minor fixup: ``SIGSYS`` support and seccomp return
value checking. Then it must just add ``CONFIG_HAVE_ARCH_SECCOMP_FILTER``
to its arch-specific Kconfig.
vDSO와 vsyscall 주의사항
332-376vDSO는 일부 system call을 전적으로 사용자 공간에서 실행할 수 있습니다. 다른 시스템에서 실제 syscall로 fallback할 때 동작이 달라질 수 있습니다. x86에서 이를 시험하려면 `/sys/devices/system/clocksource/clocksource0/current_clocksource`를 `acpi_pm` 같은 값으로 설정하십시오.
x86-64는 기본으로 legacy vDSO 변형인 vsyscall emulation을 사용합니다. emulated vsyscall도 seccomp를 따르지만 예외가 있습니다.
일반 syscall 재시작이나 ptrace 처리와 같은 방식으로 다루면 안 됩니다.
이 동작은 `addr & ~0x0C00 == 0xFFFFFFFFFF600000`으로 감지합니다. TRACE에는 rip, TRAP에는 `siginfo->si_call_addr`를 사용합니다. 미래 커널이나 `vsyscall=native`가 다르게 동작할 수 있으므로 다른 조건은 검사하지 마십시오.
현대 시스템은 legacy이고 일반 syscall보다 훨씬 느린 vsyscall을 거의 쓰지 않습니다. 새 코드는 vDSO를 사용하며 vDSO가 실제로 발행한 system call은 일반 system call과 구별되지 않습니다.
Caveats
=======
The vDSO can cause some system calls to run entirely in userspace,
leading to surprises when you run programs on different machines that
fall back to real syscalls. To minimize these surprises on x86, make
sure you test with
``/sys/devices/system/clocksource/clocksource0/current_clocksource`` set to
something like ``acpi_pm``.
On x86-64, vsyscall emulation is enabled by default. (vsyscalls are
legacy variants on vDSO calls.) Currently, emulated vsyscalls will
honor seccomp, with a few oddities:
- A return value of ``SECCOMP_RET_TRAP`` will set a ``si_call_addr`` pointing to
the vsyscall entry for the given call and not the address after the
'syscall' instruction. Any code which wants to restart the call
should be aware that (a) a ret instruction has been emulated and (b)
trying to resume the syscall will again trigger the standard vsyscall
emulation security checks, making resuming the syscall mostly
pointless.
- A return value of ``SECCOMP_RET_TRACE`` will signal the tracer as usual,
but the syscall may not be changed to another system call using the
orig_rax register. It may only be changed to -1 order to skip the
currently emulated call. Any other change MAY terminate the process.
The rip value seen by the tracer will be the syscall entry address;
this is different from normal behavior. The tracer MUST NOT modify
rip or rsp. (Do not rely on other changes terminating the process.
They might work. For example, on some kernels, choosing a syscall
that only exists in future kernels will be correctly emulated (by
returning ``-ENOSYS``).
To detect this quirky behavior, check for ``addr & ~0x0C00 ==
0xFFFFFFFFFF600000``. (For ``SECCOMP_RET_TRACE``, use rip. For
``SECCOMP_RET_TRAP``, use ``siginfo->si_call_addr``.) Do not check any other
condition: future kernels may improve vsyscall emulation and current
kernels in vsyscall=native mode will behave differently, but the
instructions at ``0xF...F600{0,4,8,C}00`` will not be system calls in these
cases.
Note that modern systems are unlikely to use vsyscalls at all -- they
are a legacy feature and they are considerably slower than standard
syscalls. New code will use the vDSO, and vDSO-issued system calls
are indistinguishable from normal system calls.
요약·해설
seccomp_filter.rst:1-376seccomp는 sandbox 전체가 아니라 syscall 공격 표면을 줄이는 도구입니다. 안전한 정책은 arch 검사, `no_new_privs`, action 우선순위, USER_NOTIF 메모리 복사 순서와 ptrace 탈출 가능성을 함께 다뤄야 합니다.