← Documents Documentation/userspace-api/seccomp_filter.rst GitHub 원문 ↗

Linux 6.18.37 · 사용자 공간 API

Seccomp BPF 필터

Seccomp BPF 필터 설치, action 우선순위, 사용자 공간 알림과 아키텍처 주의사항을 설명합니다.

Source pathDocumentation/userspace-api/seccomp_filter.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

seccomp_filter.rst:1-376

seccomp는 sandbox 전체가 아니라 syscall 공격 표면을 줄이는 도구입니다. 안전한 정책은 arch 검사, `no_new_privs`, action 우선순위, USER_NOTIF 메모리 복사 순서와 ptrace 탈출 가능성을 함께 다뤄야 합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 ===========================================
2 Seccomp BPF (SECure COMPuting with filters)
3 ===========================================
4
5 Introduction
6 ============
7
8 A large number of system calls are exposed to every userland process
9 with many of them going unused for the entire lifetime of the process.
10 As system calls change and mature, bugs are found and eradicated. A
11 certain subset of userland applications benefit by having a reduced set
12 of available system calls. The resulting set reduces the total kernel
13 surface exposed to the application. System call filtering is meant for
14 use with those applications.
15
16 Seccomp filtering provides a means for a process to specify a filter for
17 incoming system calls. The filter is expressed as a Berkeley Packet
18 Filter (BPF) program, as with socket filters, except that the data
19 operated on is related to the system call being made: system call
20 number and the system call arguments. This allows for expressive
21 filtering of system calls using a filter program language with a long
22 history of being exposed to userland and a straightforward data set.
23
24 Additionally, BPF makes it impossible for users of seccomp to fall prey
25 to time-of-check-time-of-use (TOCTOU) attacks that are common in system
26 call interposition frameworks. BPF programs may not dereference
27 pointers which constrains all filters to solely evaluating the system
28 call arguments directly.
29
30 What it isn't
31 =============
32
33 System call filtering isn't a sandbox. It provides a clearly defined
34 mechanism for minimizing the exposed kernel surface. It is meant to be
35 a tool for sandbox developers to use. Beyond that, policy for logical
36 behavior and information flow should be managed with a combination of
37 other system hardening techniques and, potentially, an LSM of your
38 choosing. Expressive, dynamic filters provide further options down this
39 path (avoiding pathological sizes or selecting which of the multiplexed
40 system calls in socketcall() is allowed, for instance) which could be
41 construed, incorrectly, as a more complete sandboxing solution.
42
43 Usage
44 =====
45
46 An additional seccomp mode is added and is enabled using the same
47 prctl(2) call as the strict seccomp. If the architecture has
48 ``CONFIG_HAVE_ARCH_SECCOMP_FILTER``, then filters may be added as below:
49
50 ``PR_SET_SECCOMP``:
51 Now takes an additional argument which specifies a new filter
52 using a BPF program.
53 The BPF program will be executed over struct seccomp_data
54 reflecting the system call number, arguments, and other
55 metadata. The BPF program must then return one of the
56 acceptable values to inform the kernel which action should be
57 taken.
58
59 Usage::
60
61 prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);
62
63 The 'prog' argument is a pointer to a struct sock_fprog which
64 will contain the filter program. If the program is invalid, the
65 call will return -1 and set errno to ``EINVAL``.
66
67 If ``fork``/``clone`` and ``execve`` are allowed by @prog, any child
68 processes will be constrained to the same filters and system
69 call ABI as the parent.
70
71 Prior to use, the task must call ``prctl(PR_SET_NO_NEW_PRIVS, 1)`` or
72 run with ``CAP_SYS_ADMIN`` privileges in its namespace. If these are not
73 true, ``-EACCES`` will be returned. This requirement ensures that filter
74 programs cannot be applied to child processes with greater privileges
75 than the task that installed them.
76
77 Additionally, if ``prctl(2)`` is allowed by the attached filter,
78 additional filters may be layered on which will increase evaluation
79 time, but allow for further decreasing the attack surface during
80 execution of a process.
81
82 The above call returns 0 on success and non-zero on error.
83
84 Return values
85 =============
86
87 A seccomp filter may return any of the following values. If multiple
88 filters exist, the return value for the evaluation of a given system
89 call will always use the highest precedent value. (For example,
90 ``SECCOMP_RET_KILL_PROCESS`` will always take precedence.)
91
92 In precedence order, they are:
93
94 ``SECCOMP_RET_KILL_PROCESS``:
95 Results in the entire process exiting immediately without executing
96 the system call. The exit status of the task (``status & 0x7f``)
97 will be ``SIGSYS``, not ``SIGKILL``.
98
99 ``SECCOMP_RET_KILL_THREAD``:
100 Results in the task exiting immediately without executing the
101 system call. The exit status of the task (``status & 0x7f``) will
102 be ``SIGSYS``, not ``SIGKILL``.
103
104 ``SECCOMP_RET_TRAP``:
105 Results in the kernel sending a ``SIGSYS`` signal to the triggering
106 task without executing the system call. ``siginfo->si_call_addr``
107 will show the address of the system call instruction, and
108 ``siginfo->si_syscall`` and ``siginfo->si_arch`` will indicate which
109 syscall was attempted. The program counter will be as though
110 the syscall happened (i.e. it will not point to the syscall
111 instruction). The return value register will contain an arch-
112 dependent value -- if resuming execution, set it to something
113 sensible. (The architecture dependency is because replacing
114 it with ``-ENOSYS`` could overwrite some useful information.)
115
116 The ``SECCOMP_RET_DATA`` portion of the return value will be passed
117 as ``si_errno``.
118
119 ``SIGSYS`` triggered by seccomp will have a si_code of ``SYS_SECCOMP``.
120
121 ``SECCOMP_RET_ERRNO``:
122 Results in the lower 16-bits of the return value being passed
123 to userland as the errno without executing the system call.
124
125 ``SECCOMP_RET_USER_NOTIF``:
126 Results in a ``struct seccomp_notif`` message sent on the userspace
127 notification fd, if it is attached, or ``-ENOSYS`` if it is not. See
128 below on discussion of how to handle user notifications.
129
130 ``SECCOMP_RET_TRACE``:
131 When returned, this value will cause the kernel to attempt to
132 notify a ``ptrace()``-based tracer prior to executing the system
133 call. If there is no tracer present, ``-ENOSYS`` is returned to
134 userland and the system call is not executed.
135
136 A tracer will be notified if it requests ``PTRACE_O_TRACESECCOMP``
137 using ``ptrace(PTRACE_SETOPTIONS)``. The tracer will be notified
138 of a ``PTRACE_EVENT_SECCOMP`` and the ``SECCOMP_RET_DATA`` portion of
139 the BPF program return value will be available to the tracer
140 via ``PTRACE_GETEVENTMSG``.
141
142 The tracer can skip the system call by changing the syscall number
143 to -1. Alternatively, the tracer can change the system call
144 requested by changing the system call to a valid syscall number. If
145 the tracer asks to skip the system call, then the system call will
146 appear to return the value that the tracer puts in the return value
147 register.
148
149 The seccomp check will not be run again after the tracer is
150 notified. (This means that seccomp-based sandboxes MUST NOT
151 allow use of ptrace, even of other sandboxed processes, without
152 extreme care; ptracers can use this mechanism to escape.)
153
154 ``SECCOMP_RET_LOG``:
155 Results in the system call being executed after it is logged. This
156 should be used by application developers to learn which syscalls their
157 application needs without having to iterate through multiple test and
158 development cycles to build the list.
159
160 This action will only be logged if "log" is present in the
161 actions_logged sysctl string.
162
163 ``SECCOMP_RET_ALLOW``:
164 Results in the system call being executed.
165
166 If multiple filters exist, the return value for the evaluation of a
167 given system call will always use the highest precedent value.
168
169 Precedence is only determined using the ``SECCOMP_RET_ACTION`` mask. When
170 multiple filters return values of the same precedence, only the
171 ``SECCOMP_RET_DATA`` from the most recently installed filter will be
172 returned.
173
174 Pitfalls
175 ========
176
177 The biggest pitfall to avoid during use is filtering on system call
178 number without checking the architecture value. Why? On any
179 architecture that supports multiple system call invocation conventions,
180 the system call numbers may vary based on the specific invocation. If
181 the numbers in the different calling conventions overlap, then checks in
182 the filters may be abused. Always check the arch value!
183
184 Example
185 =======
186
187 The ``samples/seccomp/`` directory contains both an x86-specific example
188 and a more generic example of a higher level macro interface for BPF
189 program generation.
190
191 Userspace Notification
192 ======================
193
194 The ``SECCOMP_RET_USER_NOTIF`` return code lets seccomp filters pass a
195 particular syscall to userspace to be handled. This may be useful for
196 applications like container managers, which wish to intercept particular
197 syscalls (``mount()``, ``finit_module()``, etc.) and change their behavior.
198
199 To acquire a notification FD, use the ``SECCOMP_FILTER_FLAG_NEW_LISTENER``
200 argument to the ``seccomp()`` syscall:
201
202 .. code-block:: c
203
204 fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);
205
206 which (on success) will return a listener fd for the filter, which can then be
207 passed around via ``SCM_RIGHTS`` or similar. Note that filter fds correspond to
208 a particular filter, and not a particular task. So if this task then forks,
209 notifications from both tasks will appear on the same filter fd. Reads and
210 writes to/from a filter fd are also synchronized, so a filter fd can safely
211 have many readers.
212
213 The interface for a seccomp notification fd consists of two structures:
214
215 .. code-block:: c
216
217 struct seccomp_notif_sizes {
218 __u16 seccomp_notif;
219 __u16 seccomp_notif_resp;
220 __u16 seccomp_data;
221 };
222
223 struct seccomp_notif {
224 __u64 id;
225 __u32 pid;
226 __u32 flags;
227 struct seccomp_data data;
228 };
229
230 struct seccomp_notif_resp {
231 __u64 id;
232 __s64 val;
233 __s32 error;
234 __u32 flags;
235 };
236
237 The ``struct seccomp_notif_sizes`` structure can be used to determine the size
238 of the various structures used in seccomp notifications. The size of ``struct
239 seccomp_data`` may change in the future, so code should use:
240
241 .. code-block:: c
242
243 struct seccomp_notif_sizes sizes;
244 seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);
245
246 to determine the size of the various structures to allocate. See
247 samples/seccomp/user-trap.c for an example.
248
249 Users can read via ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)`` (or ``poll()``) on a
250 seccomp notification fd to receive a ``struct seccomp_notif``, which contains
251 five members: the input length of the structure, a unique-per-filter ``id``,
252 the ``pid`` of the task which triggered this request (which may be 0 if the
253 task is in a pid ns not visible from the listener's pid namespace). The
254 notification also contains the ``data`` passed to seccomp, and a filters flag.
255 The structure should be zeroed out prior to calling the ioctl.
256
257 Userspace can then make a decision based on this information about what to do,
258 and ``ioctl(SECCOMP_IOCTL_NOTIF_SEND)`` a response, indicating what should be
259 returned to userspace. The ``id`` member of ``struct seccomp_notif_resp`` should
260 be the same ``id`` as in ``struct seccomp_notif``.
261
262 Userspace can also add file descriptors to the notifying process via
263 ``ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)``. The ``id`` member of
264 ``struct seccomp_notif_addfd`` should be the same ``id`` as in
265 ``struct seccomp_notif``. The ``newfd_flags`` flag may be used to set flags
266 like O_CLOEXEC on the file descriptor in the notifying process. If the supervisor
267 wants to inject the file descriptor with a specific number, the
268 ``SECCOMP_ADDFD_FLAG_SETFD`` flag can be used, and set the ``newfd`` member to
269 the specific number to use. If that file descriptor is already open in the
270 notifying process it will be replaced. The supervisor can also add an FD, and
271 respond atomically by using the ``SECCOMP_ADDFD_FLAG_SEND`` flag and the return
272 value will be the injected file descriptor number.
273
274 The notifying process can be preempted, resulting in the notification being
275 aborted. This can be problematic when trying to take actions on behalf of the
276 notifying process that are long-running and typically retryable (mounting a
277 filesystem). Alternatively, at filter installation time, the
278 ``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`` flag can be set. This flag makes it
279 such that when a user notification is received by the supervisor, the notifying
280 process will ignore non-fatal signals until the response is sent. Signals that
281 are sent prior to the notification being received by userspace are handled
282 normally.
283
284 It is worth noting that ``struct seccomp_data`` contains the values of register
285 arguments to the syscall, but does not contain pointers to memory. The task's
286 memory is accessible to suitably privileged traces via ``ptrace()`` or
287 ``/proc/pid/mem``. However, care should be taken to avoid the TOCTOU mentioned
288 above in this document: all arguments being read from the tracee's memory
289 should be read into the tracer's memory before any policy decisions are made.
290 This allows for an atomic decision on syscall arguments.
291
292 Sysctls
293 =======
294
295 Seccomp's sysctl files can be found in the ``/proc/sys/kernel/seccomp/``
296 directory. Here's a description of each file in that directory:
297
298 ``actions_avail``:
299 A read-only ordered list of seccomp return values (refer to the
300 ``SECCOMP_RET_*`` macros above) in string form. The ordering, from
301 left-to-right, is the least permissive return value to the most
302 permissive return value.
303
304 The list represents the set of seccomp return values supported
305 by the kernel. A userspace program may use this list to
306 determine if the actions found in the ``seccomp.h``, when the
307 program was built, differs from the set of actions actually
308 supported in the current running kernel.
309
310 ``actions_logged``:
311 A read-write ordered list of seccomp return values (refer to the
312 ``SECCOMP_RET_*`` macros above) that are allowed to be logged. Writes
313 to the file do not need to be in ordered form but reads from the file
314 will be ordered in the same way as the actions_avail sysctl.
315
316 The ``allow`` string is not accepted in the ``actions_logged`` sysctl
317 as it is not possible to log ``SECCOMP_RET_ALLOW`` actions. Attempting
318 to write ``allow`` to the sysctl will result in an EINVAL being
319 returned.
320
321 Adding architecture support
322 ===========================
323
324 See ``arch/Kconfig`` for the authoritative requirements. In general, if an
325 architecture supports both ptrace_event and seccomp, it will be able to
326 support seccomp filter with minor fixup: ``SIGSYS`` support and seccomp return
327 value checking. Then it must just add ``CONFIG_HAVE_ARCH_SECCOMP_FILTER``
328 to its arch-specific Kconfig.
329
330
331
332 Caveats
333 =======
334
335 The vDSO can cause some system calls to run entirely in userspace,
336 leading to surprises when you run programs on different machines that
337 fall back to real syscalls. To minimize these surprises on x86, make
338 sure you test with
339 ``/sys/devices/system/clocksource/clocksource0/current_clocksource`` set to
340 something like ``acpi_pm``.
341
342 On x86-64, vsyscall emulation is enabled by default. (vsyscalls are
343 legacy variants on vDSO calls.) Currently, emulated vsyscalls will
344 honor seccomp, with a few oddities:
345
346 - A return value of ``SECCOMP_RET_TRAP`` will set a ``si_call_addr`` pointing to
347 the vsyscall entry for the given call and not the address after the
348 'syscall' instruction. Any code which wants to restart the call
349 should be aware that (a) a ret instruction has been emulated and (b)
350 trying to resume the syscall will again trigger the standard vsyscall
351 emulation security checks, making resuming the syscall mostly
352 pointless.
353
354 - A return value of ``SECCOMP_RET_TRACE`` will signal the tracer as usual,
355 but the syscall may not be changed to another system call using the
356 orig_rax register. It may only be changed to -1 order to skip the
357 currently emulated call. Any other change MAY terminate the process.
358 The rip value seen by the tracer will be the syscall entry address;
359 this is different from normal behavior. The tracer MUST NOT modify
360 rip or rsp. (Do not rely on other changes terminating the process.
361 They might work. For example, on some kernels, choosing a syscall
362 that only exists in future kernels will be correctly emulated (by
363 returning ``-ENOSYS``).
364
365 To detect this quirky behavior, check for ``addr & ~0x0C00 ==
366 0xFFFFFFFFFF600000``. (For ``SECCOMP_RET_TRACE``, use rip. For
367 ``SECCOMP_RET_TRAP``, use ``siginfo->si_call_addr``.) Do not check any other
368 condition: future kernels may improve vsyscall emulation and current
369 kernels in vsyscall=native mode will behave differently, but the
370 instructions at ``0xF...F600{0,4,8,C}00`` will not be system calls in these
371 cases.
372
373 Note that modern systems are unlikely to use vsyscalls at all -- they
374 are a legacy feature and they are considerably slower than standard
375 syscalls. New code will use the vDSO, and vDSO-issued system calls
376 are indistinguishable from normal system calls.
377

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

소개

1-29

모든 사용자 공간 프로세스에는 많은 system call이 노출되지만, 상당수는 프로세스 수명 동안 한 번도 쓰이지 않습니다. system call이 변하고 성숙하는 과정에서 버그가 발견되고 제거되므로 일부 애플리케이션은 사용할 수 있는 system call 집합을 줄여 커널 공격 표면을 축소하는 것이 유리합니다. system call filtering은 이런 애플리케이션을 위한 기능입니다.

seccomp filtering을 사용하면 프로세스가 들어오는 system call에 적용할 필터를 지정할 수 있습니다. 필터는 socket filter와 같은 Berkeley Packet Filter, 즉 BPF 프로그램이지만 입력 데이터는 호출 번호, 인자와 관련 메타데이터입니다. 오래 검증된 사용자 공간 필터 언어와 단순한 입력 집합으로 표현력 있는 system call 정책을 작성할 수 있습니다.

BPF 프로그램은 포인터를 역참조할 수 없고 system call 인자 값 자체만 평가합니다. 이 제약 덕분에 seccomp 사용자는 일반적인 system call interposition 프레임워크의 time-of-check-time-of-use, 즉 TOCTOU 공격을 피할 수 있습니다.

===========================================
Seccomp BPF (SECure COMPuting with filters)
===========================================

Introduction
============

A large number of system calls are exposed to every userland process
with many of them going unused for the entire lifetime of the process.
As system calls change and mature, bugs are found and eradicated.  A
certain subset of userland applications benefit by having a reduced set
of available system calls.  The resulting set reduces the total kernel
surface exposed to the application.  System call filtering is meant for
use with those applications.

Seccomp filtering provides a means for a process to specify a filter for
incoming system calls.  The filter is expressed as a Berkeley Packet
Filter (BPF) program, as with socket filters, except that the data
operated on is related to the system call being made: system call
number and the system call arguments.  This allows for expressive
filtering of system calls using a filter program language with a long
history of being exposed to userland and a straightforward data set.

Additionally, BPF makes it impossible for users of seccomp to fall prey
to time-of-check-time-of-use (TOCTOU) attacks that are common in system
call interposition frameworks.  BPF programs may not dereference
pointers which constrains all filters to solely evaluating the system
call arguments directly.

seccomp가 아닌 것

30-42

system call filtering 자체는 sandbox가 아닙니다. 노출된 커널 표면을 최소화하는 명확한 메커니즘이며 sandbox 개발자가 사용하는 구성 요소입니다. 논리적 동작과 정보 흐름 정책은 다른 시스템 강화 기법, 필요하다면 선택한 LSM과 함께 관리해야 합니다.

동적이고 표현력 있는 필터로 비정상적으로 큰 정책을 피하거나 `socketcall()`이 multiplex하는 호출 중 일부만 허용할 수 있지만, 이를 완전한 sandbox 해법으로 오해해서는 안 됩니다.

What it isn't
=============

System call filtering isn't a sandbox.  It provides a clearly defined
mechanism for minimizing the exposed kernel surface.  It is meant to be
a tool for sandbox developers to use.  Beyond that, policy for logical
behavior and information flow should be managed with a combination of
other system hardening techniques and, potentially, an LSM of your
choosing.  Expressive, dynamic filters provide further options down this
path (avoiding pathological sizes or selecting which of the multiplexed
system calls in socketcall() is allowed, for instance) which could be
construed, incorrectly, as a more complete sandboxing solution.

필터 설치와 상속

43-83

아키텍처가 `CONFIG_HAVE_ARCH_SECCOMP_FILTER`를 제공하면 strict seccomp와 같은 `prctl(2)` 호출에 추가 모드를 지정해 필터를 설치할 수 있습니다.

`PR_SET_SECCOMP`는 BPF 프로그램으로 새 필터를 지정하는 인자를 추가로 받습니다. 프로그램은 system call 번호, 인자와 메타데이터를 담은 `struct seccomp_data`를 평가하고 커널이 수행할 action 값을 반환해야 합니다.

prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);

`prog`는 필터 프로그램을 담은 `struct sock_fprog` 포인터입니다. 프로그램이 잘못되면 호출은 -1을 반환하고 `errno`를 `EINVAL`로 설정합니다.

필터가 `fork`/`clone`과 `execve`를 허용하면 자식 프로세스도 부모와 같은 필터 및 system call ABI 제약을 받습니다.

설치 전에 태스크가 `prctl(PR_SET_NO_NEW_PRIVS, 1)`을 호출했거나 자신의 namespace에서 `CAP_SYS_ADMIN` 권한을 가져야 합니다. 아니면 `-EACCES`를 반환합니다. 이 조건은 설치자보다 권한이 큰 자식에게 필터 프로그램을 적용해 권한을 악용하는 일을 막습니다.

연결된 필터가 `prctl(2)`을 허용하면 필터를 추가로 겹쳐 적용할 수 있습니다. 평가 시간은 늘지만 실행 중 공격 표면을 더 줄일 수 있습니다. 설치 호출은 성공 시 0, 오류 시 0이 아닌 값을 반환합니다.

Usage
=====

An additional seccomp mode is added and is enabled using the same
prctl(2) call as the strict seccomp.  If the architecture has
``CONFIG_HAVE_ARCH_SECCOMP_FILTER``, then filters may be added as below:

``PR_SET_SECCOMP``:
	Now takes an additional argument which specifies a new filter
	using a BPF program.
	The BPF program will be executed over struct seccomp_data
	reflecting the system call number, arguments, and other
	metadata.  The BPF program must then return one of the
	acceptable values to inform the kernel which action should be
	taken.

	Usage::

		prctl(PR_SET_SECCOMP, SECCOMP_MODE_FILTER, prog);

	The 'prog' argument is a pointer to a struct sock_fprog which
	will contain the filter program.  If the program is invalid, the
	call will return -1 and set errno to ``EINVAL``.

	If ``fork``/``clone`` and ``execve`` are allowed by @prog, any child
	processes will be constrained to the same filters and system
	call ABI as the parent.

	Prior to use, the task must call ``prctl(PR_SET_NO_NEW_PRIVS, 1)`` or
	run with ``CAP_SYS_ADMIN`` privileges in its namespace.  If these are not
	true, ``-EACCES`` will be returned.  This requirement ensures that filter
	programs cannot be applied to child processes with greater privileges
	than the task that installed them.

	Additionally, if ``prctl(2)`` is allowed by the attached filter,
	additional filters may be layered on which will increase evaluation
	time, but allow for further decreasing the attack surface during
	execution of a process.

The above call returns 0 on success and non-zero on error.

반환값: 종료와 trap

84-120

여러 필터가 있으면 주어진 system call 평가에서 가장 우선순위가 높은 action이 항상 선택됩니다. 예를 들어 `SECCOMP_RET_KILL_PROCESS`가 언제나 우선합니다.

높은 우선순위의 seccomp action
Action결과
`SECCOMP_RET_KILL_PROCESS`system call을 실행하지 않고 전체 프로세스를 즉시 종료합니다. 종료 신호는 `SIGKILL`이 아니라 `SIGSYS`입니다.
`SECCOMP_RET_KILL_THREAD`system call을 실행하지 않고 현재 태스크를 즉시 종료합니다. 종료 신호는 `SIGSYS`입니다.
`SECCOMP_RET_TRAP`호출을 실행하지 않고 발생 태스크에 `SIGSYS`를 보냅니다.

위에서 아래 순서로 우선순위가 낮아집니다.

TRAP의 `siginfo->si_call_addr`는 system call 명령 주소를, `si_syscall`과 `si_arch`는 시도한 호출을 나타냅니다. program counter는 호출이 일어난 뒤처럼 보이며 syscall 명령 자체를 가리키지 않습니다. 반환값 register는 아키텍처 의존 값이므로 실행을 재개하려면 의미 있는 값으로 설정해야 합니다.

TRAP 반환값의 `SECCOMP_RET_DATA` 부분은 `si_errno`로 전달되고, seccomp가 발생시킨 `SIGSYS`의 `si_code`는 `SYS_SECCOMP`입니다.

Return values
=============

A seccomp filter may return any of the following values. If multiple
filters exist, the return value for the evaluation of a given system
call will always use the highest precedent value. (For example,
``SECCOMP_RET_KILL_PROCESS`` will always take precedence.)

In precedence order, they are:

``SECCOMP_RET_KILL_PROCESS``:
	Results in the entire process exiting immediately without executing
	the system call.  The exit status of the task (``status & 0x7f``)
	will be ``SIGSYS``, not ``SIGKILL``.

``SECCOMP_RET_KILL_THREAD``:
	Results in the task exiting immediately without executing the
	system call.  The exit status of the task (``status & 0x7f``) will
	be ``SIGSYS``, not ``SIGKILL``.

``SECCOMP_RET_TRAP``:
	Results in the kernel sending a ``SIGSYS`` signal to the triggering
	task without executing the system call. ``siginfo->si_call_addr``
	will show the address of the system call instruction, and
	``siginfo->si_syscall`` and ``siginfo->si_arch`` will indicate which
	syscall was attempted.  The program counter will be as though
	the syscall happened (i.e. it will not point to the syscall
	instruction).  The return value register will contain an arch-
	dependent value -- if resuming execution, set it to something
	sensible.  (The architecture dependency is because replacing
	it with ``-ENOSYS`` could overwrite some useful information.)

	The ``SECCOMP_RET_DATA`` portion of the return value will be passed
	as ``si_errno``.

	``SIGSYS`` triggered by seccomp will have a si_code of ``SYS_SECCOMP``.

반환값: errno, 알림, trace, log, allow

121-173
나머지 seccomp action
Action동작
`SECCOMP_RET_ERRNO`system call을 실행하지 않고 반환값 하위 16비트를 사용자 공간 errno로 전달합니다.
`SECCOMP_RET_USER_NOTIF`listener가 있으면 notification FD로 `struct seccomp_notif`를 보내고, 없으면 `-ENOSYS`입니다.
`SECCOMP_RET_TRACE`실행 전에 `ptrace()` tracer 통지를 시도하며 tracer가 없으면 호출을 실행하지 않고 `-ENOSYS`입니다.
`SECCOMP_RET_LOG`기록한 뒤 system call을 실행합니다. `actions_logged`에 `log`가 있을 때만 기록됩니다.
`SECCOMP_RET_ALLOW`system call을 실행합니다.

각 action은 호출 실행 여부와 관찰 주체가 다릅니다.

`SECCOMP_RET_LOG`는 애플리케이션 개발자가 여러 차례의 시험·개발 주기를 반복하지 않고 실제로 필요한 syscall 목록을 학습하도록 마련된 action입니다. 호출은 차단되지 않으며 기록 여부는 `actions_logged` 설정에도 좌우됩니다.

TRACE에서 `PTRACE_O_TRACESECCOMP`를 `ptrace(PTRACE_SETOPTIONS)`로 요청한 tracer는 `PTRACE_EVENT_SECCOMP` 통지와 `PTRACE_GETEVENTMSG`를 통한 `SECCOMP_RET_DATA`를 받습니다.

tracer는 syscall 번호를 -1로 바꿔 호출을 건너뛰거나 유효한 번호로 바꿀 수 있습니다. 건너뛰면 tracer가 반환값 register에 넣은 값을 system call 결과로 보게 됩니다.

tracer 통지 뒤 seccomp 검사를 다시 하지 않습니다. 따라서 seccomp sandbox가 ptrace를 허용하면 tracer가 이 경로로 탈출할 수 있으므로 극도로 주의해야 합니다.

여러 필터 결과의 우선순위는 `SECCOMP_RET_ACTION` mask만으로 정합니다. 같은 우선순위라면 가장 최근에 설치한 필터의 `SECCOMP_RET_DATA`만 반환합니다.

``SECCOMP_RET_ERRNO``:
	Results in the lower 16-bits of the return value being passed
	to userland as the errno without executing the system call.

``SECCOMP_RET_USER_NOTIF``:
	Results in a ``struct seccomp_notif`` message sent on the userspace
	notification fd, if it is attached, or ``-ENOSYS`` if it is not. See
	below on discussion of how to handle user notifications.

``SECCOMP_RET_TRACE``:
	When returned, this value will cause the kernel to attempt to
	notify a ``ptrace()``-based tracer prior to executing the system
	call.  If there is no tracer present, ``-ENOSYS`` is returned to
	userland and the system call is not executed.

	A tracer will be notified if it requests ``PTRACE_O_TRACESECCOMP``
	using ``ptrace(PTRACE_SETOPTIONS)``.  The tracer will be notified
	of a ``PTRACE_EVENT_SECCOMP`` and the ``SECCOMP_RET_DATA`` portion of
	the BPF program return value will be available to the tracer
	via ``PTRACE_GETEVENTMSG``.

	The tracer can skip the system call by changing the syscall number
	to -1.  Alternatively, the tracer can change the system call
	requested by changing the system call to a valid syscall number.  If
	the tracer asks to skip the system call, then the system call will
	appear to return the value that the tracer puts in the return value
	register.

	The seccomp check will not be run again after the tracer is
	notified.  (This means that seccomp-based sandboxes MUST NOT
	allow use of ptrace, even of other sandboxed processes, without
	extreme care; ptracers can use this mechanism to escape.)

``SECCOMP_RET_LOG``:
	Results in the system call being executed after it is logged. This
	should be used by application developers to learn which syscalls their
	application needs without having to iterate through multiple test and
	development cycles to build the list.

	This action will only be logged if "log" is present in the
	actions_logged sysctl string.

``SECCOMP_RET_ALLOW``:
	Results in the system call being executed.

If multiple filters exist, the return value for the evaluation of a
given system call will always use the highest precedent value.

Precedence is only determined using the ``SECCOMP_RET_ACTION`` mask.  When
multiple filters return values of the same precedence, only the
``SECCOMP_RET_DATA`` from the most recently installed filter will be
returned.

아키텍처 확인과 예제

174-190

가장 큰 함정은 architecture 값을 확인하지 않고 system call 번호만 필터링하는 것입니다. 여러 호출 규약을 지원하는 아키텍처에서는 규약에 따라 번호가 다르고 서로 겹칠 수 있어 필터 검사가 악용될 수 있습니다. 항상 arch 값을 검사해야 합니다.

`samples/seccomp/` 디렉터리에는 x86 전용 예제와 BPF 프로그램 생성을 위한 상위 수준 macro 인터페이스의 범용 예제가 있습니다.

Pitfalls
========

The biggest pitfall to avoid during use is filtering on system call
number without checking the architecture value.  Why?  On any
architecture that supports multiple system call invocation conventions,
the system call numbers may vary based on the specific invocation.  If
the numbers in the different calling conventions overlap, then checks in
the filters may be abused.  Always check the arch value!

Example
=======

The ``samples/seccomp/`` directory contains both an x86-specific example
and a more generic example of a higher level macro interface for BPF
program generation.

사용자 공간 알림 설정과 구조체

191-248

`SECCOMP_RET_USER_NOTIF`는 특정 system call을 사용자 공간으로 넘겨 처리하게 합니다. container manager가 `mount()`나 `finit_module()` 같은 호출을 가로채 동작을 바꾸는 경우에 유용합니다.

notification FD를 얻으려면 `seccomp()` 호출에 `SECCOMP_FILTER_FLAG_NEW_LISTENER`를 지정합니다.

fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);

성공하면 listener FD를 반환하며 `SCM_RIGHTS` 등으로 전달할 수 있습니다. filter FD는 특정 태스크가 아니라 특정 필터에 대응하므로 설치 태스크가 fork하면 두 태스크의 알림이 같은 FD에 나타납니다. FD read/write는 동기화되어 여러 reader가 안전하게 공유할 수 있습니다.

seccomp notification FD 인터페이스는 크기 정보와 요청·응답 구조체를 사용합니다.

struct seccomp_notif_sizes {
    __u16 seccomp_notif;
    __u16 seccomp_notif_resp;
    __u16 seccomp_data;
};

struct seccomp_notif {
    __u64 id;
    __u32 pid;
    __u32 flags;
    struct seccomp_data data;
};

struct seccomp_notif_resp {
    __u64 id;
    __s64 val;
    __s32 error;
    __u32 flags;
};

`struct seccomp_data` 크기는 미래에 바뀔 수 있으므로 `SECCOMP_GET_NOTIF_SIZES`로 각 구조체의 실제 크기를 조회해 할당해야 합니다.

struct seccomp_notif_sizes sizes;
seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);

구현 예제는 `samples/seccomp/user-trap.c`에서 볼 수 있습니다.

Userspace Notification
======================

The ``SECCOMP_RET_USER_NOTIF`` return code lets seccomp filters pass a
particular syscall to userspace to be handled. This may be useful for
applications like container managers, which wish to intercept particular
syscalls (``mount()``, ``finit_module()``, etc.) and change their behavior.

To acquire a notification FD, use the ``SECCOMP_FILTER_FLAG_NEW_LISTENER``
argument to the ``seccomp()`` syscall:

.. code-block:: c

    fd = seccomp(SECCOMP_SET_MODE_FILTER, SECCOMP_FILTER_FLAG_NEW_LISTENER, &prog);

which (on success) will return a listener fd for the filter, which can then be
passed around via ``SCM_RIGHTS`` or similar. Note that filter fds correspond to
a particular filter, and not a particular task. So if this task then forks,
notifications from both tasks will appear on the same filter fd. Reads and
writes to/from a filter fd are also synchronized, so a filter fd can safely
have many readers.

The interface for a seccomp notification fd consists of two structures:

.. code-block:: c

    struct seccomp_notif_sizes {
        __u16 seccomp_notif;
        __u16 seccomp_notif_resp;
        __u16 seccomp_data;
    };

    struct seccomp_notif {
        __u64 id;
        __u32 pid;
        __u32 flags;
        struct seccomp_data data;
    };

    struct seccomp_notif_resp {
        __u64 id;
        __s64 val;
        __s32 error;
        __u32 flags;
    };

The ``struct seccomp_notif_sizes`` structure can be used to determine the size
of the various structures used in seccomp notifications. The size of ``struct
seccomp_data`` may change in the future, so code should use:

.. code-block:: c

    struct seccomp_notif_sizes sizes;
    seccomp(SECCOMP_GET_NOTIF_SIZES, 0, &sizes);

to determine the size of the various structures to allocate. See
samples/seccomp/user-trap.c for an example.

알림 수신, 응답과 FD 주입

249-291

사용자 공간은 notification FD에서 `ioctl(SECCOMP_IOCTL_NOTIF_RECV)` 또는 `poll()`로 `struct seccomp_notif`를 받습니다. 알림에는 입력 구조체 길이, 필터별 고유 `id`, 요청을 발생시킨 태스크의 `pid`, seccomp `data`, filter flag가 들어갑니다. listener의 PID namespace에서 태스크가 보이지 않으면 pid는 0일 수 있습니다. ioctl 전 구조체를 0으로 초기화해야 합니다.

supervisor는 정보를 바탕으로 결정을 내린 뒤 `ioctl(SECCOMP_IOCTL_NOTIF_SEND)`로 사용자 공간에 반환할 결과를 보냅니다. `struct seccomp_notif_resp.id`는 요청의 `struct seccomp_notif.id`와 같아야 합니다.

`ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)`로 알림을 발생시킨 프로세스에 file descriptor를 추가할 수도 있습니다. `struct seccomp_notif_addfd.id`도 요청 id와 같아야 하고, `newfd_flags`로 O_CLOEXEC 같은 flag를 설정합니다.

특정 번호로 FD를 주입하려면 `SECCOMP_ADDFD_FLAG_SETFD`와 `newfd`를 사용합니다. 대상 번호가 열려 있으면 교체됩니다. `SECCOMP_ADDFD_FLAG_SEND`를 사용하면 FD 추가와 응답을 원자적으로 수행하고 주입된 FD 번호를 반환받습니다.

알림 프로세스가 preempt되면 notification이 중단될 수 있어 filesystem mount처럼 오래 걸리고 보통 재시도 가능한 작업에 문제가 됩니다. 필터 설치 때 `SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`를 설정하면 supervisor가 알림을 받은 뒤 응답을 보낼 때까지 알림 프로세스가 fatal이 아닌 signal을 무시합니다. 사용자 공간 수신 전의 signal은 정상 처리됩니다.

`struct seccomp_data`에는 syscall register 인자 값이 있지만 메모리 포인터가 가리키는 내용은 없습니다. 충분한 권한의 tracer는 `ptrace()`나 `/proc/pid/mem`으로 태스크 메모리를 읽을 수 있지만 TOCTOU를 피하려면 정책 결정을 내리기 전에 tracee 메모리의 모든 인자를 tracer 메모리로 먼저 복사해야 합니다. 그래야 syscall 인자에 대한 원자적 결정을 내릴 수 있습니다.

Users can read via ``ioctl(SECCOMP_IOCTL_NOTIF_RECV)``  (or ``poll()``) on a
seccomp notification fd to receive a ``struct seccomp_notif``, which contains
five members: the input length of the structure, a unique-per-filter ``id``,
the ``pid`` of the task which triggered this request (which may be 0 if the
task is in a pid ns not visible from the listener's pid namespace). The
notification also contains the ``data`` passed to seccomp, and a filters flag.
The structure should be zeroed out prior to calling the ioctl.

Userspace can then make a decision based on this information about what to do,
and ``ioctl(SECCOMP_IOCTL_NOTIF_SEND)`` a response, indicating what should be
returned to userspace. The ``id`` member of ``struct seccomp_notif_resp`` should
be the same ``id`` as in ``struct seccomp_notif``.

Userspace can also add file descriptors to the notifying process via
``ioctl(SECCOMP_IOCTL_NOTIF_ADDFD)``. The ``id`` member of
``struct seccomp_notif_addfd`` should be the same ``id`` as in
``struct seccomp_notif``. The ``newfd_flags`` flag may be used to set flags
like O_CLOEXEC on the file descriptor in the notifying process. If the supervisor
wants to inject the file descriptor with a specific number, the
``SECCOMP_ADDFD_FLAG_SETFD`` flag can be used, and set the ``newfd`` member to
the specific number to use. If that file descriptor is already open in the
notifying process it will be replaced. The supervisor can also add an FD, and
respond atomically by using the ``SECCOMP_ADDFD_FLAG_SEND`` flag and the return
value will be the injected file descriptor number.

The notifying process can be preempted, resulting in the notification being
aborted. This can be problematic when trying to take actions on behalf of the
notifying process that are long-running and typically retryable (mounting a
filesystem). Alternatively, at filter installation time, the
``SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV`` flag can be set. This flag makes it
such that when a user notification is received by the supervisor, the notifying
process will ignore non-fatal signals until the response is sent. Signals that
are sent prior to the notification being received by userspace are handled
normally.

It is worth noting that ``struct seccomp_data`` contains the values of register
arguments to the syscall, but does not contain pointers to memory. The task's
memory is accessible to suitably privileged traces via ``ptrace()`` or
``/proc/pid/mem``. However, care should be taken to avoid the TOCTOU mentioned
above in this document: all arguments being read from the tracee's memory
should be read into the tracer's memory before any policy decisions are made.
This allows for an atomic decision on syscall arguments.

Seccomp sysctl

292-320

seccomp sysctl은 `/proc/sys/kernel/seccomp/`에 있습니다.

Seccomp sysctl 파일
파일접근의미
`actions_avail`read-only지원하는 `SECCOMP_RET_*` 문자열을 가장 제한적인 값에서 가장 허용적인 값 순으로 표시합니다. 빌드 시 header와 실행 커널의 지원 차이를 확인할 수 있습니다.
`actions_logged`read-write기록이 허용된 action 목록입니다. 쓰기 순서는 자유롭지만 읽을 때는 `actions_avail`과 같은 순서입니다.

현재 커널이 지원하고 기록할 수 있는 action 집합을 제공합니다.

`SECCOMP_RET_ALLOW` action은 기록할 수 없으므로 `actions_logged`는 `allow` 문자열을 받지 않으며 쓰려고 하면 `EINVAL`을 반환합니다.

Sysctls
=======

Seccomp's sysctl files can be found in the ``/proc/sys/kernel/seccomp/``
directory. Here's a description of each file in that directory:

``actions_avail``:
	A read-only ordered list of seccomp return values (refer to the
	``SECCOMP_RET_*`` macros above) in string form. The ordering, from
	left-to-right, is the least permissive return value to the most
	permissive return value.

	The list represents the set of seccomp return values supported
	by the kernel. A userspace program may use this list to
	determine if the actions found in the ``seccomp.h``, when the
	program was built, differs from the set of actions actually
	supported in the current running kernel.

``actions_logged``:
	A read-write ordered list of seccomp return values (refer to the
	``SECCOMP_RET_*`` macros above) that are allowed to be logged. Writes
	to the file do not need to be in ordered form but reads from the file
	will be ordered in the same way as the actions_avail sysctl.

	The ``allow`` string is not accepted in the ``actions_logged`` sysctl
	as it is not possible to log ``SECCOMP_RET_ALLOW`` actions. Attempting
	to write ``allow`` to the sysctl will result in an EINVAL being
	returned.

아키텍처 지원 추가

321-331

권위 있는 요구사항은 `arch/Kconfig`에 있습니다. 일반적으로 ptrace_event와 seccomp를 모두 지원하는 아키텍처는 `SIGSYS` 지원과 seccomp 반환값 검사를 약간 보완하면 필터를 지원할 수 있습니다. 그 뒤 아키텍처별 Kconfig에 `CONFIG_HAVE_ARCH_SECCOMP_FILTER`를 추가합니다.

Adding architecture support
===========================

See ``arch/Kconfig`` for the authoritative requirements.  In general, if an
architecture supports both ptrace_event and seccomp, it will be able to
support seccomp filter with minor fixup: ``SIGSYS`` support and seccomp return
value checking.  Then it must just add ``CONFIG_HAVE_ARCH_SECCOMP_FILTER``
to its arch-specific Kconfig.


vDSO와 vsyscall 주의사항

332-376

vDSO는 일부 system call을 전적으로 사용자 공간에서 실행할 수 있습니다. 다른 시스템에서 실제 syscall로 fallback할 때 동작이 달라질 수 있습니다. x86에서 이를 시험하려면 `/sys/devices/system/clocksource/clocksource0/current_clocksource`를 `acpi_pm` 같은 값으로 설정하십시오.

x86-64는 기본으로 legacy vDSO 변형인 vsyscall emulation을 사용합니다. emulated vsyscall도 seccomp를 따르지만 예외가 있습니다.

Emulated vsyscall의 seccomp 차이
Action특이점
`SECCOMP_RET_TRAP``si_call_addr`는 syscall 다음 주소가 아니라 해당 vsyscall entry를 가리킵니다. ret이 이미 emulation되었고 재개하면 보안 검사를 다시 거치므로 syscall 재시작은 거의 의미가 없습니다.
`SECCOMP_RET_TRACE`tracer 통지는 정상이나 `orig_rax`로 다른 syscall로 바꿀 수 없고 -1로 현재 호출을 건너뛰는 것만 허용됩니다. 다른 변경은 프로세스를 종료할 수 있습니다. `rip`은 syscall entry 주소이며 tracer는 `rip` 또는 `rsp`를 바꾸면 안 됩니다.

일반 syscall 재시작이나 ptrace 처리와 같은 방식으로 다루면 안 됩니다.

이 동작은 `addr & ~0x0C00 == 0xFFFFFFFFFF600000`으로 감지합니다. TRACE에는 rip, TRAP에는 `siginfo->si_call_addr`를 사용합니다. 미래 커널이나 `vsyscall=native`가 다르게 동작할 수 있으므로 다른 조건은 검사하지 마십시오.

현대 시스템은 legacy이고 일반 syscall보다 훨씬 느린 vsyscall을 거의 쓰지 않습니다. 새 코드는 vDSO를 사용하며 vDSO가 실제로 발행한 system call은 일반 system call과 구별되지 않습니다.

Caveats
=======

The vDSO can cause some system calls to run entirely in userspace,
leading to surprises when you run programs on different machines that
fall back to real syscalls.  To minimize these surprises on x86, make
sure you test with
``/sys/devices/system/clocksource/clocksource0/current_clocksource`` set to
something like ``acpi_pm``.

On x86-64, vsyscall emulation is enabled by default.  (vsyscalls are
legacy variants on vDSO calls.)  Currently, emulated vsyscalls will
honor seccomp, with a few oddities:

- A return value of ``SECCOMP_RET_TRAP`` will set a ``si_call_addr`` pointing to
  the vsyscall entry for the given call and not the address after the
  'syscall' instruction.  Any code which wants to restart the call
  should be aware that (a) a ret instruction has been emulated and (b)
  trying to resume the syscall will again trigger the standard vsyscall
  emulation security checks, making resuming the syscall mostly
  pointless.

- A return value of ``SECCOMP_RET_TRACE`` will signal the tracer as usual,
  but the syscall may not be changed to another system call using the
  orig_rax register. It may only be changed to -1 order to skip the
  currently emulated call. Any other change MAY terminate the process.
  The rip value seen by the tracer will be the syscall entry address;
  this is different from normal behavior.  The tracer MUST NOT modify
  rip or rsp.  (Do not rely on other changes terminating the process.
  They might work.  For example, on some kernels, choosing a syscall
  that only exists in future kernels will be correctly emulated (by
  returning ``-ENOSYS``).

To detect this quirky behavior, check for ``addr & ~0x0C00 ==
0xFFFFFFFFFF600000``.  (For ``SECCOMP_RET_TRACE``, use rip.  For
``SECCOMP_RET_TRAP``, use ``siginfo->si_call_addr``.)  Do not check any other
condition: future kernels may improve vsyscall emulation and current
kernels in vsyscall=native mode will behave differently, but the
instructions at ``0xF...F600{0,4,8,C}00`` will not be system calls in these
cases.

Note that modern systems are unlikely to use vsyscalls at all -- they
are a legacy feature and they are considerably slower than standard
syscalls.  New code will use the vDSO, and vDSO-issued system calls
are indistinguishable from normal system calls.