요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
.. SPDX-License-Identifier: GPL-2.0
=====================
Introduction of mseal
=====================
:Author: Jeff Xu <jeffxu@chromium.org>
Modern CPUs support memory permissions such as RW and NX bits. The memory
permission feature improves security stance on memory corruption bugs, i.e.
the attacker can’t just write to arbitrary memory and point the code to it,
the memory has to be marked with X bit, or else an exception will happen.
Memory sealing additionally protects the mapping itself against
modifications. This is useful to mitigate memory corruption issues where a
corrupted pointer is passed to a memory management system. For example,
such an attacker primitive can break control-flow integrity guarantees
since read-only memory that is supposed to be trusted can become writable
or .text pages can get remapped. Memory sealing can automatically be
applied by the runtime loader to seal .text and .rodata pages and
applications can additionally seal security critical data at runtime.
A similar feature already exists in the XNU kernel with the
VM_FLAGS_PERMANENT flag [1] and on OpenBSD with the mimmutable syscall [2].
SYSCALL
=======
mseal syscall signature
-----------------------
``int mseal(void *addr, size_t len, unsigned long flags)``
**addr**/**len**: virtual memory address range.
The address range set by **addr**/**len** must meet:
- The start address must be in an allocated VMA.
- The start address must be page aligned.
- The end address (**addr** + **len**) must be in an allocated VMA.
- no gap (unallocated memory) between start and end address.
The ``len`` will be paged aligned implicitly by the kernel.
**flags**: reserved for future use.
**Return values**:
- **0**: Success.
- **-EINVAL**:
* Invalid input ``flags``.
* The start address (``addr``) is not page aligned.
* Address range (``addr`` + ``len``) overflow.
- **-ENOMEM**:
* The start address (``addr``) is not allocated.
* The end address (``addr`` + ``len``) is not allocated.
* A gap (unallocated memory) between start and end address.
- **-EPERM**:
* sealing is supported only on 64-bit CPUs, 32-bit is not supported.
**Note about error return**:
- For above error cases, users can expect the given memory range is
unmodified, i.e. no partial update.
- There might be other internal errors/cases not listed here, e.g.
error during merging/splitting VMAs, or the process reaching the maximum
number of supported VMAs. In those cases, partial updates to the given
memory range could happen. However, those cases should be rare.
**Architecture support**:
mseal only works on 64-bit CPUs, not 32-bit CPUs.
**Idempotent**:
users can call mseal multiple times. mseal on an already sealed memory
is a no-action (not error).
**no munseal**
Once mapping is sealed, it can't be unsealed. The kernel should never
have munseal, this is consistent with other sealing feature, e.g.
F_SEAL_SEAL for file.
Blocked mm syscall for sealed mapping
-------------------------------------
It might be important to note: **once the mapping is sealed, it will
stay in the process's memory until the process terminates**.
Example::
*ptr = mmap(0, 4096, PROT_READ, MAP_ANONYMOUS | MAP_PRIVATE, 0, 0);
rc = mseal(ptr, 4096, 0);
/* munmap will fail */
rc = munmap(ptr, 4096);
assert(rc < 0);
Blocked mm syscall:
- munmap
- mmap
- mremap
- mprotect and pkey_mprotect
- some destructive madvise behaviors: MADV_DONTNEED, MADV_FREE,
MADV_DONTNEED_LOCKED, MADV_FREE, MADV_DONTFORK, MADV_WIPEONFORK
The first set of syscalls to block is munmap, mremap, mmap. They can
either leave an empty space in the address space, therefore allowing
replacement with a new mapping with new set of attributes, or can
overwrite the existing mapping with another mapping.
mprotect and pkey_mprotect are blocked because they changes the
protection bits (RWX) of the mapping.
Certain destructive madvise behaviors, specifically MADV_DONTNEED,
MADV_FREE, MADV_DONTNEED_LOCKED, and MADV_WIPEONFORK, can introduce
risks when applied to anonymous memory by threads lacking write
permissions. Consequently, these operations are prohibited under such
conditions. The aforementioned behaviors have the potential to modify
region contents by discarding pages, effectively performing a memset(0)
operation on the anonymous memory.
Kernel will return -EPERM for blocked syscalls.
When blocked syscall return -EPERM due to sealing, the memory regions may
or may not be changed, depends on the syscall being blocked:
- munmap: munmap is atomic. If one of VMAs in the given range is
sealed, none of VMAs are updated.
- mprotect, pkey_mprotect, madvise: partial update might happen, e.g.
when mprotect over multiple VMAs, mprotect might update the beginning
VMAs before reaching the sealed VMA and return -EPERM.
- mmap and mremap: undefined behavior.
Use cases
=========
- glibc:
The dynamic linker, during loading ELF executables, can apply sealing to
mapping segments.
- Chrome browser: protect some security sensitive data structures.
- System mappings:
The system mappings are created by the kernel and includes vdso, vvar,
vvar_vclock, vectors (arm compat-mode), sigpage (arm compat-mode), uprobes.
Those system mappings are readonly only or execute only, memory sealing can
protect them from ever changing to writable or unmmap/remapped as different
attributes. This is useful to mitigate memory corruption issues where a
corrupted pointer is passed to a memory management system.
If supported by an architecture (CONFIG_ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS),
the CONFIG_MSEAL_SYSTEM_MAPPINGS seals all system mappings of this
architecture.
The following architectures currently support this feature: x86-64, arm64,
loongarch and s390.
WARNING: This feature breaks programs which rely on relocating
or unmapping system mappings. Known broken software at the time
of writing includes CHECKPOINT_RESTORE, UML, gVisor, rr. Therefore
this config can't be enabled universally.
When not to use mseal
=====================
Applications can apply sealing to any virtual memory region from userspace,
but it is *crucial to thoroughly analyze the mapping's lifetime* prior to
apply the sealing. This is because the sealed mapping *won’t be unmapped*
until the process terminates or the exec system call is invoked.
For example:
- aio/shm
aio/shm can call mmap and munmap on behalf of userspace, e.g.
ksys_shmdt() in shm.c. The lifetimes of those mapping are not tied to
the lifetime of the process. If those memories are sealed from userspace,
then munmap will fail, causing leaks in VMA address space during the
lifetime of the process.
- ptr allocated by malloc (heap)
Don't use mseal on the memory ptr return from malloc().
malloc() is implemented by allocator, e.g. by glibc. Heap manager might
allocate a ptr from brk or mapping created by mmap.
If an app calls mseal on a ptr returned from malloc(), this can affect
the heap manager's ability to manage the mappings; the outcome is
non-deterministic.
Example::
ptr = malloc(size);
/* don't call mseal on ptr return from malloc. */
mseal(ptr, size);
/* free will success, allocator can't shrink heap lower than ptr */
free(ptr);
mseal doesn't block
===================
In a nutshell, mseal blocks certain mm syscall from modifying some of VMA's
attributes, such as protection bits (RWX). Sealed mappings doesn't mean the
memory is immutable.
As Jann Horn pointed out in [3], there are still a few ways to write
to RO memory, which is, in a way, by design. And those could be blocked
by different security measures.
Those cases are:
- Write to read-only memory through /proc/self/mem interface (FOLL_FORCE).
- Write to read-only memory through ptrace (such as PTRACE_POKETEXT).
- userfaultfd.
The idea that inspired this patch comes from Stephen Röttger’s work in V8
CFI [4]. Chrome browser in ChromeOS will be the first user of this API.
Reference
=========
- [1] https://github.com/apple-oss-distributions/xnu/blob/1031c584a5e37aff177559b9f69dbd3c8c3fd30a/osfmk/mach/vm_statistics.h#L274
- [2] https://man.openbsd.org/mimmutable.2
- [3] https://lore.kernel.org/lkml/CAG48ez3ShUYey+ZAFsU2i1RpQn0a5eOs2hzQ426FkcgnfUGLvA@mail.gmail.com
- [4] https://docs.google.com/document/d/1O2jwK4dxI3nRcOJuPYkonhTkNQfbmwdvxQMyXgeaRHo/edit#heading=h.bvaojj9fu6hc
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
메모리 봉인의 목적
1-25현대 CPU의 RW 및 NX 권한은 공격자가 임의 메모리에 코드를 쓴 뒤 곧바로 실행하는 일을 어렵게 만듭니다. 코드를 실행하려면 메모리에 X 비트가 있어야 하며, 그렇지 않으면 예외가 발생합니다.
메모리 봉인은 한 단계 더 나아가 매핑 자체가 변경되지 않도록 보호합니다. 손상된 포인터가 메모리 관리 인터페이스에 전달되더라도 신뢰해야 할 읽기 전용 메모리를 쓰기 가능하게 바꾸거나 `.text` 페이지를 다시 매핑하여 제어 흐름 무결성을 깨뜨리는 공격을 줄일 수 있습니다.
런타임 로더는 `.text`와 `.rodata` 페이지를 자동으로 봉인할 수 있고, 애플리케이션은 실행 중 보안상 중요한 데이터를 추가로 봉인할 수 있습니다. 유사한 기능으로 XNU의 `VM_FLAGS_PERMANENT`와 OpenBSD의 `mimmutable` 시스템 호출이 있습니다.
.. SPDX-License-Identifier: GPL-2.0
=====================
Introduction of mseal
=====================
:Author: Jeff Xu <jeffxu@chromium.org>
Modern CPUs support memory permissions such as RW and NX bits. The memory
permission feature improves security stance on memory corruption bugs, i.e.
the attacker can’t just write to arbitrary memory and point the code to it,
the memory has to be marked with X bit, or else an exception will happen.
Memory sealing additionally protects the mapping itself against
modifications. This is useful to mitigate memory corruption issues where a
corrupted pointer is passed to a memory management system. For example,
such an attacker primitive can break control-flow integrity guarantees
since read-only memory that is supposed to be trusted can become writable
or .text pages can get remapped. Memory sealing can automatically be
applied by the runtime loader to seal .text and .rodata pages and
applications can additionally seal security critical data at runtime.
A similar feature already exists in the XNU kernel with the
VM_FLAGS_PERMANENT flag [1] and on OpenBSD with the mimmutable syscall [2].
시스템 호출 계약과 오류
26-75`mseal(void *addr, size_t len, unsigned long flags)`은 `addr`부터 `len` 길이의 가상 메모리 범위를 봉인합니다. 시작 주소는 할당된 VMA 안에 있고 페이지 경계에 맞아야 하며, 끝 주소도 할당된 VMA 안에 있어야 합니다. 시작과 끝 사이에 할당되지 않은 틈이 있어서는 안 됩니다. `len`은 커널이 암묵적으로 페이지 경계에 맞춥니다.
입력 범위와 아키텍처 조건에 따른 결과입니다.
나열된 입력 오류에서는 범위가 부분적으로 바뀌지 않았다고 기대할 수 있습니다. 다만 VMA 병합·분할 오류나 프로세스가 지원 가능한 최대 VMA 수에 도달하는 등의 드문 내부 오류에서는 일부만 갱신될 수 있습니다.
호출자가 수명과 되돌릴 수 없음까지 고려해야 하는 계약입니다.
SYSCALL
=======
mseal syscall signature
-----------------------
``int mseal(void *addr, size_t len, unsigned long flags)``
**addr**/**len**: virtual memory address range.
The address range set by **addr**/**len** must meet:
- The start address must be in an allocated VMA.
- The start address must be page aligned.
- The end address (**addr** + **len**) must be in an allocated VMA.
- no gap (unallocated memory) between start and end address.
The ``len`` will be paged aligned implicitly by the kernel.
**flags**: reserved for future use.
**Return values**:
- **0**: Success.
- **-EINVAL**:
* Invalid input ``flags``.
* The start address (``addr``) is not page aligned.
* Address range (``addr`` + ``len``) overflow.
- **-ENOMEM**:
* The start address (``addr``) is not allocated.
* The end address (``addr`` + ``len``) is not allocated.
* A gap (unallocated memory) between start and end address.
- **-EPERM**:
* sealing is supported only on 64-bit CPUs, 32-bit is not supported.
**Note about error return**:
- For above error cases, users can expect the given memory range is
unmodified, i.e. no partial update.
- There might be other internal errors/cases not listed here, e.g.
error during merging/splitting VMAs, or the process reaching the maximum
number of supported VMAs. In those cases, partial updates to the given
memory range could happen. However, those cases should be rare.
**Architecture support**:
mseal only works on 64-bit CPUs, not 32-bit CPUs.
**Idempotent**:
users can call mseal multiple times. mseal on an already sealed memory
is a no-action (not error).
**no munseal**
Once mapping is sealed, it can't be unsealed. The kernel should never
have munseal, this is consistent with other sealing feature, e.g.
F_SEAL_SEAL for file.
차단되는 메모리 관리 작업
76-124봉인된 매핑은 프로세스가 종료될 때까지 메모리에 남습니다. 예제처럼 `mmap()`으로 만든 범위를 `mseal()`한 뒤 `munmap()`하면 해제 호출은 실패합니다.
주소 공간의 교체, 권한 변경, 내용 파괴를 막습니다.
차단된 호출은 `-EPERM`을 반환합니다. 그러나 반환 시 전체 범위의 원자성은 호출마다 다르므로, 실패했다는 사실만으로 모든 VMA가 원래 상태라고 가정하면 안 됩니다.
봉인된 VMA를 만났을 때 앞부분이 이미 바뀌었을 가능성을 구분합니다.
Blocked mm syscall for sealed mapping
-------------------------------------
It might be important to note: **once the mapping is sealed, it will
stay in the process's memory until the process terminates**.
Example::
*ptr = mmap(0, 4096, PROT_READ, MAP_ANONYMOUS | MAP_PRIVATE, 0, 0);
rc = mseal(ptr, 4096, 0);
/* munmap will fail */
rc = munmap(ptr, 4096);
assert(rc < 0);
Blocked mm syscall:
- munmap
- mmap
- mremap
- mprotect and pkey_mprotect
- some destructive madvise behaviors: MADV_DONTNEED, MADV_FREE,
MADV_DONTNEED_LOCKED, MADV_FREE, MADV_DONTFORK, MADV_WIPEONFORK
The first set of syscalls to block is munmap, mremap, mmap. They can
either leave an empty space in the address space, therefore allowing
replacement with a new mapping with new set of attributes, or can
overwrite the existing mapping with another mapping.
mprotect and pkey_mprotect are blocked because they changes the
protection bits (RWX) of the mapping.
Certain destructive madvise behaviors, specifically MADV_DONTNEED,
MADV_FREE, MADV_DONTNEED_LOCKED, and MADV_WIPEONFORK, can introduce
risks when applied to anonymous memory by threads lacking write
permissions. Consequently, these operations are prohibited under such
conditions. The aforementioned behaviors have the potential to modify
region contents by discarding pages, effectively performing a memset(0)
operation on the anonymous memory.
Kernel will return -EPERM for blocked syscalls.
When blocked syscall return -EPERM due to sealing, the memory regions may
or may not be changed, depends on the syscall being blocked:
- munmap: munmap is atomic. If one of VMAs in the given range is
sealed, none of VMAs are updated.
- mprotect, pkey_mprotect, madvise: partial update might happen, e.g.
when mprotect over multiple VMAs, mprotect might update the beginning
VMAs before reaching the sealed VMA and return -EPERM.
- mmap and mremap: undefined behavior.
사용 사례와 시스템 매핑
125-153glibc 동적 링커는 ELF 실행 파일을 적재할 때 매핑 세그먼트를 봉인할 수 있고, Chrome 브라우저는 보안에 민감한 데이터 구조를 보호할 수 있습니다.
커널이 만드는 시스템 매핑에는 `vdso`, `vvar`, `vvar_vclock`, ARM 호환 모드의 `vectors`와 `sigpage`, `uprobes`가 포함됩니다. 읽기 전용 또는 실행 전용인 이 매핑을 봉인하면 쓰기 가능 상태로 바꾸거나 다른 속성으로 해제·재매핑하는 일을 막을 수 있습니다.
아키텍처 지원과 실제 활성화를 구분합니다.
이 설정은 시스템 매핑을 옮기거나 해제하는 프로그램을 깨뜨립니다. 문서 작성 시점의 알려진 영향 대상은 `CHECKPOINT_RESTORE`, UML, gVisor, rr이므로 모든 구성에서 보편적으로 켤 수 없습니다.
Use cases
=========
- glibc:
The dynamic linker, during loading ELF executables, can apply sealing to
mapping segments.
- Chrome browser: protect some security sensitive data structures.
- System mappings:
The system mappings are created by the kernel and includes vdso, vvar,
vvar_vclock, vectors (arm compat-mode), sigpage (arm compat-mode), uprobes.
Those system mappings are readonly only or execute only, memory sealing can
protect them from ever changing to writable or unmmap/remapped as different
attributes. This is useful to mitigate memory corruption issues where a
corrupted pointer is passed to a memory management system.
If supported by an architecture (CONFIG_ARCH_SUPPORTS_MSEAL_SYSTEM_MAPPINGS),
the CONFIG_MSEAL_SYSTEM_MAPPINGS seals all system mappings of this
architecture.
The following architectures currently support this feature: x86-64, arm64,
loongarch and s390.
WARNING: This feature breaks programs which rely on relocating
or unmapping system mappings. Known broken software at the time
of writing includes CHECKPOINT_RESTORE, UML, gVisor, rr. Therefore
this config can't be enabled universally.
mseal을 사용하지 말아야 할 경우
154-184사용자 공간은 어떤 가상 메모리 영역에도 sealing을 적용할 수 있지만, 먼저 매핑의 수명을 철저히 분석해야 합니다. 봉인된 매핑은 프로세스 종료 또는 `exec` 시스템 호출 전에는 해제되지 않습니다.
프로세스 수명과 매핑 관리자의 기대를 깨뜨리는 대표 사례입니다.
heap 포인터를 봉인한 뒤 `free()` 자체는 성공할 수 있지만 allocator는 그 포인터보다 아래로 heap을 축소하지 못할 수 있습니다. 따라서 할당기 소유 메모리의 일부 주소만 보고 `mseal()`을 호출해서는 안 됩니다.
When not to use mseal
=====================
Applications can apply sealing to any virtual memory region from userspace,
but it is *crucial to thoroughly analyze the mapping's lifetime* prior to
apply the sealing. This is because the sealed mapping *won’t be unmapped*
until the process terminates or the exec system call is invoked.
For example:
- aio/shm
aio/shm can call mmap and munmap on behalf of userspace, e.g.
ksys_shmdt() in shm.c. The lifetimes of those mapping are not tied to
the lifetime of the process. If those memories are sealed from userspace,
then munmap will fail, causing leaks in VMA address space during the
lifetime of the process.
- ptr allocated by malloc (heap)
Don't use mseal on the memory ptr return from malloc().
malloc() is implemented by allocator, e.g. by glibc. Heap manager might
allocate a ptr from brk or mapping created by mmap.
If an app calls mseal on a ptr returned from malloc(), this can affect
the heap manager's ability to manage the mappings; the outcome is
non-deterministic.
Example::
ptr = malloc(size);
/* don't call mseal on ptr return from malloc. */
mseal(ptr, size);
/* free will success, allocator can't shrink heap lower than ptr */
free(ptr);
mseal이 막지 않는 쓰기
185-203`mseal`은 일부 mm 시스템 호출이 보호 비트 같은 VMA 속성을 바꾸는 것을 막을 뿐이며, 봉인된 메모리가 불변이라는 뜻은 아닙니다. 다른 설계상 경로는 별도의 보안 수단으로 통제해야 합니다.
읽기 전용 메모리에 영향을 줄 수 있지만 mseal의 차단 범위 밖에 있습니다.
이 API의 발상은 Stephen Röttger의 V8 CFI 작업에서 왔으며, ChromeOS의 Chrome 브라우저가 첫 사용자가 될 예정이라고 문서는 설명합니다.
mseal doesn't block
===================
In a nutshell, mseal blocks certain mm syscall from modifying some of VMA's
attributes, such as protection bits (RWX). Sealed mappings doesn't mean the
memory is immutable.
As Jann Horn pointed out in [3], there are still a few ways to write
to RO memory, which is, in a way, by design. And those could be blocked
by different security measures.
Those cases are:
- Write to read-only memory through /proc/self/mem interface (FOLL_FORCE).
- Write to read-only memory through ptrace (such as PTRACE_POKETEXT).
- userfaultfd.
The idea that inspired this patch comes from Stephen Röttger’s work in V8
CFI [4]. Chrome browser in ChromeOS will be the first user of this API.
참고 자료
204-209참고 자료는 XNU의 영구 매핑 플래그, OpenBSD `mimmutable(2)`, 읽기 전용 메모리 쓰기 경로에 관한 LKML 논의, V8 CFI 설계 문서로 이어집니다.
Reference
=========
- [1] https://github.com/apple-oss-distributions/xnu/blob/1031c584a5e37aff177559b9f69dbd3c8c3fd30a/osfmk/mach/vm_statistics.h#L274
- [2] https://man.openbsd.org/mimmutable.2
- [3] https://lore.kernel.org/lkml/CAG48ez3ShUYey+ZAFsU2i1RpQn0a5eOs2hzQ426FkcgnfUGLvA@mail.gmail.com
- [4] https://docs.google.com/document/d/1O2jwK4dxI3nRcOJuPYkonhTkNQfbmwdvxQMyXgeaRHo/edit#heading=h.bvaojj9fu6hc
요약·해설
mseal.rst:1-209`mseal`은 메모리 내용 자체를 불변으로 만드는 장치가 아니라 특정 mm 시스템 호출이 매핑 속성과 수명을 바꾸는 일을 차단합니다. 봉인 후 해제할 수 없으므로 보호 가치뿐 아니라 매핑 소유자와 전체 수명을 먼저 확인하는 것이 핵심입니다.