요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
===================================
Documentation for /proc/sys/kernel/
===================================
.. See scripts/check-sysctl-docs to keep this up to date
Copyright (c) 1998, 1999, Rik van Riel <riel@nl.linux.org>
Copyright (c) 2009, Shen Feng<shen@cn.fujitsu.com>
For general info and legal blurb, please look in
Documentation/admin-guide/sysctl/index.rst.
------------------------------------------------------------------------------
This file contains documentation for the sysctl files in
``/proc/sys/kernel/``.
The files in this directory can be used to tune and monitor
miscellaneous and general things in the operation of the Linux
kernel. Since some of the files *can* be used to screw up your
system, it is advisable to read both documentation and source
before actually making adjustments.
Currently, these files might (depending on your configuration)
show up in ``/proc/sys/kernel``:
.. contents:: :local:
acct
====
::
highwater lowwater frequency
If BSD-style process accounting is enabled these values control
its behaviour. If free space on filesystem where the log lives
goes below ``lowwater``\ % accounting suspends. If free space gets
above ``highwater``\ % accounting resumes. ``frequency`` determines
how often do we check the amount of free space (value is in
seconds). Default:
::
4 2 30
That is, suspend accounting if free space drops below 2%; resume it
if it increases to at least 4%; consider information about amount of
free space valid for 30 seconds.
acpi_video_flags
================
See Documentation/power/video.rst. This allows the video resume mode to be set,
in a similar fashion to the ``acpi_sleep`` kernel parameter, by
combining the following values:
= =======
1 s3_bios
2 s3_mode
4 s3_beep
= =======
arch
====
The machine hardware name, the same output as ``uname -m``
(e.g. ``x86_64`` or ``aarch64``).
auto_msgmni
===========
This variable has no effect and may be removed in future kernel
releases. Reading it always returns 0.
Up to Linux 3.17, it enabled/disabled automatic recomputing of
`msgmni`_
upon memory add/remove or upon IPC namespace creation/removal.
Echoing "1" into this file enabled msgmni automatic recomputing.
Echoing "0" turned it off. The default value was 1.
bootloader_type (x86 only)
==========================
This gives the bootloader type number as indicated by the bootloader,
shifted left by 4, and OR'd with the low four bits of the bootloader
version. The reason for this encoding is that this used to match the
``type_of_loader`` field in the kernel header; the encoding is kept for
backwards compatibility. That is, if the full bootloader type number
is 0x15 and the full version number is 0x234, this file will contain
the value 340 = 0x154.
See the ``type_of_loader`` and ``ext_loader_type`` fields in
Documentation/arch/x86/boot.rst for additional information.
bootloader_version (x86 only)
=============================
The complete bootloader version number. In the example above, this
file will contain the value 564 = 0x234.
See the ``type_of_loader`` and ``ext_loader_ver`` fields in
Documentation/arch/x86/boot.rst for additional information.
bpf_stats_enabled
=================
Controls whether the kernel should collect statistics on BPF programs
(total time spent running, number of times run...). Enabling
statistics causes a slight reduction in performance on each program
run. The statistics can be seen using ``bpftool``.
= ===================================
0 Don't collect statistics (default).
1 Collect statistics.
= ===================================
cad_pid
=======
This is the pid which will be signalled on reboot (notably, by
Ctrl-Alt-Delete). Writing a value to this file which doesn't
correspond to a running process will result in ``-ESRCH``.
See also `ctrl-alt-del`_.
cap_last_cap
============
Highest valid capability of the running kernel. Exports
``CAP_LAST_CAP`` from the kernel.
.. _core_pattern:
core_pattern
============
``core_pattern`` is used to specify a core dumpfile pattern name.
* max length 127 characters; default value is "core"
* ``core_pattern`` is used as a pattern template for the output
filename; certain string patterns (beginning with '%') are
substituted with their actual values.
* backward compatibility with ``core_uses_pid``:
If ``core_pattern`` does not include "%p" (default does not)
and ``core_uses_pid`` is set, then .PID will be appended to
the filename.
* corename format specifiers
======== ==========================================
%<NUL> '%' is dropped
%% output one '%'
%p pid
%P global pid (init PID namespace)
%i tid
%I global tid (init PID namespace)
%u uid (in initial user namespace)
%g gid (in initial user namespace)
%d dump mode, matches ``PR_SET_DUMPABLE`` and
``/proc/sys/fs/suid_dumpable``
%s signal number
%t UNIX time of dump
%h hostname
%e executable filename (may be shortened, could be changed by prctl etc)
%f executable filename
%E executable path
%c maximum size of core file by resource limit RLIMIT_CORE
%C CPU the task ran on
%F pidfd number
%<OTHER> both are dropped
======== ==========================================
* If the first character of the pattern is a '|', the kernel will treat
the rest of the pattern as a command to run. The core dump will be
written to the standard input of that program instead of to a file.
core_pipe_limit
===============
This sysctl is only applicable when `core_pattern`_ is configured to
pipe core files to a user space helper (when the first character of
``core_pattern`` is a '|', see above).
When collecting cores via a pipe to an application, it is occasionally
useful for the collecting application to gather data about the
crashing process from its ``/proc/pid`` directory.
In order to do this safely, the kernel must wait for the collecting
process to exit, so as not to remove the crashing processes proc files
prematurely.
This in turn creates the possibility that a misbehaving userspace
collecting process can block the reaping of a crashed process simply
by never exiting.
This sysctl defends against that.
It defines how many concurrent crashing processes may be piped to user
space applications in parallel.
If this value is exceeded, then those crashing processes above that
value are noted via the kernel log and their cores are skipped.
0 is a special value, indicating that unlimited processes may be
captured in parallel, but that no waiting will take place (i.e. the
collecting process is not guaranteed access to ``/proc/<crashing
pid>/``).
This value defaults to 0.
core_sort_vma
=============
The default coredump writes VMAs in address order. By setting
``core_sort_vma`` to 1, VMAs will be written from smallest size
to largest size. This is known to break at least elfutils, but
can be handy when dealing with very large (and truncated)
coredumps where the more useful debugging details are included
in the smaller VMAs.
core_uses_pid
=============
The default coredump filename is "core". By setting
``core_uses_pid`` to 1, the coredump filename becomes core.PID.
If `core_pattern`_ does not include "%p" (default does not)
and ``core_uses_pid`` is set, then .PID will be appended to
the filename.
ctrl-alt-del
============
When the value in this file is 0, ctrl-alt-del is trapped and
sent to the ``init(1)`` program to handle a graceful restart.
When, however, the value is > 0, Linux's reaction to a Vulcan
Nerve Pinch (tm) will be an immediate reboot, without even
syncing its dirty buffers.
Note:
when a program (like dosemu) has the keyboard in 'raw'
mode, the ctrl-alt-del is intercepted by the program before it
ever reaches the kernel tty layer, and it's up to the program
to decide what to do with it.
dmesg_restrict
==============
This toggle indicates whether unprivileged users are prevented
from using ``dmesg(8)`` to view messages from the kernel's log
buffer.
When ``dmesg_restrict`` is set to 0 there are no restrictions.
When ``dmesg_restrict`` is set to 1, users must have
``CAP_SYSLOG`` to use ``dmesg(8)``.
The kernel config option ``CONFIG_SECURITY_DMESG_RESTRICT`` sets the
default value of ``dmesg_restrict``.
domainname & hostname
=====================
These files can be used to set the NIS/YP domainname and the
hostname of your box in exactly the same way as the commands
domainname and hostname, i.e.::
# echo "darkstar" > /proc/sys/kernel/hostname
# echo "mydomain" > /proc/sys/kernel/domainname
has the same effect as::
# hostname "darkstar"
# domainname "mydomain"
Note, however, that the classic darkstar.frop.org has the
hostname "darkstar" and DNS (Internet Domain Name Server)
domainname "frop.org", not to be confused with the NIS (Network
Information Service) or YP (Yellow Pages) domainname. These two
domain names are in general different. For a detailed discussion
see the ``hostname(1)`` man page.
firmware_config
===============
See Documentation/driver-api/firmware/fallback-mechanisms.rst.
The entries in this directory allow the firmware loader helper
fallback to be controlled:
* ``force_sysfs_fallback``, when set to 1, forces the use of the
fallback;
* ``ignore_sysfs_fallback``, when set to 1, ignores any fallback.
ftrace_dump_on_oops
===================
Determines whether ``ftrace_dump()`` should be called on an oops (or
kernel panic). This will output the contents of the ftrace buffers to
the console. This is very useful for capturing traces that lead to
crashes and outputting them to a serial console.
======================= ===========================================
0 Disabled (default).
1 Dump buffers of all CPUs.
2(orig_cpu) Dump the buffer of the CPU that triggered the
oops.
<instance> Dump the specific instance buffer on all CPUs.
<instance>=2(orig_cpu) Dump the specific instance buffer on the CPU
that triggered the oops.
======================= ===========================================
Multiple instance dump is also supported, and instances are separated
by commas. If global buffer also needs to be dumped, please specify
the dump mode (1/2/orig_cpu) first for global buffer.
So for example to dump "foo" and "bar" instance buffer on all CPUs,
user can::
echo "foo,bar" > /proc/sys/kernel/ftrace_dump_on_oops
To dump global buffer and "foo" instance buffer on all
CPUs along with the "bar" instance buffer on CPU that triggered the
oops, user can::
echo "1,foo,bar=2" > /proc/sys/kernel/ftrace_dump_on_oops
ftrace_enabled, stack_tracer_enabled
====================================
See Documentation/trace/ftrace.rst.
hardlockup_all_cpu_backtrace
============================
This value controls the hard lockup detector behavior when a hard
lockup condition is detected as to whether or not to gather further
debug information. If enabled, arch-specific all-CPU stack dumping
will be initiated.
= ============================================
0 Do nothing. This is the default behavior.
1 On detection capture more debug information.
= ============================================
hardlockup_panic
================
This parameter can be used to control whether the kernel panics
when a hard lockup is detected.
= ===========================
0 Don't panic on hard lockup.
1 Panic on hard lockup.
= ===========================
See Documentation/admin-guide/lockup-watchdogs.rst for more information.
This can also be set using the nmi_watchdog kernel parameter.
hotplug
=======
Path for the hotplug policy agent.
Default value is ``CONFIG_UEVENT_HELPER_PATH``, which in turn defaults
to the empty string.
This file only exists when ``CONFIG_UEVENT_HELPER`` is enabled. Most
modern systems rely exclusively on the netlink-based uevent source and
don't need this.
hung_task_all_cpu_backtrace
===========================
If this option is set, the kernel will send an NMI to all CPUs to dump
their backtraces when a hung task is detected. This file shows up if
CONFIG_DETECT_HUNG_TASK and CONFIG_SMP are enabled.
0: Won't show all CPUs backtraces when a hung task is detected.
This is the default behavior.
1: Will non-maskably interrupt all CPUs and dump their backtraces when
a hung task is detected.
hung_task_panic
===============
Controls the kernel's behavior when a hung task is detected.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
= =================================================
0 Continue operation. This is the default behavior.
1 Panic immediately.
= =================================================
hung_task_check_count
=====================
The upper bound on the number of tasks that are checked.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
hung_task_detect_count
======================
Indicates the total number of tasks that have been detected as hung since
the system boot.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
hung_task_timeout_secs
======================
When a task in D state did not get scheduled
for more than this value report a warning.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
0 means infinite timeout, no checking is done.
Possible values to set are in range {0:``LONG_MAX``/``HZ``}.
hung_task_check_interval_secs
=============================
Hung task check interval. If hung task checking is enabled
(see `hung_task_timeout_secs`_), the check is done every
``hung_task_check_interval_secs`` seconds.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
0 (default) means use ``hung_task_timeout_secs`` as checking
interval.
Possible values to set are in range {0:``LONG_MAX``/``HZ``}.
hung_task_warnings
==================
The maximum number of warnings to report. During a check interval
if a hung task is detected, this value is decreased by 1.
When this value reaches 0, no more warnings will be reported.
This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled.
-1: report an infinite number of warnings.
hyperv_record_panic_msg
=======================
Controls whether the panic kmsg data should be reported to Hyper-V.
= =========================================================
0 Do not report panic kmsg data.
1 Report the panic kmsg data. This is the default behavior.
= =========================================================
ignore-unaligned-usertrap
=========================
On architectures where unaligned accesses cause traps, and where this
feature is supported (``CONFIG_SYSCTL_ARCH_UNALIGN_NO_WARN``;
currently, ``arc``, ``parisc`` and ``loongarch``), controls whether all
unaligned traps are logged.
= =============================================================
0 Log all unaligned accesses.
1 Only warn the first time a process traps. This is the default
setting.
= =============================================================
See also `unaligned-trap`_.
io_uring_disabled
=================
Prevents all processes from creating new io_uring instances. Enabling this
shrinks the kernel's attack surface.
= ======================================================================
0 All processes can create io_uring instances as normal. This is the
default setting.
1 io_uring creation is disabled (io_uring_setup() will fail with
-EPERM) for unprivileged processes not in the io_uring_group group.
Existing io_uring instances can still be used. See the
documentation for io_uring_group for more information.
2 io_uring creation is disabled for all processes. io_uring_setup()
always fails with -EPERM. Existing io_uring instances can still be
used.
= ======================================================================
io_uring_group
==============
When io_uring_disabled is set to 1, a process must either be
privileged (CAP_SYS_ADMIN) or be in the io_uring_group group in order
to create an io_uring instance. If io_uring_group is set to -1 (the
default), only processes with the CAP_SYS_ADMIN capability may create
io_uring instances.
kexec_load_disabled
===================
A toggle indicating if the syscalls ``kexec_load`` and
``kexec_file_load`` have been disabled.
This value defaults to 0 (false: ``kexec_*load`` enabled), but can be
set to 1 (true: ``kexec_*load`` disabled).
Once true, kexec can no longer be used, and the toggle cannot be set
back to false.
This allows a kexec image to be loaded before disabling the syscall,
allowing a system to set up (and later use) an image without it being
altered.
Generally used together with the `modules_disabled`_ sysctl.
kexec_load_limit_panic
======================
This parameter specifies a limit to the number of times the syscalls
``kexec_load`` and ``kexec_file_load`` can be called with a crash
image. It can only be set with a more restrictive value than the
current one.
== ======================================================
-1 Unlimited calls to kexec. This is the default setting.
N Number of calls left.
== ======================================================
kexec_load_limit_reboot
=======================
Similar functionality as ``kexec_load_limit_panic``, but for a normal
image.
kptr_restrict
=============
This toggle indicates whether restrictions are placed on
exposing kernel addresses via ``/proc`` and other interfaces.
When ``kptr_restrict`` is set to 0 (the default) the address is hashed
before printing.
(This is the equivalent to %p.)
When ``kptr_restrict`` is set to 1, kernel pointers printed using the
%pK format specifier will be replaced with 0s unless the user has
``CAP_SYSLOG`` and effective user and group ids are equal to the real
ids.
This is because %pK checks are done at read() time rather than open()
time, so if permissions are elevated between the open() and the read()
(e.g via a setuid binary) then %pK will not leak kernel pointers to
unprivileged users.
Note, this is a temporary solution only.
The correct long-term solution is to do the permission checks at
open() time.
Consider removing world read permissions from files that use %pK, and
using `dmesg_restrict`_ to protect against uses of %pK in ``dmesg(8)``
if leaking kernel pointer values to unprivileged users is a concern.
When ``kptr_restrict`` is set to 2, kernel pointers printed using
%pK will be replaced with 0s regardless of privileges.
modprobe
========
The full path to the usermode helper for autoloading kernel modules,
by default ``CONFIG_MODPROBE_PATH``, which in turn defaults to
"/sbin/modprobe". This binary is executed when the kernel requests a
module. For example, if userspace passes an unknown filesystem type
to mount(), then the kernel will automatically request the
corresponding filesystem module by executing this usermode helper.
This usermode helper should insert the needed module into the kernel.
This sysctl only affects module autoloading. It has no effect on the
ability to explicitly insert modules.
This sysctl can be used to debug module loading requests::
echo '#! /bin/sh' > /tmp/modprobe
echo 'echo "$@" >> /tmp/modprobe.log' >> /tmp/modprobe
echo 'exec /sbin/modprobe "$@"' >> /tmp/modprobe
chmod a+x /tmp/modprobe
echo /tmp/modprobe > /proc/sys/kernel/modprobe
Alternatively, if this sysctl is set to the empty string, then module
autoloading is completely disabled. The kernel will not try to
execute a usermode helper at all, nor will it call the
kernel_module_request LSM hook.
If CONFIG_STATIC_USERMODEHELPER=y is set in the kernel configuration,
then the configured static usermode helper overrides this sysctl,
except that the empty string is still accepted to completely disable
module autoloading as described above.
modules_disabled
================
A toggle value indicating if modules are allowed to be loaded
in an otherwise modular kernel. This toggle defaults to off
(0), but can be set true (1). Once true, modules can be
neither loaded nor unloaded, and the toggle cannot be set back
to false. Generally used with the `kexec_load_disabled`_ toggle.
.. _msgmni:
msgmax, msgmnb, and msgmni
==========================
``msgmax`` is the maximum size of an IPC message, in bytes. 8192 by
default (``MSGMAX``).
``msgmnb`` is the maximum size of an IPC queue, in bytes. 16384 by
default (``MSGMNB``).
``msgmni`` is the maximum number of IPC queues. 32000 by default
(``MSGMNI``).
All of these parameters are set per ipc namespace. The maximum number of bytes
in POSIX message queues is limited by ``RLIMIT_MSGQUEUE``. This limit is
respected hierarchically in the each user namespace.
msg_next_id, sem_next_id, and shm_next_id (System V IPC)
========================================================
These three toggles allows to specify desired id for next allocated IPC
object: message, semaphore or shared memory respectively.
By default they are equal to -1, which means generic allocation logic.
Possible values to set are in range {0:``INT_MAX``}.
Notes:
1) kernel doesn't guarantee, that new object will have desired id. So,
it's up to userspace, how to handle an object with "wrong" id.
2) Toggle with non-default value will be set back to -1 by kernel after
successful IPC object allocation. If an IPC object allocation syscall
fails, it is undefined if the value remains unmodified or is reset to -1.
ngroups_max
===========
Maximum number of supplementary groups, _i.e._ the maximum size which
``setgroups`` will accept. Exports ``NGROUPS_MAX`` from the kernel.
nmi_watchdog
============
This parameter can be used to control the NMI watchdog
(i.e. the hard lockup detector) on x86 systems.
= =================================
0 Disable the hard lockup detector.
1 Enable the hard lockup detector.
= =================================
The hard lockup detector monitors each CPU for its ability to respond to
timer interrupts. The mechanism utilizes CPU performance counter registers
that are programmed to generate Non-Maskable Interrupts (NMIs) periodically
while a CPU is busy. Hence, the alternative name 'NMI watchdog'.
The NMI watchdog is disabled by default if the kernel is running as a guest
in a KVM virtual machine. This default can be overridden by adding::
nmi_watchdog=1
to the guest kernel command line (see
Documentation/admin-guide/kernel-parameters.rst).
nmi_wd_lpm_factor (PPC only)
============================
Factor to apply to the NMI watchdog timeout (only when ``nmi_watchdog`` is
set to 1). This factor represents the percentage added to
``watchdog_thresh`` when calculating the NMI watchdog timeout during an
LPM. The soft lockup timeout is not impacted.
A value of 0 means no change. The default value is 200 meaning the NMI
watchdog is set to 30s (based on ``watchdog_thresh`` equal to 10).
numa_balancing
==============
Enables/disables and configures automatic page fault based NUMA memory
balancing. Memory is moved automatically to nodes that access it often.
The value to set can be the result of ORing the following:
= =================================
0 NUMA_BALANCING_DISABLED
1 NUMA_BALANCING_NORMAL
2 NUMA_BALANCING_MEMORY_TIERING
= =================================
Or NUMA_BALANCING_NORMAL to optimize page placement among different
NUMA nodes to reduce remote accessing. On NUMA machines, there is a
performance penalty if remote memory is accessed by a CPU. When this
feature is enabled the kernel samples what task thread is accessing
memory by periodically unmapping pages and later trapping a page
fault. At the time of the page fault, it is determined if the data
being accessed should be migrated to a local memory node.
The unmapping of pages and trapping faults incur additional overhead that
ideally is offset by improved memory locality but there is no universal
guarantee. If the target workload is already bound to NUMA nodes then this
feature should be disabled.
Or NUMA_BALANCING_MEMORY_TIERING to optimize page placement among
different types of memory (represented as different NUMA nodes) to
place the hot pages in the fast memory. This is implemented based on
unmapping and page fault too.
numa_balancing_promote_rate_limit_MBps
======================================
Too high promotion/demotion throughput between different memory types
may hurt application latency. This can be used to rate limit the
promotion throughput. The per-node max promotion throughput in MB/s
will be limited to be no more than the set value.
A rule of thumb is to set this to less than 1/10 of the PMEM node
write bandwidth.
oops_all_cpu_backtrace
======================
If this option is set, the kernel will send an NMI to all CPUs to dump
their backtraces when an oops event occurs. It should be used as a last
resort in case a panic cannot be triggered (to protect VMs running, for
example) or kdump can't be collected. This file shows up if CONFIG_SMP
is enabled.
0: Won't show all CPUs backtraces when an oops is detected.
This is the default behavior.
1: Will non-maskably interrupt all CPUs and dump their backtraces when
an oops event is detected.
oops_limit
==========
Number of kernel oopses after which the kernel should panic when
``panic_on_oops`` is not set. Setting this to 0 disables checking
the count. Setting this to 1 has the same effect as setting
``panic_on_oops=1``. The default value is 10000.
osrelease, ostype & version
===========================
::
# cat osrelease
2.1.88
# cat ostype
Linux
# cat version
#5 Wed Feb 25 21:49:24 MET 1998
The files ``osrelease`` and ``ostype`` should be clear enough.
``version``
needs a little more clarification however. The '#5' means that
this is the fifth kernel built from this source base and the
date behind it indicates the time the kernel was built.
The only way to tune these values is to rebuild the kernel :-)
overflowgid & overflowuid
=========================
if your architecture did not always support 32-bit UIDs (i.e. arm,
i386, m68k, sh, and sparc32), a fixed UID and GID will be returned to
applications that use the old 16-bit UID/GID system calls, if the
actual UID or GID would exceed 65535.
These sysctls allow you to change the value of the fixed UID and GID.
The default is 65534.
panic
=====
The value in this file determines the behaviour of the kernel on a
panic:
* if zero, the kernel will loop forever;
* if negative, the kernel will reboot immediately;
* if positive, the kernel will reboot after the corresponding number
of seconds.
When you use the software watchdog, the recommended setting is 60.
panic_on_io_nmi
===============
Controls the kernel's behavior when a CPU receives an NMI caused by
an IO error.
= ==================================================================
0 Try to continue operation (default).
1 Panic immediately. The IO error triggered an NMI. This indicates a
serious system condition which could result in IO data corruption.
Rather than continuing, panicking might be a better choice. Some
servers issue this sort of NMI when the dump button is pushed,
and you can use this option to take a crash dump.
= ==================================================================
panic_on_oops
=============
Controls the kernel's behaviour when an oops or BUG is encountered.
= ===================================================================
0 Try to continue operation.
1 Panic immediately. If the `panic` sysctl is also non-zero then the
machine will be rebooted.
= ===================================================================
panic_on_stackoverflow
======================
Controls the kernel's behavior when detecting the overflows of
kernel, IRQ and exception stacks except a user stack.
This file shows up if ``CONFIG_DEBUG_STACKOVERFLOW`` is enabled.
= ==========================
0 Try to continue operation.
1 Panic immediately.
= ==========================
panic_on_unrecovered_nmi
========================
The default Linux behaviour on an NMI of either memory or unknown is
to continue operation. For many environments such as scientific
computing it is preferable that the box is taken out and the error
dealt with than an uncorrected parity/ECC error get propagated.
A small number of systems do generate NMIs for bizarre random reasons
such as power management so the default is off. That sysctl works like
the existing panic controls already in that directory.
panic_on_warn
=============
Calls panic() in the WARN() path when set to 1. This is useful to avoid
a kernel rebuild when attempting to kdump at the location of a WARN().
= ================================================
0 Only WARN(), default behaviour.
1 Call panic() after printing out WARN() location.
= ================================================
panic_print
===========
Bitmask for printing system info when panic happens. User can chose
combination of the following bits:
===== ============================================
bit 0 print all tasks info
bit 1 print system memory info
bit 2 print timer info
bit 3 print locks info if ``CONFIG_LOCKDEP`` is on
bit 4 print ftrace buffer
bit 5 replay all kernel messages on consoles at the end of panic
bit 6 print all CPUs backtrace (if available in the arch)
bit 7 print only tasks in uninterruptible (blocked) state
===== ============================================
So for example to print tasks and memory info on panic, user can::
echo 3 > /proc/sys/kernel/panic_print
panic_sys_info
==============
A comma separated list of extra information to be dumped on panic,
for example, "tasks,mem,timers,...". It is a human readable alternative
to 'panic_print'. Possible values are:
============= ===================================================
tasks print all tasks info
mem print system memory info
timer print timers info
lock print locks info if CONFIG_LOCKDEP is on
ftrace print ftrace buffer
all_bt print all CPUs backtrace (if available in the arch)
blocked_tasks print only tasks in uninterruptible (blocked) state
============= ===================================================
panic_on_rcu_stall
==================
When set to 1, calls panic() after RCU stall detection messages. This
is useful to define the root cause of RCU stalls using a vmcore.
= ============================================================
0 Do not panic() when RCU stall takes place, default behavior.
1 panic() after printing RCU stall messages.
= ============================================================
max_rcu_stall_to_panic
======================
When ``panic_on_rcu_stall`` is set to 1, this value determines the
number of times that RCU can stall before panic() is called.
When ``panic_on_rcu_stall`` is set to 0, this value is has no effect.
perf_cpu_time_max_percent
=========================
Hints to the kernel how much CPU time it should be allowed to
use to handle perf sampling events. If the perf subsystem
is informed that its samples are exceeding this limit, it
will drop its sampling frequency to attempt to reduce its CPU
usage.
Some perf sampling happens in NMIs. If these samples
unexpectedly take too long to execute, the NMIs can become
stacked up next to each other so much that nothing else is
allowed to execute.
===== ========================================================
0 Disable the mechanism. Do not monitor or correct perf's
sampling rate no matter how CPU time it takes.
1-100 Attempt to throttle perf's sample rate to this
percentage of CPU. Note: the kernel calculates an
"expected" length of each sample event. 100 here means
100% of that expected length. Even if this is set to
100, you may still see sample throttling if this
length is exceeded. Set to 0 if you truly do not care
how much CPU is consumed.
===== ========================================================
perf_event_paranoid
===================
Controls use of the performance events system by unprivileged
users (without CAP_PERFMON). The default value is 2.
For backward compatibility reasons access to system performance
monitoring and observability remains open for CAP_SYS_ADMIN
privileged processes but CAP_SYS_ADMIN usage for secure system
performance monitoring and observability operations is discouraged
with respect to CAP_PERFMON use cases.
=== ==================================================================
-1 Allow use of (almost) all events by all users.
Ignore mlock limit after perf_event_mlock_kb without
``CAP_IPC_LOCK``.
>=0 Disallow ftrace function tracepoint by users without
``CAP_PERFMON``.
Disallow raw tracepoint access by users without ``CAP_PERFMON``.
>=1 Disallow CPU event access by users without ``CAP_PERFMON``.
>=2 Disallow kernel profiling by users without ``CAP_PERFMON``.
=== ==================================================================
perf_event_max_stack
====================
Controls maximum number of stack frames to copy for (``attr.sample_type &
PERF_SAMPLE_CALLCHAIN``) configured events, for instance, when using
'``perf record -g``' or '``perf trace --call-graph fp``'.
This can only be done when no events are in use that have callchains
enabled, otherwise writing to this file will return ``-EBUSY``.
The default value is 127.
perf_event_mlock_kb
===================
Control size of per-cpu ring buffer not counted against mlock limit.
The default value is 512 + 1 page
perf_event_max_contexts_per_stack
=================================
Controls maximum number of stack frame context entries for
(``attr.sample_type & PERF_SAMPLE_CALLCHAIN``) configured events, for
instance, when using '``perf record -g``' or '``perf trace --call-graph fp``'.
This can only be done when no events are in use that have callchains
enabled, otherwise writing to this file will return ``-EBUSY``.
The default value is 8.
perf_user_access (arm64 and riscv only)
=======================================
Controls user space access for reading perf event counters.
* for arm64
The default value is 0 (access disabled).
When set to 1, user space can read performance monitor counter registers
directly.
See Documentation/arch/arm64/perf.rst for more information.
* for riscv
When set to 0, user space access is disabled.
The default value is 1, user space can read performance monitor counter
registers through perf, any direct access without perf intervention will trigger
an illegal instruction.
When set to 2, which enables legacy mode (user space has direct access to cycle
and insret CSRs only). Note that this legacy value is deprecated and will be
removed once all user space applications are fixed.
Note that the time CSR is always directly accessible to all modes.
pid_max
=======
PID allocation wrap value. When the kernel's next PID value
reaches this value, it wraps back to a minimum PID value.
PIDs of value ``pid_max`` or larger are not allocated.
ns_last_pid
===========
The last pid allocated in the current (the one task using this sysctl
lives in) pid namespace. When selecting a pid for a next task on fork
kernel tries to allocate a number starting from this one.
powersave-nap (PPC only)
========================
If set, Linux-PPC will use the 'nap' mode of powersaving,
otherwise the 'doze' mode will be used.
==============================================================
printk
======
The four values in printk denote: ``console_loglevel``,
``default_message_loglevel``, ``minimum_console_loglevel`` and
``default_console_loglevel`` respectively.
These values influence printk() behavior when printing or
logging error messages. See '``man 2 syslog``' for more info on
the different loglevels.
======================== =====================================
console_loglevel messages with a higher priority than
this will be printed to the console
default_message_loglevel messages without an explicit priority
will be printed with this priority
minimum_console_loglevel minimum (highest) value to which
console_loglevel can be set
default_console_loglevel default value for console_loglevel
======================== =====================================
printk_delay
============
Delay each printk message in ``printk_delay`` milliseconds
Value from 0 - 10000 is allowed.
printk_ratelimit
================
Some warning messages are rate limited. ``printk_ratelimit`` specifies
the minimum length of time between these messages (in seconds).
The default value is 5 seconds.
A value of 0 will disable rate limiting.
printk_ratelimit_burst
======================
While long term we enforce one message per `printk_ratelimit`_
seconds, we do allow a burst of messages to pass through.
``printk_ratelimit_burst`` specifies the number of messages we can
send before ratelimiting kicks in. After `printk_ratelimit`_ seconds
have elapsed, another burst of messages may be sent.
The default value is 10 messages.
printk_devkmsg
==============
Control the logging to ``/dev/kmsg`` from userspace:
========= =============================================
ratelimit default, ratelimited
on unlimited logging to /dev/kmsg from userspace
off logging to /dev/kmsg disabled
========= =============================================
The kernel command line parameter ``printk.devkmsg=`` overrides this and is
a one-time setting until next reboot: once set, it cannot be changed by
this sysctl interface anymore.
==============================================================
pty
===
See Documentation/filesystems/devpts.rst.
random
======
This is a directory, with the following entries:
* ``boot_id``: a UUID generated the first time this is retrieved, and
unvarying after that;
* ``uuid``: a UUID generated every time this is retrieved (this can
thus be used to generate UUIDs at will);
* ``entropy_avail``: the pool's entropy count, in bits;
* ``poolsize``: the entropy pool size, in bits;
* ``urandom_min_reseed_secs``: obsolete (used to determine the minimum
number of seconds between urandom pool reseeding). This file is
writable for compatibility purposes, but writing to it has no effect
on any RNG behavior;
* ``write_wakeup_threshold``: when the entropy count drops below this
(as a number of bits), processes waiting to write to ``/dev/random``
are woken up. This file is writable for compatibility purposes, but
writing to it has no effect on any RNG behavior.
randomize_va_space
==================
This option can be used to select the type of process address
space randomization that is used in the system, for architectures
that support this feature.
== ===========================================================================
0 Turn the process address space randomization off. This is the
default for architectures that do not support this feature anyways,
and kernels that are booted with the "norandmaps" parameter.
1 Make the addresses of mmap base, stack and VDSO page randomized.
This, among other things, implies that shared libraries will be
loaded to random addresses. Also for PIE-linked binaries, the
location of code start is randomized. This is the default if the
``CONFIG_COMPAT_BRK`` option is enabled.
2 Additionally enable heap randomization. This is the default if
``CONFIG_COMPAT_BRK`` is disabled.
There are a few legacy applications out there (such as some ancient
versions of libc.so.5 from 1996) that assume that brk area starts
just after the end of the code+bss. These applications break when
start of the brk area is randomized. There are however no known
non-legacy applications that would be broken this way, so for most
systems it is safe to choose full randomization.
Systems with ancient and/or broken binaries should be configured
with ``CONFIG_COMPAT_BRK`` enabled, which excludes the heap from process
address space randomization.
== ===========================================================================
real-root-dev
=============
See Documentation/admin-guide/initrd.rst.
reboot-cmd (SPARC only)
=======================
??? This seems to be a way to give an argument to the Sparc
ROM/Flash boot loader. Maybe to tell it what to do after
rebooting. ???
sched_energy_aware
==================
Enables/disables Energy Aware Scheduling (EAS). EAS starts
automatically on platforms where it can run (that is,
platforms with asymmetric CPU topologies and having an Energy
Model available). If your platform happens to meet the
requirements for EAS but you do not want to use it, change
this value to 0. On Non-EAS platforms, write operation fails and
read doesn't return anything.
task_delayacct
===============
Enables/disables task delay accounting (see
Documentation/accounting/delay-accounting.rst. Enabling this feature incurs
a small amount of overhead in the scheduler but is useful for debugging
and performance tuning. It is required by some tools such as iotop.
sched_schedstats
================
Enables/disables scheduler statistics. Enabling this feature
incurs a small amount of overhead in the scheduler but is
useful for debugging and performance tuning.
sched_util_clamp_min
====================
Max allowed *minimum* utilization.
Default value is 1024, which is the maximum possible value.
It means that any requested uclamp.min value cannot be greater than
sched_util_clamp_min, i.e., it is restricted to the range
[0:sched_util_clamp_min].
sched_util_clamp_max
====================
Max allowed *maximum* utilization.
Default value is 1024, which is the maximum possible value.
It means that any requested uclamp.max value cannot be greater than
sched_util_clamp_max, i.e., it is restricted to the range
[0:sched_util_clamp_max].
sched_util_clamp_min_rt_default
===============================
By default Linux is tuned for performance. Which means that RT tasks always run
at the highest frequency and most capable (highest capacity) CPU (in
heterogeneous systems).
Uclamp achieves this by setting the requested uclamp.min of all RT tasks to
1024 by default, which effectively boosts the tasks to run at the highest
frequency and biases them to run on the biggest CPU.
This knob allows admins to change the default behavior when uclamp is being
used. In battery powered devices particularly, running at the maximum
capacity and frequency will increase energy consumption and shorten the battery
life.
This knob is only effective for RT tasks which the user hasn't modified their
requested uclamp.min value via sched_setattr() syscall.
This knob will not escape the range constraint imposed by sched_util_clamp_min
defined above.
For example if
sched_util_clamp_min_rt_default = 800
sched_util_clamp_min = 600
Then the boost will be clamped to 600 because 800 is outside of the permissible
range of [0:600]. This could happen for instance if a powersave mode will
restrict all boosts temporarily by modifying sched_util_clamp_min. As soon as
this restriction is lifted, the requested sched_util_clamp_min_rt_default
will take effect.
seccomp
=======
See Documentation/userspace-api/seccomp_filter.rst.
sg-big-buff
===========
This file shows the size of the generic SCSI (sg) buffer.
You can't tune it just yet, but you could change it on
compile time by editing ``include/scsi/sg.h`` and changing
the value of ``SG_BIG_BUFF``.
There shouldn't be any reason to change this value. If
you can come up with one, you probably know what you
are doing anyway :)
shmall
======
This parameter sets the total amount of shared memory pages that can be used
inside ipc namespace. The shared memory pages counting occurs for each ipc
namespace separately and is not inherited. Hence, ``shmall`` should always be at
least ``ceil(shmmax/PAGE_SIZE)``.
If you are not sure what the default ``PAGE_SIZE`` is on your Linux
system, you can run the following command::
# getconf PAGE_SIZE
To reduce or disable the ability to allocate shared memory, you must create a
new ipc namespace, set this parameter to the required value and prohibit the
creation of a new ipc namespace in the current user namespace or cgroups can
be used.
shmmax
======
This value can be used to query and set the run time limit
on the maximum shared memory segment size that can be created.
Shared memory segments up to 1Gb are now supported in the
kernel. This value defaults to ``SHMMAX``.
shmmni
======
This value determines the maximum number of shared memory segments.
4096 by default (``SHMMNI``).
shm_rmid_forced
===============
Linux lets you set resource limits, including how much memory one
process can consume, via ``setrlimit(2)``. Unfortunately, shared memory
segments are allowed to exist without association with any process, and
thus might not be counted against any resource limits. If enabled,
shared memory segments are automatically destroyed when their attach
count becomes zero after a detach or a process termination. It will
also destroy segments that were created, but never attached to, on exit
from the process. The only use left for ``IPC_RMID`` is to immediately
destroy an unattached segment. Of course, this breaks the way things are
defined, so some applications might stop working. Note that this
feature will do you no good unless you also configure your resource
limits (in particular, ``RLIMIT_AS`` and ``RLIMIT_NPROC``). Most systems don't
need this.
Note that if you change this from 0 to 1, already created segments
without users and with a dead originative process will be destroyed.
sysctl_writes_strict
====================
Control how file position affects the behavior of updating sysctl values
via the ``/proc/sys`` interface:
== ======================================================================
-1 Legacy per-write sysctl value handling, with no printk warnings.
Each write syscall must fully contain the sysctl value to be
written, and multiple writes on the same sysctl file descriptor
will rewrite the sysctl value, regardless of file position.
0 Same behavior as above, but warn about processes that perform writes
to a sysctl file descriptor when the file position is not 0.
1 (default) Respect file position when writing sysctl strings. Multiple
writes will append to the sysctl value buffer. Anything past the max
length of the sysctl value buffer will be ignored. Writes to numeric
sysctl entries must always be at file position 0 and the value must
be fully contained in the buffer sent in the write syscall.
== ======================================================================
softlockup_all_cpu_backtrace
============================
This value controls the soft lockup detector thread's behavior
when a soft lockup condition is detected as to whether or not
to gather further debug information. If enabled, each cpu will
be issued an NMI and instructed to capture stack trace.
This feature is only applicable for architectures which support
NMI.
= ============================================
0 Do nothing. This is the default behavior.
1 On detection capture more debug information.
= ============================================
softlockup_panic
=================
This parameter can be used to control whether the kernel panics
when a soft lockup is detected.
= ============================================
0 Don't panic on soft lockup.
1 Panic on soft lockup.
= ============================================
This can also be set using the softlockup_panic kernel parameter.
soft_watchdog
=============
This parameter can be used to control the soft lockup detector.
= =================================
0 Disable the soft lockup detector.
1 Enable the soft lockup detector.
= =================================
The soft lockup detector monitors CPUs for threads that are hogging the CPUs
without rescheduling voluntarily, and thus prevent the 'migration/N' threads
from running, causing the watchdog work fail to execute. The mechanism depends
on the CPUs ability to respond to timer interrupts which are needed for the
watchdog work to be queued by the watchdog timer function, otherwise the NMI
watchdog — if enabled — can detect a hard lockup condition.
split_lock_mitigate (x86 only)
==============================
On x86, each "split lock" imposes a system-wide performance penalty. On larger
systems, large numbers of split locks from unprivileged users can result in
denials of service to well-behaved and potentially more important users.
The kernel mitigates these bad users by detecting split locks and imposing
penalties: forcing them to wait and only allowing one core to execute split
locks at a time.
These mitigations can make those bad applications unbearably slow. Setting
split_lock_mitigate=0 may restore some application performance, but will also
increase system exposure to denial of service attacks from split lock users.
= ===================================================================
0 Disable the mitigation mode - just warns the split lock on kernel log
and exposes the system to denials of service from the split lockers.
1 Enable the mitigation mode (this is the default) - penalizes the split
lockers with intentional performance degradation.
= ===================================================================
stack_erasing
=============
This parameter can be used to control kernel stack erasing at the end
of syscalls for kernels built with ``CONFIG_KSTACK_ERASE``.
That erasing reduces the information which kernel stack leak bugs
can reveal and blocks some uninitialized stack variable attacks.
The tradeoff is the performance impact: on a single CPU system kernel
compilation sees a 1% slowdown, other systems and workloads may vary.
= ====================================================================
0 Kernel stack erasing is disabled, KSTACK_ERASE_METRICS are not updated.
1 Kernel stack erasing is enabled (default), it is performed before
returning to the userspace at the end of syscalls.
= ====================================================================
stop-a (SPARC only)
===================
Controls Stop-A:
= ====================================
0 Stop-A has no effect.
1 Stop-A breaks to the PROM (default).
= ====================================
Stop-A is always enabled on a panic, so that the user can return to
the boot PROM.
sysrq
=====
See Documentation/admin-guide/sysrq.rst.
tainted
=======
Non-zero if the kernel has been tainted. Numeric values, which can be
ORed together. The letters are seen in "Tainted" line of Oops reports.
====== ===== ==============================================================
1 `(P)` proprietary module was loaded
2 `(F)` module was force loaded
4 `(S)` kernel running on an out of specification system
8 `(R)` module was force unloaded
16 `(M)` processor reported a Machine Check Exception (MCE)
32 `(B)` bad page referenced or some unexpected page flags
64 `(U)` taint requested by userspace application
128 `(D)` kernel died recently, i.e. there was an OOPS or BUG
256 `(A)` an ACPI table was overridden by user
512 `(W)` kernel issued warning
1024 `(C)` staging driver was loaded
2048 `(I)` workaround for bug in platform firmware applied
4096 `(O)` externally-built ("out-of-tree") module was loaded
8192 `(E)` unsigned module was loaded
16384 `(L)` soft lockup occurred
32768 `(K)` kernel has been live patched
65536 `(X)` Auxiliary taint, defined and used by for distros
131072 `(T)` The kernel was built with the struct randomization plugin
====== ===== ==============================================================
See Documentation/admin-guide/tainted-kernels.rst for more information.
Note:
writes to this sysctl interface will fail with ``EINVAL`` if the kernel is
booted with the command line option ``panic_on_taint=<bitmask>,nousertaint``
and any of the ORed together values being written to ``tainted`` match with
the bitmask declared on panic_on_taint.
See Documentation/admin-guide/kernel-parameters.rst for more details on
that particular kernel command line option and its optional
``nousertaint`` switch.
threads-max
===========
This value controls the maximum number of threads that can be created
using ``fork()``.
During initialization the kernel sets this value such that even if the
maximum number of threads is created, the thread structures occupy only
a part (1/8th) of the available RAM pages.
The minimum value that can be written to ``threads-max`` is 1.
The maximum value that can be written to ``threads-max`` is given by the
constant ``FUTEX_TID_MASK`` (0x3fffffff).
If a value outside of this range is written to ``threads-max`` an
``EINVAL`` error occurs.
timer_migration
===============
When set to a non-zero value, attempt to migrate timers away from idle cpus to
allow them to remain in low power states longer.
Default is set (1).
traceoff_on_warning
===================
When set, disables tracing (see Documentation/trace/ftrace.rst) when a
``WARN()`` is hit.
tracepoint_printk
=================
When tracepoints are sent to printk() (enabled by the ``tp_printk``
boot parameter), this entry provides runtime control::
echo 0 > /proc/sys/kernel/tracepoint_printk
will stop tracepoints from being sent to printk(), and::
echo 1 > /proc/sys/kernel/tracepoint_printk
will send them to printk() again.
This only works if the kernel was booted with ``tp_printk`` enabled.
See Documentation/admin-guide/kernel-parameters.rst and
Documentation/trace/boottime-trace.rst.
unaligned-trap
==============
On architectures where unaligned accesses cause traps, and where this
feature is supported (``CONFIG_SYSCTL_ARCH_UNALIGN_ALLOW``; currently,
``arc``, ``parisc`` and ``loongarch``), controls whether unaligned traps
are caught and emulated (instead of failing).
= ========================================================
0 Do not emulate unaligned accesses.
1 Emulate unaligned accesses. This is the default setting.
= ========================================================
See also `ignore-unaligned-usertrap`_.
unknown_nmi_panic
=================
The value in this file affects behavior of handling NMI. When the
value is non-zero, unknown NMI is trapped and then panic occurs. At
that time, kernel debugging information is displayed on console.
NMI switch that most IA32 servers have fires unknown NMI up, for
example. If a system hangs up, try pressing the NMI switch.
unprivileged_bpf_disabled
=========================
Writing 1 to this entry will disable unprivileged calls to ``bpf()``;
once disabled, calling ``bpf()`` without ``CAP_SYS_ADMIN`` or ``CAP_BPF``
will return ``-EPERM``. Once set to 1, this can't be cleared from the
running kernel anymore.
Writing 2 to this entry will also disable unprivileged calls to ``bpf()``,
however, an admin can still change this setting later on, if needed, by
writing 0 or 1 to this entry.
If ``BPF_UNPRIV_DEFAULT_OFF`` is enabled in the kernel config, then this
entry will default to 2 instead of 0.
= =============================================================
0 Unprivileged calls to ``bpf()`` are enabled
1 Unprivileged calls to ``bpf()`` are disabled without recovery
2 Unprivileged calls to ``bpf()`` are disabled
= =============================================================
warn_limit
==========
Number of kernel warnings after which the kernel should panic when
``panic_on_warn`` is not set. Setting this to 0 disables checking
the warning count. Setting this to 1 has the same effect as setting
``panic_on_warn=1``. The default value is 0.
watchdog
========
This parameter can be used to disable or enable the soft lockup detector
*and* the NMI watchdog (i.e. the hard lockup detector) at the same time.
= ==============================
0 Disable both lockup detectors.
1 Enable both lockup detectors.
= ==============================
The soft lockup detector and the NMI watchdog can also be disabled or
enabled individually, using the ``soft_watchdog`` and ``nmi_watchdog``
parameters.
If the ``watchdog`` parameter is read, for example by executing::
cat /proc/sys/kernel/watchdog
the output of this command (0 or 1) shows the logical OR of
``soft_watchdog`` and ``nmi_watchdog``.
watchdog_cpumask
================
This value can be used to control on which cpus the watchdog may run.
The default cpumask is all possible cores, but if ``NO_HZ_FULL`` is
enabled in the kernel config, and cores are specified with the
``nohz_full=`` boot argument, those cores are excluded by default.
Offline cores can be included in this mask, and if the core is later
brought online, the watchdog will be started based on the mask value.
Typically this value would only be touched in the ``nohz_full`` case
to re-enable cores that by default were not running the watchdog,
if a kernel lockup was suspected on those cores.
The argument value is the standard cpulist format for cpumasks,
so for example to enable the watchdog on cores 0, 2, 3, and 4 you
might say::
echo 0,2-4 > /proc/sys/kernel/watchdog_cpumask
watchdog_thresh
===============
This value can be used to control the frequency of hrtimer and NMI
events and the soft and hard lockup thresholds. The default threshold
is 10 seconds.
The softlockup threshold is (``2 * watchdog_thresh``). Setting this
tunable to zero will disable lockup detection altogether.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
문서 범위와 주의 사항
1-31이 문서는 Rik van Riel과 Shen Feng이 작성한 `/proc/sys/kernel/` sysctl 설명서입니다. 일반 정보와 법적 안내는 `Documentation/admin-guide/sysctl/index.rst`를 참조합니다.
이 디렉터리의 파일은 Linux 커널의 여러 일반 동작을 감시하고 조정합니다. 일부 값은 시스템을 망가뜨릴 수 있으므로 실제 변경 전에 문서와 소스를 모두 읽어야 합니다. 나타나는 항목은 커널 구성에 따라 달라집니다.
acct
32-54highwater lowwater frequency
BSD 방식 프로세스 회계가 활성화됐을 때 세 값이 동작을 제어합니다. 로그 파일 시스템의 여유 공간이 `lowwater`% 아래로 내려가면 회계를 중지하고, `highwater`% 위로 올라가면 재개합니다. `frequency`는 여유 공간을 다시 검사하는 간격이며 단위는 초입니다.
4 2 30
기본값은 여유 공간이 2% 미만이면 중지하고 4% 이상이면 재개하며, 여유 공간 정보를 30초 동안 유효한 것으로 간주한다는 뜻입니다.
acpi_video_flags
55-67`Documentation/power/video.rst`를 참조하십시오. `acpi_sleep` 커널 매개변수와 비슷하게 다음 값을 조합해 비디오 resume 모드를 설정합니다.
| 값 | 플래그 |
|---|---|
| `1` | `s3_bios` |
| `2` | `s3_mode` |
| `4` | `s3_beep` |
arch
68-73머신 하드웨어 이름입니다. `uname -m`과 같은 결과를 내며 예로 `x86_64`, `aarch64`가 있습니다.
auto_msgmni
74-85현재는 아무 효과가 없고 앞으로 제거될 수 있으며 읽으면 항상 `0`을 반환합니다. Linux 3.17까지는 메모리 추가·제거 또는 IPC namespace 생성·제거 때 `msgmni`를 자동 재계산할지 제어했습니다. `1`은 활성화, `0`은 비활성화였고 기본값은 `1`이었습니다.
bootloader_type (x86 only)
86-100bootloader가 알린 형식 번호를 4비트 왼쪽으로 이동한 뒤 bootloader 버전의 하위 4비트와 OR한 값입니다. 예전 커널 헤더의 `type_of_loader` 필드와 일치하던 인코딩을 하위 호환성 때문에 유지합니다. 전체 형식 번호가 `0x15`, 전체 버전이 `0x234`라면 값은 `340 = 0x154`입니다.
자세한 내용은 `Documentation/arch/x86/boot.rst`의 `type_of_loader`와 `ext_loader_type` 필드를 참조하십시오.
bootloader_version (x86 only)
101-110완전한 bootloader 버전 번호입니다. 앞 예에서는 `564 = 0x234`입니다. 자세한 내용은 `Documentation/arch/x86/boot.rst`의 `type_of_loader`와 `ext_loader_ver` 필드를 참조하십시오.
bpf_stats_enabled
111-124BPF 프로그램의 총 실행 시간과 실행 횟수 같은 통계를 커널이 수집할지 제어합니다. 수집을 켜면 프로그램을 실행할 때마다 성능이 약간 떨어지며, 통계는 `bpftool`로 볼 수 있습니다.
| 값 | 동작 |
|---|---|
| `0` | 통계를 수집하지 않음(기본값) |
| `1` | 통계를 수집함 |
cad_pid
125-134재부팅 때, 특히 Ctrl-Alt-Delete로 재부팅할 때 signal을 받을 PID입니다. 실행 중인 프로세스와 일치하지 않는 값을 쓰면 `-ESRCH`가 발생합니다. `ctrl-alt-del` 절도 참조하십시오.
cap_last_cap
135-143실행 중인 커널에서 유효한 가장 높은 capability이며 커널의 `CAP_LAST_CAP`을 내보냅니다.
core_pattern
144-188`core_pattern`은 core dump 파일 이름 패턴을 지정합니다. 최대 길이는 127자이고 기본값은 `core`입니다. `%`로 시작하는 형식 지정자는 실제 값으로 치환됩니다.
하위 호환성을 위해 패턴에 `%p`가 없고 `core_uses_pid`가 설정돼 있으면 파일 이름 뒤에 `.PID`를 붙입니다.
| 지정자 | 치환 값 |
|---|---|
| `%<NUL>` | `%`를 버림 |
| `%%` | `%` 하나 출력 |
| `%p` | PID |
| `%P` | init PID namespace의 전역 PID |
| `%i` | TID |
| `%I` | init PID namespace의 전역 TID |
| `%u` | 초기 user namespace의 UID |
| `%g` | 초기 user namespace의 GID |
| `%d` | `PR_SET_DUMPABLE` 및 `/proc/sys/fs/suid_dumpable`과 일치하는 dump 모드 |
| `%s` | signal 번호 |
| `%t` | dump의 UNIX 시간 |
| `%h` | hostname |
| `%e` | 실행 파일 이름. 줄어들거나 `prctl` 등으로 바뀔 수 있음 |
| `%f` | 실행 파일 이름 |
| `%E` | 실행 파일 경로 |
| `%c` | `RLIMIT_CORE`가 정한 최대 core 파일 크기 |
| `%C` | task가 실행된 CPU |
| `%F` | pidfd 번호 |
| `%<OTHER>` | `%`와 뒤 문자를 모두 버림 |
패턴의 첫 문자가 `|`이면 나머지를 실행할 명령으로 취급하고, core dump를 파일 대신 그 프로그램의 표준 입력으로 씁니다.
core_pipe_limit
189-215이 sysctl은 `core_pattern`이 `|`로 시작해 core를 사용자 공간 helper로 pipe하는 경우에만 적용됩니다. 수집기가 충돌 프로세스의 `/proc/pid` 정보를 읽을 수 있도록 커널은 수집 프로세스가 종료될 때까지 기다려야 합니다.
잘못된 수집기가 종료하지 않으면 충돌 프로세스의 회수를 막을 수 있으므로, 이 값은 동시에 사용자 공간으로 pipe할 수 있는 충돌 프로세스 수를 제한합니다. 초과한 프로세스는 커널 로그에 기록하고 core를 건너뜁니다.
`0`은 병렬 수집 수가 무제한이지만 기다리지 않는 특별한 값입니다. 따라서 수집기가 `/proc/<crashing pid>/`에 접근할 수 있다는 보장이 없습니다. 기본값은 `0`입니다.
core_sort_vma
216-226기본 coredump는 VMA를 주소 순서로 씁니다. `1`이면 작은 VMA부터 큰 VMA 순으로 씁니다. 적어도 elfutils가 이 형식을 처리하지 못하는 것으로 알려졌지만, 매우 크고 잘린 coredump에서 유용한 디버깅 정보가 작은 VMA에 있을 때 도움이 될 수 있습니다.
core_uses_pid
227-236기본 coredump 이름은 `core`입니다. `core_uses_pid=1`이면 `core.PID`가 됩니다. `core_pattern`에 `%p`가 없고 이 값이 설정돼 있으면 `.PID`를 붙입니다.
ctrl-alt-del
237-252`0`이면 Ctrl-Alt-Delete를 가로채 `init(1)`에 보내 정상 재시작을 처리합니다. `0`보다 크면 dirty buffer를 동기화하지 않고 즉시 재부팅합니다.
dosemu 같은 프로그램이 키보드를 raw 모드로 사용하면 커널 tty 계층에 도달하기 전에 그 프로그램이 Ctrl-Alt-Delete를 가로채므로 처리 방식도 해당 프로그램이 결정합니다.
dmesg_restrict
253-266권한 없는 사용자가 `dmesg(8)`로 커널 로그 버퍼를 볼 수 있는지 제어합니다. `0`은 제한 없음, `1`은 `CAP_SYSLOG`가 있어야 접근 가능하다는 뜻입니다. `CONFIG_SECURITY_DMESG_RESTRICT`가 기본값을 정합니다.
domainname & hostname
267-289NIS/YP domainname과 hostname을 명령과 같은 방식으로 설정할 수 있습니다.
# echo "darkstar" > /proc/sys/kernel/hostname
# echo "mydomain" > /proc/sys/kernel/domainname
위 명령은 다음과 같은 효과를 냅니다.
# hostname "darkstar"
# domainname "mydomain"
`darkstar.frop.org`의 hostname은 `darkstar`, DNS domainname은 `frop.org`입니다. DNS domainname은 NIS 또는 YP domainname과 혼동하면 안 되며 일반적으로 서로 다릅니다. 자세한 설명은 `hostname(1)`을 참조하십시오.
firmware_config
290-302`Documentation/driver-api/firmware/fallback-mechanisms.rst`를 참조하십시오. 이 디렉터리는 firmware loader helper fallback을 제어합니다.
| 항목 | 값 `1`의 의미 |
|---|---|
| `force_sysfs_fallback` | fallback 사용을 강제함 |
| `ignore_sysfs_fallback` | 모든 fallback을 무시함 |
ftrace_dump_on_oops
303-335oops 또는 kernel panic 때 `ftrace_dump()`를 호출해 ftrace buffer를 console로 출력할지 정합니다. 충돌까지 이어진 trace를 serial console에 남길 때 유용합니다.
| 값 | 동작 |
|---|---|
| `0` | 비활성화(기본값) |
| `1` | 모든 CPU buffer dump |
| `2` 또는 `orig_cpu` | oops를 일으킨 CPU buffer dump |
| `<instance>` | 모든 CPU에서 지정 instance buffer dump |
| `<instance>=2` 또는 `<instance>=orig_cpu` | oops를 일으킨 CPU에서 지정 instance buffer dump |
쉼표로 여러 instance를 지정할 수 있습니다. 전역 buffer도 필요하면 전역 모드 `1`, `2`, `orig_cpu`를 맨 앞에 둡니다. 다음은 `foo`와 `bar` instance를 모든 CPU에서 dump합니다.
echo "foo,bar" > /proc/sys/kernel/ftrace_dump_on_oops
다음은 전역 buffer와 `foo`를 모든 CPU에서, `bar`를 oops 발생 CPU에서 dump합니다.
echo "1,foo,bar=2" > /proc/sys/kernel/ftrace_dump_on_oops
ftrace_enabled, stack_tracer_enabled
336-341`Documentation/trace/ftrace.rst`를 참조하십시오.
hardlockup_all_cpu_backtrace
342-355hard lockup 감지 때 추가 디버깅 정보를 수집할지 제어합니다. 활성화하면 아키텍처별 all-CPU stack dump를 시작합니다.
| 값 | 동작 |
|---|---|
| `0` | 아무 작업도 하지 않음(기본값) |
| `1` | 감지 시 추가 디버깅 정보 수집 |
hardlockup_panic
356-370hard lockup 감지 때 kernel panic을 일으킬지 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | panic하지 않음 |
| `1` | panic함 |
자세한 내용은 `Documentation/admin-guide/lockup-watchdogs.rst`를 참조하십시오. `nmi_watchdog` 커널 매개변수로도 설정할 수 있습니다.
hotplug
371-382hotplug 정책 agent의 경로입니다. 기본값은 `CONFIG_UEVENT_HELPER_PATH`이며, 그 기본값은 빈 문자열입니다. `CONFIG_UEVENT_HELPER`가 켜졌을 때만 존재합니다. 현대 시스템 대부분은 netlink 기반 uevent만 사용하므로 필요하지 않습니다.
hung_task_all_cpu_backtrace
383-396hung task 감지 때 모든 CPU에 NMI를 보내 backtrace를 dump할지 정합니다. `CONFIG_DETECT_HUNG_TASK`와 `CONFIG_SMP`가 켜졌을 때 나타납니다.
| 값 | 동작 |
|---|---|
| `0` | 모든 CPU backtrace를 표시하지 않음(기본값) |
| `1` | 모든 CPU를 non-maskable interrupt하고 backtrace를 dump함 |
hung_task_panic
397-408hung task 감지 때의 커널 동작을 제어하며 `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.
| 값 | 동작 |
|---|---|
| `0` | 계속 실행(기본값) |
| `1` | 즉시 panic |
hung_task_check_count
409-415검사할 task 수의 상한입니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.
hung_task_detect_count
416-424부팅 이후 hung 상태로 감지된 task의 총수입니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다.
hung_task_timeout_secs
425-436D 상태 task가 이 값보다 오래 schedule되지 않으면 경고합니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다. `0`은 무한 timeout이라 검사하지 않습니다. 설정 범위는 `0`부터 `LONG_MAX/HZ`까지입니다.
hung_task_check_interval_secs
437-450hung task 검사가 활성화됐을 때 `hung_task_check_interval_secs`초마다 검사합니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타납니다. 기본값 `0`은 검사 간격으로 `hung_task_timeout_secs`를 사용한다는 뜻이며 범위는 `0`부터 `LONG_MAX/HZ`까지입니다.
hung_task_warnings
451-461보고할 최대 경고 수입니다. 검사 구간에 hung task를 찾을 때마다 1씩 줄고, `0`이면 더는 경고하지 않습니다. `CONFIG_DETECT_HUNG_TASK`가 켜졌을 때 나타나며 `-1`은 무제한 경고입니다.
hyperv_record_panic_msg
462-472panic kmsg 데이터를 Hyper-V에 보고할지 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | 보고하지 않음 |
| `1` | 보고함(기본값) |
ignore-unaligned-usertrap
473-488unaligned 접근이 trap을 일으키고 `CONFIG_SYSCTL_ARCH_UNALIGN_NO_WARN`을 지원하는 아키텍처, 현재 `arc`, `parisc`, `loongarch`에서 모든 unaligned trap을 기록할지 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | 모든 unaligned 접근 기록 |
| `1` | 프로세스가 처음 trap될 때만 경고(기본값) |
`unaligned-trap` 절도 참조하십시오.
io_uring_disabled
489-507새 `io_uring` instance 생성을 제한해 커널 공격 표면을 줄입니다.
| 값 | 동작 |
|---|---|
| `0` | 모든 프로세스가 정상적으로 생성 가능(기본값) |
| `1` | `io_uring_group`에 속하지 않은 비특권 프로세스의 `io_uring_setup()`이 `-EPERM`으로 실패함. 기존 instance는 사용 가능 |
| `2` | 모든 프로세스의 `io_uring_setup()`이 `-EPERM`으로 실패함. 기존 instance는 사용 가능 |
io_uring_group
508-517`io_uring_disabled=1`일 때 새 instance를 만들려면 `CAP_SYS_ADMIN`이 있거나 `io_uring_group` 그룹에 속해야 합니다. 기본값 `-1`이면 `CAP_SYS_ADMIN`이 있는 프로세스만 만들 수 있습니다.
kexec_load_disabled
518-531`kexec_load`와 `kexec_file_load` syscall을 비활성화했는지 나타냅니다. 기본값 `0`은 `kexec_*load` 활성, `1`은 비활성입니다. 한 번 `1`이 되면 되돌릴 수 없고 kexec를 더는 사용할 수 없습니다.
syscall을 막기 전에 kexec image를 적재하면 나중에 사용할 image를 변경하지 못하게 보호할 수 있습니다. 일반적으로 `modules_disabled`와 함께 사용합니다.
kexec_load_limit_panic
532-544`kexec_load`와 `kexec_file_load`를 crash image와 함께 호출할 수 있는 횟수입니다. 현재보다 더 제한적인 값으로만 바꿀 수 있습니다.
| 값 | 동작 |
|---|---|
| `-1` | kexec 호출 무제한(기본값) |
| `N` | 남은 호출 횟수 |
kexec_load_limit_reboot
545-550`kexec_load_limit_panic`과 같은 기능을 normal image에 적용합니다.
kptr_restrict
551-579`/proc`와 다른 인터페이스에서 커널 주소를 노출할 때 적용할 제한을 정합니다.
| 값 | 동작 |
|---|---|
| `0` | 출력 전에 주소를 hash함(기본값, `%p`와 같음) |
| `1` | `%pK` 포인터를 출력할 때 사용자가 `CAP_SYSLOG`를 갖고 effective UID/GID가 real UID/GID와 같지 않으면 0으로 바꿈 |
| `2` | 권한과 무관하게 `%pK` 포인터를 0으로 바꿈 |
`%pK` 검사는 `open()`이 아니라 `read()` 때 수행하므로, open과 read 사이에 setuid binary 등으로 권한이 상승해도 비특권 사용자에게 포인터가 새지 않게 합니다. 이는 임시 해결책이며 장기적으로는 `open()` 때 검사해야 합니다.
포인터 노출이 우려되면 `%pK`를 사용하는 파일의 world-read 권한을 제거하고, `dmesg(8)`의 `%pK` 노출은 `dmesg_restrict`로 보호하는 방안을 고려하십시오.
modprobe
580-611커널 module을 자동 적재하는 usermode helper의 전체 경로입니다. 기본값은 `CONFIG_MODPROBE_PATH`이며 그 기본값은 `/sbin/modprobe`입니다. 사용자 공간이 `mount()`에 알 수 없는 파일 시스템 형식을 넘기는 경우처럼 커널이 module을 요청하면 이 binary를 실행하며, helper는 필요한 module을 커널에 삽입해야 합니다.
이 sysctl은 module 자동 적재에만 영향을 주고 명시적인 module 삽입 능력에는 영향을 주지 않습니다. 다음처럼 module 적재 요청을 디버깅할 수 있습니다.
echo '#! /bin/sh' > /tmp/modprobe
echo 'echo "$@" >> /tmp/modprobe.log' >> /tmp/modprobe
echo 'exec /sbin/modprobe "$@"' >> /tmp/modprobe
chmod a+x /tmp/modprobe
echo /tmp/modprobe > /proc/sys/kernel/modprobe
빈 문자열로 설정하면 module 자동 적재를 완전히 끕니다. 커널은 usermode helper를 실행하지 않고 `kernel_module_request` LSM hook도 호출하지 않습니다.
커널이 `CONFIG_STATIC_USERMODEHELPER=y`로 구성됐다면 정적 helper가 이 sysctl보다 우선합니다. 다만 빈 문자열로 자동 적재를 완전히 끄는 동작은 그대로 허용됩니다.
modules_disabled
612-623모듈식 커널에서 module 적재를 허용할지 나타내는 toggle입니다. 기본값 `0`은 허용이며 `1`로 바꿀 수 있습니다. 한 번 `1`이 되면 module을 적재하거나 제거할 수 없고 `0`으로 되돌릴 수도 없습니다. 일반적으로 `kexec_load_disabled`와 함께 사용합니다.
msgmax, msgmnb, and msgmni
624-639| 항목 | 의미와 기본값 |
|---|---|
| `msgmax` | IPC message 하나의 최대 크기(byte), 기본 `8192` (`MSGMAX`) |
| `msgmnb` | IPC queue 하나의 최대 크기(byte), 기본 `16384` (`MSGMNB`) |
| `msgmni` | IPC queue 최대 개수, 기본 `32000` (`MSGMNI`) |
세 매개변수는 IPC namespace별로 설정됩니다. POSIX message queue의 최대 byte 수는 `RLIMIT_MSGQUEUE`가 제한하며, 각 user namespace에서 계층적으로 적용됩니다.
msg_next_id, sem_next_id, and shm_next_id (System V IPC)
640-656다음에 할당할 message, semaphore, shared memory IPC 객체의 원하는 ID를 각각 지정합니다. 기본값 `-1`은 일반 할당 로직을 사용한다는 뜻이며 설정 범위는 `0`부터 `INT_MAX`까지입니다.
커널은 새 객체가 요청한 ID를 얻는다고 보장하지 않으므로 잘못된 ID를 처리하는 책임은 사용자 공간에 있습니다. 기본값이 아닌 toggle은 IPC 객체를 성공적으로 할당한 뒤 커널이 `-1`로 되돌립니다. 할당 syscall이 실패하면 값이 유지될지 `-1`로 초기화될지는 정의돼 있지 않습니다.
ngroups_max
657-664`setgroups`가 허용하는 supplementary group 수의 최댓값이며 커널의 `NGROUPS_MAX`를 내보냅니다.
nmi_watchdog
665-689x86 시스템에서 NMI watchdog, 즉 hard lockup detector를 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | hard lockup detector 비활성화 |
| `1` | hard lockup detector 활성화 |
detector는 각 CPU가 timer interrupt에 응답하는지 감시합니다. CPU가 바쁜 동안에도 주기적으로 NMI를 발생하도록 CPU performance counter register를 설정하므로 NMI watchdog이라고도 부릅니다.
KVM virtual machine의 guest로 실행되는 커널에서는 기본적으로 꺼집니다. guest 커널 command line에 다음 값을 추가해 기본 동작을 덮어쓸 수 있습니다.
nmi_watchdog=1
자세한 내용은 `Documentation/admin-guide/kernel-parameters.rst`를 참조하십시오.
nmi_wd_lpm_factor (PPC only)
690-701`nmi_watchdog=1`일 때 NMI watchdog timeout에 적용하는 계수입니다. LPM 중 timeout을 계산할 때 `watchdog_thresh`에 더할 백분율을 뜻하며 soft lockup timeout에는 영향을 주지 않습니다.
`0`은 변경 없음입니다. 기본값 `200`은 `watchdog_thresh=10`일 때 NMI watchdog을 30초로 설정합니다.
numa_balancing
702-732page fault 기반 자동 NUMA memory balancing을 활성화·비활성화하고 모드를 정합니다. 자주 접근하는 node로 memory를 자동 이동하며 다음 값을 OR해 설정할 수 있습니다.
| 값 | 모드 |
|---|---|
| `0` | `NUMA_BALANCING_DISABLED` |
| `1` | `NUMA_BALANCING_NORMAL` |
| `2` | `NUMA_BALANCING_MEMORY_TIERING` |
`NUMA_BALANCING_NORMAL`은 원격 memory 접근 비용을 줄이기 위해 NUMA node 사이에서 page 배치를 최적화합니다. 커널이 주기적으로 page mapping을 제거하고 뒤따르는 page fault를 trap해 어떤 task thread가 memory에 접근하는지 표본화한 뒤, 해당 데이터를 local memory node로 옮길지 결정합니다.
unmapping과 fault trapping에는 추가 비용이 들며 memory locality 개선이 이를 항상 상쇄한다는 보장은 없습니다. workload가 이미 NUMA node에 bind돼 있다면 이 기능을 꺼야 합니다.
`NUMA_BALANCING_MEMORY_TIERING`은 서로 다른 NUMA node로 표현되는 memory 유형 사이에서 hot page를 빠른 memory에 두도록 최적화하며 역시 unmapping과 page fault를 사용합니다.
numa_balancing_promote_rate_limit_MBps
733-743서로 다른 memory 유형 사이의 promotion·demotion 처리량이 지나치면 application latency가 나빠질 수 있습니다. 이 값은 node별 최대 promotion 처리량을 MB/s 단위로 제한합니다. 경험적으로 PMEM node 쓰기 bandwidth의 1/10보다 작게 설정합니다.
oops_all_cpu_backtrace
744-759oops가 발생하면 모든 CPU에 NMI를 보내 backtrace를 dump할지 정합니다. VM 보호 때문에 panic을 일으킬 수 없거나 kdump를 수집할 수 없을 때 마지막 수단으로 사용합니다. `CONFIG_SMP`가 켜졌을 때 나타납니다.
| 값 | 동작 |
|---|---|
| `0` | 모든 CPU backtrace를 표시하지 않음(기본값) |
| `1` | 모든 CPU를 non-maskable interrupt하고 backtrace를 dump함 |
oops_limit
760-768`panic_on_oops`가 설정되지 않았을 때 몇 번째 kernel oops 뒤에 panic할지 정합니다. `0`은 횟수 검사를 끄고 `1`은 `panic_on_oops=1`과 같습니다. 기본값은 `10000`입니다.
osrelease, ostype & version
769-788# cat osrelease
2.1.88
# cat ostype
Linux
# cat version
#5 Wed Feb 25 21:49:24 MET 1998
`osrelease`와 `ostype`의 의미는 명확합니다. `version`의 `#5`는 이 source base로 다섯 번째 빌드한 커널이라는 뜻이며 뒤 날짜는 빌드 시각입니다. 이 값은 커널을 다시 빌드해야만 바꿀 수 있습니다.
overflowgid & overflowuid
789-800항상 32비트 UID를 지원하지 않았던 `arm`, `i386`, `m68k`, `sh`, `sparc32` 아키텍처에서 실제 UID/GID가 `65535`를 넘을 때 구형 16비트 UID/GID syscall을 쓰는 application에 반환할 고정 UID/GID를 정합니다. 기본값은 `65534`입니다.
panic
801-814kernel panic 뒤의 동작을 정합니다.
| 값 | 동작 |
|---|---|
| `0` | 영원히 loop |
| 음수 | 즉시 재부팅 |
| 양수 | 해당 초가 지난 뒤 재부팅 |
software watchdog을 사용할 때 권장값은 `60`입니다.
panic_on_io_nmi
815-830CPU가 IO error로 발생한 NMI를 받았을 때의 동작을 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | 계속 실행 시도(기본값) |
| `1` | 즉시 panic. 데이터 손상 가능성이 있는 심각한 IO 상태를 계속 실행하지 않고 crash dump를 수집할 수 있음 |
일부 server는 dump button을 누를 때 이런 NMI를 발생시키므로 이 옵션을 crash dump 수집에 이용할 수 있습니다.
panic_on_oops
831-842oops 또는 BUG를 만났을 때의 커널 동작을 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | 계속 실행 시도 |
| `1` | 즉시 panic. `panic` sysctl도 0이 아니면 machine 재부팅 |
panic_on_stackoverflow
843-855user stack을 제외한 kernel, IRQ, exception stack overflow를 감지했을 때의 동작입니다. `CONFIG_DEBUG_STACKOVERFLOW`가 켜졌을 때 나타납니다.
| 값 | 동작 |
|---|---|
| `0` | 계속 실행 시도 |
| `1` | 즉시 panic |
panic_on_unrecovered_nmi
856-868memory 또는 알 수 없는 NMI에 대해 Linux 기본 동작은 계속 실행하는 것입니다. 과학 계산처럼 교정되지 않은 parity/ECC error의 전파보다 시스템을 격리하고 오류를 처리하는 편이 나은 환경에서는 panic이 적합할 수 있습니다.
일부 시스템은 전원 관리 같은 예상 밖의 이유로 NMI를 만들 수 있으므로 기본값은 off입니다. 이 sysctl은 같은 디렉터리의 다른 panic 제어와 같은 방식으로 동작합니다.
panic_on_warn
869-880`1`이면 `WARN()` 경로에서 `panic()`을 호출합니다. WARN 위치에서 kdump를 얻기 위해 커널을 다시 빌드하지 않아도 됩니다.
| 값 | 동작 |
|---|---|
| `0` | `WARN()`만 수행(기본값) |
| `1` | WARN 위치를 출력한 뒤 `panic()` 호출 |
panic_print
881-902panic 때 출력할 시스템 정보의 bitmask입니다. 다음 bit를 조합할 수 있습니다.
| 비트 | 출력 |
|---|---|
| bit 0 | 모든 task 정보 |
| bit 1 | 시스템 memory 정보 |
| bit 2 | timer 정보 |
| bit 3 | `CONFIG_LOCKDEP`가 켜졌으면 lock 정보 |
| bit 4 | ftrace buffer |
| bit 5 | panic 마지막에 모든 kernel message를 console에 재생 |
| bit 6 | 아키텍처가 지원하면 모든 CPU backtrace |
| bit 7 | uninterruptible(blocked) 상태 task만 출력 |
예를 들어 task와 memory 정보를 출력하려면 다음과 같이 설정합니다.
echo 3 > /proc/sys/kernel/panic_print
panic_sys_info
903-920panic 때 추가로 dump할 정보를 쉼표로 구분한 목록입니다. `tasks,mem,timers,...`처럼 읽기 쉬운 `panic_print` 대안입니다.
| 값 | 출력 |
|---|---|
| `tasks` | 모든 task 정보 |
| `mem` | 시스템 memory 정보 |
| `timer` | timer 정보 |
| `lock` | `CONFIG_LOCKDEP`가 켜졌으면 lock 정보 |
| `ftrace` | ftrace buffer |
| `all_bt` | 아키텍처가 지원하면 모든 CPU backtrace |
| `blocked_tasks` | uninterruptible(blocked) 상태 task만 출력 |
panic_on_rcu_stall
921-931`1`이면 RCU stall 감지 메시지를 출력한 뒤 `panic()`을 호출합니다. vmcore로 RCU stall의 root cause를 밝힐 때 유용합니다.
| 값 | 동작 |
|---|---|
| `0` | RCU stall 때 panic하지 않음(기본값) |
| `1` | RCU stall 메시지 뒤 panic |
max_rcu_stall_to_panic
932-939`panic_on_rcu_stall=1`일 때 몇 번의 RCU stall 뒤에 `panic()`을 호출할지 정합니다. `panic_on_rcu_stall=0`이면 효과가 없습니다.
perf_cpu_time_max_percent
940-967perf sampling event 처리에 허용할 CPU 시간의 비율을 커널에 알려 줍니다. sample이 한도를 넘으면 perf subsystem이 sampling frequency를 낮춰 CPU 사용량을 줄입니다.
일부 perf sample은 NMI에서 처리됩니다. 예상보다 오래 걸리면 NMI가 연속으로 쌓여 다른 작업이 실행되지 못할 수 있습니다.
| 값 | 동작 |
|---|---|
| `0` | 감시와 sample rate 보정을 끔 |
| `1-100` | perf sample rate가 해당 CPU 비율을 넘지 않도록 throttle을 시도 |
커널은 sample event의 예상 길이를 계산하므로 `100`은 그 예상 길이의 100%라는 뜻입니다. 실제 길이가 넘으면 `100`에서도 throttle될 수 있으며 CPU 소비량을 전혀 제한하지 않으려면 `0`을 사용합니다.
perf_event_paranoid
968-996`CAP_PERFMON`이 없는 비특권 사용자의 performance event system 사용을 제어합니다. 기본값은 `2`입니다. 하위 호환성을 위해 `CAP_SYS_ADMIN` 프로세스에도 시스템 performance monitoring 접근을 허용하지만, 보안 목적에는 `CAP_PERFMON` 사용이 권장됩니다.
| 값 | 제한 |
|---|---|
| `-1` | 거의 모든 event를 모든 사용자에게 허용하고 `CAP_IPC_LOCK` 없이 `perf_event_mlock_kb` 이후 mlock 한도를 무시 |
| `>=0` | `CAP_PERFMON` 없는 사용자의 ftrace function tracepoint와 raw tracepoint 접근 금지 |
| `>=1` | `CAP_PERFMON` 없는 사용자의 CPU event 접근 금지 |
| `>=2` | `CAP_PERFMON` 없는 사용자의 kernel profiling 금지 |
perf_event_max_stack
997-1009`attr.sample_type & PERF_SAMPLE_CALLCHAIN` event에서 복사할 최대 stack frame 수를 정합니다. 예로 `perf record -g`, `perf trace --call-graph fp`가 있습니다.
callchain이 활성화된 event가 사용 중이면 바꿀 수 없고 쓰기가 `-EBUSY`를 반환합니다. 기본값은 `127`입니다.
perf_event_mlock_kb
1010-1017mlock 한도에 계산하지 않는 CPU별 ring buffer 크기를 제어합니다. 기본값은 `512 + 1 page`입니다.
perf_event_max_contexts_per_stack
1018-1030`attr.sample_type & PERF_SAMPLE_CALLCHAIN` event의 최대 stack frame context entry 수를 정합니다. 예로 `perf record -g`, `perf trace --call-graph fp`가 있습니다.
callchain이 활성화된 event가 사용 중이면 바꿀 수 없고 쓰기가 `-EBUSY`를 반환합니다. 기본값은 `8`입니다.
perf_user_access (arm64 and riscv only)
1031-1056performance event counter를 읽는 사용자 공간 접근을 제어합니다.
arm64의 기본값은 `0`으로 접근 비활성화입니다. `1`이면 사용자 공간이 performance monitor counter register를 직접 읽을 수 있습니다. 자세한 내용은 `Documentation/arch/arm64/perf.rst`를 참조하십시오.
RISC-V에서 `0`은 사용자 공간 접근 비활성화입니다. 기본값 `1`은 perf를 통해 counter register를 읽게 하며 perf 개입 없는 직접 접근은 illegal instruction을 일으킵니다.
RISC-V 값 `2`는 legacy mode로 cycle과 `insret` CSR만 직접 접근할 수 있습니다. 이 값은 deprecated이며 사용자 공간 application이 모두 수정되면 제거됩니다. time CSR은 모든 모드에서 항상 직접 접근할 수 있습니다.
pid_max
1057-1064PID 할당이 되감기는 값입니다. 다음 PID가 이 값에 이르면 최소 PID로 돌아가며 `pid_max` 이상인 PID는 할당하지 않습니다.
ns_last_pid
1065-1072이 sysctl을 사용하는 task가 속한 현재 PID namespace에서 마지막으로 할당한 PID입니다. fork로 다음 task의 PID를 고를 때 커널은 이 번호부터 할당을 시도합니다.
powersave-nap (PPC only)
1073-1081설정하면 Linux-PPC가 절전 `nap` 모드를 사용하고, 설정하지 않으면 `doze` 모드를 사용합니다.
printk
1082-1103네 값은 순서대로 `console_loglevel`, `default_message_loglevel`, `minimum_console_loglevel`, `default_console_loglevel`입니다. 오류 message를 출력하거나 기록할 때 `printk()` 동작에 영향을 줍니다. loglevel은 `man 2 syslog`를 참조하십시오.
| 항목 | 의미 |
|---|---|
| `console_loglevel` | 이 값보다 우선순위가 높은 message를 console에 출력 |
| `default_message_loglevel` | 명시적 우선순위가 없는 message에 적용할 우선순위 |
| `minimum_console_loglevel` | `console_loglevel`에 설정할 수 있는 최소(가장 높은) 값 |
| `default_console_loglevel` | `console_loglevel`의 기본값 |
printk_delay
1104-1111각 printk message를 `printk_delay` millisecond만큼 지연합니다. 허용 범위는 `0-10000`입니다.
printk_ratelimit
1112-1121일부 warning message에는 rate limit이 적용됩니다. 이 값은 message 사이의 최소 시간(초)을 정하며 기본값은 5초입니다. `0`은 rate limiting을 끕니다.
printk_ratelimit_burst
1122-1133장기적으로 `printk_ratelimit`초마다 message 하나를 허용하지만 짧은 burst도 허용합니다. 이 값은 rate limiting이 시작되기 전에 보낼 수 있는 message 수이며, `printk_ratelimit`초 뒤 다시 같은 수의 burst를 보낼 수 있습니다. 기본값은 10개입니다.
printk_devkmsg
1134-1151사용자 공간에서 `/dev/kmsg`로 기록하는 동작을 제어합니다.
| 값 | 동작 |
|---|---|
| `ratelimit` | 기본값, rate limit 적용 |
| `on` | 사용자 공간의 `/dev/kmsg` 기록 무제한 |
| `off` | 사용자 공간의 `/dev/kmsg` 기록 비활성화 |
kernel command line의 `printk.devkmsg=`가 이 값을 덮어쓰며 다음 재부팅까지 한 번만 설정됩니다. 한 번 command line으로 설정하면 이 sysctl로 더는 바꿀 수 없습니다.
pty
1152-1157`Documentation/filesystems/devpts.rst`를 참조하십시오.
random
1158-1183다음 항목을 포함하는 디렉터리입니다.
| 항목 | 의미 |
|---|---|
| `boot_id` | 처음 읽을 때 생성되고 이후 바뀌지 않는 UUID |
| `uuid` | 읽을 때마다 생성되는 UUID. 필요할 때 UUID를 만드는 데 사용 가능 |
| `entropy_avail` | pool의 entropy 수(bit) |
| `poolsize` | entropy pool 크기(bit) |
| `urandom_min_reseed_secs` | obsolete. 예전에는 urandom pool reseed 사이의 최소 초를 정함. 호환성을 위해 쓸 수 있지만 RNG 동작에는 영향 없음 |
| `write_wakeup_threshold` | entropy가 이 bit 수 아래로 내려갈 때 `/dev/random`에 쓰려고 기다리는 프로세스를 깨움. 호환성을 위해 쓸 수 있지만 RNG 동작에는 영향 없음 |
randomize_va_space
1184-1217기능을 지원하는 아키텍처에서 사용할 프로세스 주소 공간 randomization 유형을 고릅니다.
| 값 | 동작 |
|---|---|
| `0` | 주소 공간 randomization을 끔. 기능 미지원 아키텍처와 `norandmaps`로 부팅한 커널의 기본값 |
| `1` | mmap base, stack, VDSO page 주소를 randomize함. shared library와 PIE binary의 code 시작 위치도 randomize함. `CONFIG_COMPAT_BRK`가 켜졌을 때 기본값 |
| `2` | heap randomization도 추가함. `CONFIG_COMPAT_BRK`가 꺼졌을 때 기본값 |
1996년의 일부 오래된 libc.so.5처럼 brk 영역이 code+bss 바로 뒤에서 시작한다고 가정하는 legacy application은 brk randomization으로 고장날 수 있습니다. 알려진 비legacy application 문제는 없으므로 대부분의 시스템에서는 full randomization이 안전합니다.
오래됐거나 잘못된 binary가 있는 시스템은 `CONFIG_COMPAT_BRK`를 켜서 heap을 프로세스 주소 공간 randomization에서 제외해야 합니다.
real-root-dev
1218-1223`Documentation/admin-guide/initrd.rst`를 참조하십시오.
reboot-cmd (SPARC only)
1224-1231문서 원문도 확정하지 못한 항목입니다. SPARC ROM/Flash boot loader에 인수를 전달해 재부팅 뒤 동작을 지시하는 방법으로 보입니다.
sched_energy_aware
1232-1242Energy Aware Scheduling(EAS)을 켜거나 끕니다. asymmetric CPU topology와 Energy Model이 있어 EAS를 실행할 수 있는 platform에서는 자동으로 시작합니다. 요건을 충족하지만 사용하지 않으려면 `0`으로 바꿉니다. EAS를 지원하지 않는 platform에서는 쓰기가 실패하고 읽기는 아무것도 반환하지 않습니다.
task_delayacct
1243-1250task delay accounting을 켜거나 끕니다. 자세한 내용은 `Documentation/accounting/delay-accounting.rst`를 참조하십시오. scheduler에 작은 overhead가 생기지만 디버깅과 성능 조정에 유용하고 `iotop` 같은 일부 도구가 필요로 합니다.
sched_schedstats
1251-1257scheduler 통계를 켜거나 끕니다. scheduler에 작은 overhead가 생기지만 디버깅과 성능 조정에 유용합니다.
sched_util_clamp_min
1258-1268허용할 minimum utilization의 최댓값입니다. 기본값 `1024`는 가능한 최대치입니다. 요청한 `uclamp.min`은 이 값보다 클 수 없으며 `[0:sched_util_clamp_min]` 범위로 제한됩니다.
sched_util_clamp_max
1269-1279허용할 maximum utilization의 최댓값입니다. 기본값 `1024`는 가능한 최대치입니다. 요청한 `uclamp.max`는 이 값보다 클 수 없으며 `[0:sched_util_clamp_max]` 범위로 제한됩니다.
sched_util_clamp_min_rt_default
1280-1312Linux는 기본적으로 성능에 맞춰 조정돼 있어 RT task가 항상 가장 높은 frequency와 heterogeneous system에서 가장 큰 capacity의 CPU에서 실행됩니다. uclamp는 모든 RT task의 요청 `uclamp.min`을 기본 `1024`로 설정해 이를 구현합니다.
이 knob는 uclamp를 사용할 때 관리자가 기본 동작을 바꿀 수 있게 합니다. 특히 battery 장치에서 최대 capacity와 frequency는 에너지 소비를 늘리고 battery 수명을 줄입니다.
사용자가 `sched_setattr()` syscall로 요청 `uclamp.min`을 바꾸지 않은 RT task에만 효과가 있으며, 위의 `sched_util_clamp_min` 범위 제약을 넘을 수 없습니다.
예를 들어 다음과 같이 설정했다고 가정합니다.
sched_util_clamp_min_rt_default = 800
sched_util_clamp_min = 600
`800`은 허용 범위 `[0:600]` 밖이므로 boost는 `600`으로 clamp됩니다. 절전 모드가 `sched_util_clamp_min`을 바꿔 모든 boost를 임시 제한할 때 이런 상황이 생길 수 있습니다. 제한을 해제하면 요청한 `sched_util_clamp_min_rt_default`가 다시 적용됩니다.
seccomp
1313-1318`Documentation/userspace-api/seccomp_filter.rst`를 참조하십시오.
sg-big-buff
1319-1331generic SCSI(sg) buffer 크기를 표시합니다. 아직 런타임에 조정할 수 없지만 빌드할 때 `include/scsi/sg.h`의 `SG_BIG_BUFF`를 바꿀 수 있습니다. 보통 바꿀 이유는 없습니다.
shmall
1332-1349IPC namespace 안에서 사용할 수 있는 shared memory page 총량입니다. page 수는 namespace마다 따로 계산되고 상속되지 않으므로 `shmall`은 항상 `ceil(shmmax/PAGE_SIZE)` 이상이어야 합니다.
시스템의 기본 `PAGE_SIZE`를 모르면 다음 명령으로 확인합니다.
# getconf PAGE_SIZE
shared memory 할당 능력을 줄이거나 없애려면 새 IPC namespace를 만들고 이 값을 필요한 수준으로 설정한 뒤 현재 user namespace에서 새 IPC namespace 생성을 금지해야 합니다. cgroup을 사용할 수도 있습니다.
shmmax
1350-1358생성 가능한 shared memory segment 최대 크기의 런타임 한도를 조회하고 설정합니다. 커널은 최대 1GiB segment를 지원하며 기본값은 `SHMMAX`입니다.
shmmni
1359-1365shared memory segment의 최대 개수입니다. 기본값은 `4096` (`SHMMNI`)입니다.
shm_rmid_forced
1366-1386`setrlimit(2)`은 한 프로세스의 memory 사용량 같은 자원 한도를 설정하지만, shared memory segment는 어떤 프로세스와도 연결되지 않은 채 존재할 수 있어 한도에 계산되지 않을 수 있습니다.
활성화하면 detach나 프로세스 종료 뒤 attach count가 0이 된 segment를 자동 파괴합니다. 생성했지만 한 번도 attach하지 않은 segment도 생성 프로세스가 종료하면 파괴합니다. `IPC_RMID`는 attach되지 않은 segment를 즉시 파괴하는 용도로만 남습니다.
정의된 System V 동작을 깨므로 일부 application이 멈출 수 있습니다. `RLIMIT_AS`, `RLIMIT_NPROC` 같은 자원 한도도 함께 구성하지 않으면 유용하지 않으며 대부분의 시스템에는 필요하지 않습니다.
`0`에서 `1`로 바꾸면 사용자가 없고 생성한 프로세스가 이미 죽은 기존 segment도 파괴됩니다.
sysctl_writes_strict
1387-1407`/proc/sys`로 sysctl 값을 갱신할 때 file position이 쓰기 동작에 미치는 영향을 제어합니다.
| 값 | 동작 |
|---|---|
| `-1` | legacy per-write 처리. printk 경고 없음. 각 `write` syscall이 값 전체를 포함해야 하며 같은 descriptor의 여러 쓰기는 position과 무관하게 값을 다시 씀 |
| `0` | `-1`과 같지만 file position이 0이 아닐 때 쓰는 프로세스를 경고 |
| `1` | 기본값. 문자열 sysctl은 file position을 존중하고 여러 쓰기를 buffer에 append함. 최대 길이 뒤는 무시함. 숫자 sysctl은 항상 position 0에서 값 전체를 한 write buffer에 담아야 함 |
softlockup_all_cpu_backtrace
1408-1424soft lockup 감지 때 detector thread가 추가 디버깅 정보를 수집할지 제어합니다. 활성화하면 각 CPU에 NMI를 보내 stack trace를 수집합니다. NMI를 지원하는 아키텍처에만 적용됩니다.
| 값 | 동작 |
|---|---|
| `0` | 아무 작업도 하지 않음(기본값) |
| `1` | 감지 시 추가 디버깅 정보 수집 |
softlockup_panic
1425-1438soft lockup 감지 때 kernel panic을 일으킬지 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | panic하지 않음 |
| `1` | panic함 |
`softlockup_panic` kernel parameter로도 설정할 수 있습니다.
soft_watchdog
1439-1456soft lockup detector를 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | soft lockup detector 비활성화 |
| `1` | soft lockup detector 활성화 |
detector는 자발적으로 reschedule하지 않고 CPU를 독점해 `migration/N` thread 실행과 watchdog work를 막는 thread를 감시합니다. watchdog timer function이 work를 queue하려면 CPU가 timer interrupt에 응답해야 합니다. 그렇지 않으면 NMI watchdog이 활성화된 경우 hard lockup을 감지할 수 있습니다.
split_lock_mitigate (x86 only)
1457-1479x86의 split lock 하나마다 시스템 전체 성능 비용이 생깁니다. 큰 시스템에서 비특권 사용자가 많은 split lock을 만들면 정상 사용자에게 denial of service를 일으킬 수 있습니다.
커널은 split lock을 감지해 기다리게 하고 한 번에 core 하나만 split lock을 실행하도록 벌점을 줍니다. 문제가 있는 application이 지나치게 느려질 수 있지만 완화를 끄면 split lock 사용자의 denial of service에 더 노출됩니다.
| 값 | 동작 |
|---|---|
| `0` | 완화 비활성화. kernel log에 경고만 남기며 denial of service에 노출 |
| `1` | 완화 활성화(기본값). 의도적 성능 저하로 split lock 사용자를 제약 |
stack_erasing
1480-1497`CONFIG_KSTACK_ERASE`로 빌드한 커널에서 syscall 끝의 kernel stack 지우기를 제어합니다. stack leak bug가 드러낼 수 있는 정보를 줄이고 초기화되지 않은 stack 변수 공격 일부를 막습니다.
대가로 성능이 떨어집니다. single-CPU 시스템의 커널 컴파일에서는 약 1% 느려졌으며 다른 시스템과 workload는 달라질 수 있습니다.
| 값 | 동작 |
|---|---|
| `0` | kernel stack 지우기 비활성화, `KSTACK_ERASE_METRICS`도 갱신하지 않음 |
| `1` | 활성화(기본값). syscall 끝에 사용자 공간으로 돌아가기 전에 수행 |
stop-a (SPARC only)
1498-1511Stop-A 동작을 제어합니다.
| 값 | 동작 |
|---|---|
| `0` | Stop-A가 효과 없음 |
| `1` | PROM으로 진입(기본값) |
panic 때는 사용자가 boot PROM으로 돌아갈 수 있도록 Stop-A가 항상 활성화됩니다.
sysrq
1512-1517`Documentation/admin-guide/sysrq.rst`를 참조하십시오.
tainted
1518-1555커널이 taint됐으면 0이 아닌 값입니다. 숫자는 OR할 수 있고 문자는 Oops 보고서의 `Tainted` 줄에 나타납니다.
| 값 | 문자 | 의미 |
|---|---|---|
| `1` | `P` | proprietary module 적재 |
| `2` | `F` | module 강제 적재 |
| `4` | `S` | 규격 밖 시스템에서 실행 |
| `8` | `R` | module 강제 제거 |
| `16` | `M` | processor가 Machine Check Exception(MCE) 보고 |
| `32` | `B` | 잘못된 page 참조 또는 예상 밖 page flag |
| `64` | `U` | 사용자 공간 application이 taint 요청 |
| `128` | `D` | 최근 OOPS 또는 BUG로 kernel 사망 |
| `256` | `A` | 사용자가 ACPI table 덮어씀 |
| `512` | `W` | kernel warning 발생 |
| `1024` | `C` | staging driver 적재 |
| `2048` | `I` | platform firmware bug workaround 적용 |
| `4096` | `O` | 외부 빌드(out-of-tree) module 적재 |
| `8192` | `E` | 서명되지 않은 module 적재 |
| `16384` | `L` | soft lockup 발생 |
| `32768` | `K` | kernel live patch 적용 |
| `65536` | `X` | 배포판이 정의하고 사용하는 auxiliary taint |
| `131072` | `T` | struct randomization plugin으로 kernel 빌드 |
자세한 내용은 `Documentation/admin-guide/tainted-kernels.rst`를 참조하십시오.
command line에 `panic_on_taint=<bitmask>,nousertaint`를 지정해 부팅했고 `tainted`에 쓰려는 OR 값이 `panic_on_taint` bitmask와 겹치면 쓰기가 `EINVAL`로 실패합니다. 자세한 내용은 `Documentation/admin-guide/kernel-parameters.rst`의 해당 command line 옵션과 `nousertaint` switch를 참조하십시오.
threads-max
1556-1573`fork()`로 만들 수 있는 thread의 최대 개수입니다. 초기화 때 커널은 최대 thread를 만들어도 thread 구조체가 사용 가능한 RAM page의 1/8만 차지하도록 값을 정합니다.
쓸 수 있는 최솟값은 `1`, 최댓값은 `FUTEX_TID_MASK` (`0x3fffffff`)입니다. 범위 밖 값을 쓰면 `EINVAL`이 발생합니다.
timer_migration
1574-15810이 아니면 idle CPU의 timer를 다른 CPU로 옮겨 저전력 상태를 더 오래 유지하도록 시도합니다. 기본값은 `1`입니다.
traceoff_on_warning
1582-1588설정하면 `WARN()` 발생 때 tracing을 끕니다. tracing은 `Documentation/trace/ftrace.rst`를 참조하십시오.
tracepoint_printk
1589-1608`tp_printk` boot parameter로 tracepoint를 `printk()`에 보내도록 켰을 때 런타임 동작을 제어합니다. 다음은 전송을 중지합니다.
echo 0 > /proc/sys/kernel/tracepoint_printk
다음은 다시 `printk()`로 보냅니다.
echo 1 > /proc/sys/kernel/tracepoint_printk
커널을 `tp_printk` 활성 상태로 부팅한 경우에만 동작합니다. `Documentation/admin-guide/kernel-parameters.rst`와 `Documentation/trace/boottime-trace.rst`를 참조하십시오.
unaligned-trap
1609-1624unaligned 접근이 trap을 일으키고 `CONFIG_SYSCTL_ARCH_UNALIGN_ALLOW`를 지원하는 아키텍처, 현재 `arc`, `parisc`, `loongarch`에서 trap을 잡아 실패 대신 emulate할지 정합니다.
| 값 | 동작 |
|---|---|
| `0` | unaligned 접근을 emulate하지 않음 |
| `1` | unaligned 접근을 emulate함(기본값) |
`ignore-unaligned-usertrap` 절도 참조하십시오.
unknown_nmi_panic
1625-1635NMI 처리 동작에 영향을 줍니다. 0이 아니면 알 수 없는 NMI를 trap한 뒤 panic하고 kernel 디버깅 정보를 console에 표시합니다. 많은 IA32 server의 NMI switch가 이런 NMI를 발생시키므로 시스템이 멈췄을 때 switch를 눌러 볼 수 있습니다.
unprivileged_bpf_disabled
1636-1657`1`을 쓰면 비특권 `bpf()` 호출을 비활성화합니다. 이후 `CAP_SYS_ADMIN` 또는 `CAP_BPF` 없는 호출은 `-EPERM`을 반환하며 실행 중인 커널에서 다시 해제할 수 없습니다.
`2`도 비특권 호출을 끄지만 관리자가 나중에 `0` 또는 `1`로 바꿀 수 있습니다. 커널 구성에 `BPF_UNPRIV_DEFAULT_OFF`가 켜졌다면 기본값은 `0` 대신 `2`입니다.
| 값 | 동작 |
|---|---|
| `0` | 비특권 `bpf()` 호출 활성화 |
| `1` | 비특권 `bpf()` 호출을 복구 불가능하게 비활성화 |
| `2` | 비특권 `bpf()` 호출 비활성화, 관리자가 변경 가능 |
warn_limit
1658-1666`panic_on_warn`이 설정되지 않았을 때 몇 번째 kernel warning 뒤에 panic할지 정합니다. `0`은 횟수 검사를 끄고 `1`은 `panic_on_warn=1`과 같습니다. 기본값은 `0`입니다.
watchdog
1667-1688soft lockup detector와 NMI watchdog(hard lockup detector)을 동시에 켜거나 끕니다.
| 값 | 동작 |
|---|---|
| `0` | 두 lockup detector 모두 비활성화 |
| `1` | 두 lockup detector 모두 활성화 |
`soft_watchdog`와 `nmi_watchdog`로 각각 따로 제어할 수도 있습니다. 다음처럼 `watchdog`를 읽으면 두 값의 논리 OR 결과인 `0` 또는 `1`을 출력합니다.
cat /proc/sys/kernel/watchdog
watchdog_cpumask
1689-1709watchdog가 실행될 CPU를 제어합니다. 기본 cpumask는 가능한 모든 core입니다. 커널에 `NO_HZ_FULL`이 켜졌고 `nohz_full=` boot argument로 core를 지정했다면 해당 core는 기본적으로 제외됩니다.
offline core도 mask에 넣을 수 있으며 나중에 online이 되면 mask에 따라 watchdog을 시작합니다. 보통 `nohz_full` 환경에서 watchdog이 기본적으로 없던 core의 kernel lockup이 의심될 때 다시 활성화하는 용도로만 바꿉니다.
값은 표준 cpumask cpulist 형식입니다. 예를 들어 core 0, 2, 3, 4에서 watchdog을 켜려면 다음과 같이 씁니다.
echo 0,2-4 > /proc/sys/kernel/watchdog_cpumask
watchdog_thresh
1710-1718hrtimer와 NMI event frequency, soft·hard lockup threshold를 제어합니다. 기본 threshold는 10초이고 soft lockup threshold는 `2 * watchdog_thresh`입니다. `0`으로 설정하면 lockup detection을 모두 끕니다.
요약과 해설
kernel.rst:1-1718이 문서는 Linux 커널의 전역 동작을 런타임에 바꾸는 `kernel` sysctl을 다룹니다. 값 하나가 장애 진단, 보안 경계, 재부팅 정책, 성능 계측과 IPC 자원에 직접 영향을 줄 수 있으므로 각 항목의 단위와 되돌릴 수 있는지 여부를 함께 확인해야 합니다.
`kexec_load_disabled`, `modules_disabled`, `unprivileged_bpf_disabled=1`처럼 한 번 강화하면 실행 중 되돌릴 수 없는 설정이 있습니다. 운영 변경 전에는 문서의 기본값, 커널 구성 조건, capability 요구 사항과 장애 시 복구 경로를 반드시 검토해야 합니다.