요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
Race and topology tests
memcg_test.rst:140-244Small limit, shmem, migration, hotplug와 nested test를 다룹니다.
Swap, OOM and notifications
memcg_test.rst:245-344Swapoff, OOM, charge migration과 threshold notification을 정리합니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
=====================================================
Memory Resource Controller(Memcg) Implementation Memo
=====================================================
Last Updated: 2010/2
Base Kernel Version: based on 2.6.33-rc7-mm(candidate for 34).
Because VM is getting complex (one of reasons is memcg...), memcg's behavior
is complex. This is a document for memcg's internal behavior.
Please note that implementation details can be changed.
(*) Topics on API should be in Documentation/admin-guide/cgroup-v1/memory.rst)
0. How to record usage ?
========================
2 objects are used.
page_cgroup ....an object per page.
Allocated at boot or memory hotplug. Freed at memory hot removal.
swap_cgroup ... an entry per swp_entry.
Allocated at swapon(). Freed at swapoff().
The page_cgroup has USED bit and double count against a page_cgroup never
occurs. swap_cgroup is used only when a charged page is swapped-out.
1. Charge
=========
a page/swp_entry may be charged (usage += PAGE_SIZE) at
mem_cgroup_try_charge()
2. Uncharge
===========
a page/swp_entry may be uncharged (usage -= PAGE_SIZE) by
mem_cgroup_uncharge()
Called when a page's refcount goes down to 0.
mem_cgroup_uncharge_swap()
Called when swp_entry's refcnt goes down to 0. A charge against swap
disappears.
3. charge-commit-cancel
=======================
Memcg pages are charged in two steps:
- mem_cgroup_try_charge()
- mem_cgroup_commit_charge() or mem_cgroup_cancel_charge()
At try_charge(), there are no flags to say "this page is charged".
at this point, usage += PAGE_SIZE.
At commit(), the page is associated with the memcg.
At cancel(), simply usage -= PAGE_SIZE.
Under below explanation, we assume CONFIG_SWAP=y.
4. Anonymous
============
Anonymous page is newly allocated at
- page fault into MAP_ANONYMOUS mapping.
- Copy-On-Write.
4.1 Swap-in.
At swap-in, the page is taken from swap-cache. There are 2 cases.
(a) If the SwapCache is newly allocated and read, it has no charges.
(b) If the SwapCache has been mapped by processes, it has been
charged already.
4.2 Swap-out.
At swap-out, typical state transition is below.
(a) add to swap cache. (marked as SwapCache)
swp_entry's refcnt += 1.
(b) fully unmapped.
swp_entry's refcnt += # of ptes.
(c) write back to swap.
(d) delete from swap cache. (remove from SwapCache)
swp_entry's refcnt -= 1.
Finally, at task exit,
(e) zap_pte() is called and swp_entry's refcnt -=1 -> 0.
5. Page Cache
=============
Page Cache is charged at
- filemap_add_folio().
The logic is very clear. (About migration, see below)
Note:
__filemap_remove_folio() is called by filemap_remove_folio()
and __remove_mapping().
6. Shmem(tmpfs) Page Cache
===========================
The best way to understand shmem's page state transition is to read
mm/shmem.c.
But brief explanation of the behavior of memcg around shmem will be
helpful to understand the logic.
Shmem's page (just leaf page, not direct/indirect block) can be on
- radix-tree of shmem's inode.
- SwapCache.
- Both on radix-tree and SwapCache. This happens at swap-in
and swap-out,
It's charged when...
- A new page is added to shmem's radix-tree.
- A swp page is read. (move a charge from swap_cgroup to page_cgroup)
7. Page Migration
=================
mem_cgroup_migrate()
8. LRU
======
Each memcg has its own vector of LRUs (inactive anon, active anon,
inactive file, active file, unevictable) of pages from each node,
each LRU handled under a single lru_lock for that memcg and node.
9. Typical Tests.
=================
Tests for racy cases.
9.1 Small limit to memcg.
-------------------------
When you do test to do racy case, it's good test to set memcg's limit
to be very small rather than GB. Many races found in the test under
xKB or xxMB limits.
(Memory behavior under GB and Memory behavior under MB shows very
different situation.)
9.2 Shmem
---------
Historically, memcg's shmem handling was poor and we saw some amount
of troubles here. This is because shmem is page-cache but can be
SwapCache. Test with shmem/tmpfs is always good test.
9.3 Migration
-------------
For NUMA, migration is an another special case. To do easy test, cpuset
is useful. Following is a sample script to do migration::
mount -t cgroup -o cpuset none /opt/cpuset
mkdir /opt/cpuset/01
echo 1 > /opt/cpuset/01/cpuset.cpus
echo 0 > /opt/cpuset/01/cpuset.mems
echo 1 > /opt/cpuset/01/cpuset.memory_migrate
mkdir /opt/cpuset/02
echo 1 > /opt/cpuset/02/cpuset.cpus
echo 1 > /opt/cpuset/02/cpuset.mems
echo 1 > /opt/cpuset/02/cpuset.memory_migrate
In above set, when you moves a task from 01 to 02, page migration to
node 0 to node 1 will occur. Following is a script to migrate all
under cpuset.::
--
move_task()
{
for pid in $1
do
/bin/echo $pid >$2/tasks 2>/dev/null
echo -n $pid
echo -n " "
done
echo END
}
G1_TASK=`cat ${G1}/tasks`
G2_TASK=`cat ${G2}/tasks`
move_task "${G1_TASK}" ${G2} &
--
9.4 Memory hotplug
------------------
memory hotplug test is one of good test.
to offline memory, do following::
# echo offline > /sys/devices/system/memory/memoryXXX/state
(XXX is the place of memory)
This is an easy way to test page migration, too.
9.5 nested cgroups
------------------
Use tests like the following for testing nested cgroups::
mkdir /opt/cgroup/01/child_a
mkdir /opt/cgroup/01/child_b
set limit to 01.
add limit to 01/child_b
run jobs under child_a and child_b
create/delete following groups at random while jobs are running::
/opt/cgroup/01/child_a/child_aa
/opt/cgroup/01/child_b/child_bb
/opt/cgroup/01/child_c
running new jobs in new group is also good.
9.6 Mount with other subsystems
-------------------------------
Mounting with other subsystems is a good test because there is a
race and lock dependency with other cgroup subsystems.
example::
# mount -t cgroup none /cgroup -o cpuset,memory,cpu,devices
and do task move, mkdir, rmdir etc...under this.
9.7 swapoff
-----------
Besides management of swap is one of complicated parts of memcg,
call path of swap-in at swapoff is not same as usual swap-in path..
It's worth to be tested explicitly.
For example, test like following is good:
(Shell-A)::
# mount -t cgroup none /cgroup -o memory
# mkdir /cgroup/test
# echo 40M > /cgroup/test/memory.limit_in_bytes
# echo 0 > /cgroup/test/tasks
Run malloc(100M) program under this. You'll see 60M of swaps.
(Shell-B)::
# move all tasks in /cgroup/test to /cgroup
# /sbin/swapoff -a
# rmdir /cgroup/test
# kill malloc task.
Of course, tmpfs v.s. swapoff test should be tested, too.
9.8 OOM-Killer
--------------
Out-of-memory caused by memcg's limit will kill tasks under
the memcg. When hierarchy is used, a task under hierarchy
will be killed by the kernel.
In this case, panic_on_oom shouldn't be invoked and tasks
in other groups shouldn't be killed.
It's not difficult to cause OOM under memcg as following.
Case A) when you can swapoff::
#swapoff -a
#echo 50M > /memory.limit_in_bytes
run 51M of malloc
Case B) when you use mem+swap limitation::
#echo 50M > memory.limit_in_bytes
#echo 50M > memory.memsw.limit_in_bytes
run 51M of malloc
9.9 Move charges at task migration
----------------------------------
Charges associated with a task can be moved along with task migration.
(Shell-A)::
#mkdir /cgroup/A
#echo $$ >/cgroup/A/tasks
run some programs which uses some amount of memory in /cgroup/A.
(Shell-B)::
#mkdir /cgroup/B
#echo 1 >/cgroup/B/memory.move_charge_at_immigrate
#echo "pid of the program running in group A" >/cgroup/B/tasks
You can see charges have been moved by reading ``*.usage_in_bytes`` or
memory.stat of both A and B.
See 8.2 of Documentation/admin-guide/cgroup-v1/memory.rst to see what value should
be written to move_charge_at_immigrate.
9.10 Memory thresholds
----------------------
Memory controller implements memory thresholds using cgroups notification
API. You can use tools/cgroup/cgroup_event_listener.c to test it.
(Shell-A) Create cgroup and run event listener::
# mkdir /cgroup/A
# ./cgroup_event_listener /cgroup/A/memory.usage_in_bytes 5M
(Shell-B) Add task to cgroup and try to allocate and free memory::
# echo $$ >/cgroup/A/tasks
# a="$(dd if=/dev/zero bs=1M count=10)"
# a=
You will see message from cgroup_event_listener every time you cross
the thresholds.
Use /cgroup/A/memory.memsw.usage_in_bytes to test memsw thresholds.
It's good idea to test root cgroup as well.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
구현 메모의 기준과 범위
1-14이 문서는 2010년 2월에 갱신된 Memory Resource Controller(Memcg) 구현 메모이며 base kernel은 `2.6.33-rc7-mm(candidate for 34)`입니다. VM과 memcg behavior가 복잡해 internal 동작을 설명하지만 implementation detail은 바뀔 수 있습니다.
이 문서와 API 문서의 역할을 구분합니다.
page_cgroup과 swap_cgroup
15-30Usage 기록에는 page마다 하나인 `page_cgroup` object와 `swp_entry`마다 하나인 `swap_cgroup` entry를 사용합니다.
Object의 allocation·free 시점과 역할입니다.
`USED` bit가 page double charge를 막고 swap-out 때 swap_cgroup으로 추적합니다.
Charge와 uncharge entry points
31-49Page 또는 `swp_entry` charge는 `mem_cgroup_try_charge()`에서 `usage += PAGE_SIZE`로 수행합니다.
Reference가 사라질 때 object 유형에 맞춰 usage를 내립니다.
Page와 swap entry는 각각 reference lifetime 끝에서 uncharge됩니다.
try·commit·cancel 두 단계 charge
50-66Memcg page charge는 `mem_cgroup_try_charge()` 뒤 `mem_cgroup_commit_charge()` 또는 `mem_cgroup_cancel_charge()`로 끝나는 두 단계 transaction입니다. 아래 설명은 `CONFIG_SWAP=y`를 가정합니다.
Try 시 usage를 먼저 올리고 성공 여부에 따라 page association 또는 rollback을 수행합니다.
Anonymous page swap-in/out state
67-95Anonymous page는 `MAP_ANONYMOUS` mapping의 page fault 또는 Copy-On-Write에서 새로 allocate됩니다. Swap-in 때 page는 swap cache에서 오며, 새로 allocate해 read한 SwapCache라 charge가 없거나 이미 process에 map되어 charge된 두 경우가 있습니다.
Swap entry reference count가 cache와 PTE lifetime을 함께 추적합니다.
Page cache와 shmem charge
96-128일반 page cache는 `filemap_add_folio()`에서 charge됩니다. Migration은 뒤 절에서 다룹니다. `__filemap_remove_folio()`는 `filemap_remove_folio()`와 `__remove_mapping()`에서 호출됩니다.
일반 file page와 shmem page의 charge event입니다.
Shmem의 자세한 state transition은 `mm/shmem.c`에서 확인합니다. Leaf shmem page는 inode radix-tree, SwapCache 또는 swap-in/out 중 두 곳 모두에 존재할 수 있습니다.
Swap-backed page가 memory로 돌아올 때 charge owner object도 이동합니다.
Page migration과 per-memcg LRU
129-139Page migration entry point는 `mem_cgroup_migrate()`입니다.
각 memcg는 node별 LRU vector를 가지며 memcg·node 조합마다 하나의 lru_lock으로 보호합니다.
작은 limit·shmem·NUMA migration test
140-199Race test에서는 GB보다 xKB 또는 xxMB처럼 작은 memcg limit이 더 효과적입니다. GB와 MB 범위의 memory behavior가 매우 달라 작은 limit test에서 많은 race가 발견됐습니다.
Shmem은 page cache이면서 SwapCache가 될 수 있어 역사적으로 memcg handling 문제가 많았습니다. 따라서 shmem/tmpfs test는 항상 가치가 있습니다.
초기 세 test가 겨냥하는 취약 지점입니다.
is useful. Following is a sample script to do migration::
mount -t cgroup -o cpuset none /opt/cpuset
mkdir /opt/cpuset/01
echo 1 > /opt/cpuset/01/cpuset.cpus
echo 0 > /opt/cpuset/01/cpuset.mems
echo 1 > /opt/cpuset/01/cpuset.memory_migrate
mkdir /opt/cpuset/02
echo 1 > /opt/cpuset/02/cpuset.cpus
echo 1 > /opt/cpuset/02/cpuset.mems
echo 1 > /opt/cpuset/02/cpuset.memory_migrate
In above set, when you moves a task from 01 to 02, page migration to
node 0 to node 1 will occur. Following is a script to migrate all
under cpuset.::
--
move_task()
{
for pid in $1
do
/bin/echo $pid >$2/tasks 2>/dev/null
echo -n $pid
echo -n " "
done
echo END
}
G1_TASK=`cat ${G1}/tasks`
G2_TASK=`cat ${G2}/tasks`
move_task "${G1_TASK}" ${G2} &
Task를 cpuset 01에서 02로 옮겨 node 0 page를 node 1로 migration합니다.
Memory hotplug·nested cgroup·mixed controller test
200-244Memory hotplug은 page migration을 쉽게 시험하는 방법입니다. `/sys/devices/system/memory/memoryXXX/state`에 `offline`을 써서 XXX 위치 memory를 offline합니다.
memory hotplug test is one of good test.
to offline memory, do following::
# echo offline > /sys/devices/system/memory/memoryXXX/state
Nested test는 `/opt/cgroup/01/child_a`와 `child_b`를 만들고 parent 01과 child_b에 limit을 둔 뒤 두 child에서 job을 실행합니다. Job이 도는 동안 `child_aa`, `child_bb`, `child_c`를 무작위 생성·삭제하고 새 group에도 job을 시작합니다.
Dynamic hierarchy와 다른 controller lock dependency를 함께 자극합니다.
9.6 Mount with other subsystems
-------------------------------
Mounting with other subsystems is a good test because there is a
race and lock dependency with other cgroup subsystems.
example::
# mount -t cgroup none /cgroup -o cpuset,memory,cpu,devices
and do task move, mkdir, rmdir etc...under this.
Task move와 hierarchy mutation이 controller 간 lock dependency를 드러냅니다.
swapoff와 memcg OOM test
245-297Swap management는 memcg의 복잡한 부분이며 `swapoff` 중 swap-in call path는 일반 swap-in과 다르므로 명시적으로 test해야 합니다.
(Shell-A)::
# mount -t cgroup none /cgroup -o memory
# mkdir /cgroup/test
# echo 40M > /cgroup/test/memory.limit_in_bytes
# echo 0 > /cgroup/test/tasks
Run malloc(100M) program under this. You'll see 60M of swaps.
(Shell-B)::
# move all tasks in /cgroup/test to /cgroup
# /sbin/swapoff -a
# rmdir /cgroup/test
# kill malloc task.
Of course, tmpfs v.s. swapoff test should be tested, too.
40M limit 아래 100M allocation으로 60M swap을 만든 뒤 task 이동과 swapoff cleanup을 시험합니다.
Tmpfs와 swapoff 조합도 함께 test해야 합니다. Memcg limit으로 발생한 OOM은 해당 memcg 또는 hierarchy 아래 task를 kill해야 하며 `panic_on_oom`을 호출하거나 다른 group의 task를 kill하면 안 됩니다.
9.8 OOM-Killer
--------------
Out-of-memory caused by memcg's limit will kill tasks under
the memcg. When hierarchy is used, a task under hierarchy
will be killed by the kernel.
In this case, panic_on_oom shouldn't be invoked and tasks
in other groups shouldn't be killed.
It's not difficult to cause OOM under memcg as following.
Case A) when you can swapoff::
#swapoff -a
#echo 50M > /memory.limit_in_bytes
run 51M of malloc
Case B) when you use mem+swap limitation::
#echo 50M > memory.limit_in_bytes
#echo 50M > memory.memsw.limit_in_bytes
run 51M of malloc
Swap을 없애거나 mem+swap limit을 함께 제한해 51M allocation으로 OOM을 만듭니다.
Charge migration과 threshold notification
298-344Task migration과 함께 task에 연결된 charge도 옮길 수 있습니다. Group A에서 memory를 사용한 program을 실행하고 group B의 `memory.move_charge_at_immigrate`를 1로 설정한 뒤 PID를 B의 `tasks`에 쓰면 charge가 이동합니다.
(Shell-A)::
#mkdir /cgroup/A
#echo $$ >/cgroup/A/tasks
run some programs which uses some amount of memory in /cgroup/A.
(Shell-B)::
#mkdir /cgroup/B
#echo 1 >/cgroup/B/memory.move_charge_at_immigrate
#echo "pid of the program running in group A" >/cgroup/B/tasks
You can see charges have been moved by reading ``*.usage_in_bytes`` or
memory.stat of both A and B.
`*.usage_in_bytes`와 `memory.stat`에서 A 감소와 B 증가를 확인합니다.
`move_charge_at_immigrate`에 쓸 값의 의미는 `Documentation/admin-guide/cgroup-v1/memory.rst` 8.2절을 참고합니다.
Memory controller threshold는 cgroup notification API로 구현됩니다. `tools/cgroup/cgroup_event_listener.c`를 사용해 memory usage가 5M threshold를 오르내릴 때마다 event message가 오는지 시험합니다.
9.10 Memory thresholds
----------------------
Memory controller implements memory thresholds using cgroups notification
API. You can use tools/cgroup/cgroup_event_listener.c to test it.
(Shell-A) Create cgroup and run event listener::
# mkdir /cgroup/A
# ./cgroup_event_listener /cgroup/A/memory.usage_in_bytes 5M
(Shell-B) Add task to cgroup and try to allocate and free memory::
# echo $$ >/cgroup/A/tasks
# a="$(dd if=/dev/zero bs=1M count=10)"
# a=
10M를 allocate했다가 free해 threshold 양방향 crossing을 만듭니다.
같은 notification path를 다른 accounting scope에도 적용합니다.
Accounting internals
memcg_test.rst:1-139Charge object와 page/swap/shmem state transition을 설명합니다.