← Documents Documentation/bpf/map_array.rst GitHub 원문 ↗

Linux 6.18.37 · BPF

BPF_MAP_TYPE_ARRAY and BPF_MAP_TYPE_PERCPU_ARRAY

일반·per-CPU BPF array map의 allocation, mmap, kernel helper, concurrency와 userspace 생성·초기화·조회 semantics를 설명합니다.

Source pathDocumentation/bpf/map_array.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약과 해설

map_array.rst:1-262

Array map은 unsigned 32-bit index와 고정된 `max_entries`를 사용하며 모든 element를 미리 zero-initialize합니다. 일반 array는 shared storage, per-CPU array는 CPU별 storage를 제공합니다.

Kernel BPF helper는 value pointer를 직접 반환하므로 shared update에는 atomic primitive나 `bpf_spin_lock`이 필요합니다. Userspace libbpf API는 fd로 map을 식별하고 per-CPU value를 `ncpus` 길이의 array로 주고받습니다.

`BPF_F_MMAPABLE`은 userspace의 direct access를 가능하게 하고, 고정 size array는 delete 대신 zero update로 element를 clear합니다. Per-CPU array update에는 `BPF_NOEXIST`를 사용할 수 없습니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0-only
2 .. Copyright (C) 2022 Red Hat, Inc.
3
4 ================================================
5 BPF_MAP_TYPE_ARRAY and BPF_MAP_TYPE_PERCPU_ARRAY
6 ================================================
7
8 .. note::
9 - ``BPF_MAP_TYPE_ARRAY`` was introduced in kernel version 3.19
10 - ``BPF_MAP_TYPE_PERCPU_ARRAY`` was introduced in version 4.6
11
12 ``BPF_MAP_TYPE_ARRAY`` and ``BPF_MAP_TYPE_PERCPU_ARRAY`` provide generic array
13 storage. The key type is an unsigned 32-bit integer (4 bytes) and the map is
14 of constant size. The size of the array is defined in ``max_entries`` at
15 creation time. All array elements are pre-allocated and zero initialized when
16 created. ``BPF_MAP_TYPE_PERCPU_ARRAY`` uses a different memory region for each
17 CPU whereas ``BPF_MAP_TYPE_ARRAY`` uses the same memory region. The value
18 stored can be of any size, however, all array elements are aligned to 8
19 bytes.
20
21 Since kernel 5.5, memory mapping may be enabled for ``BPF_MAP_TYPE_ARRAY`` by
22 setting the flag ``BPF_F_MMAPABLE``. The map definition is page-aligned and
23 starts on the first page. Sufficient page-sized and page-aligned blocks of
24 memory are allocated to store all array values, starting on the second page,
25 which in some cases will result in over-allocation of memory. The benefit of
26 using this is increased performance and ease of use since userspace programs
27 would not be required to use helper functions to access and mutate data.
28
29 Usage
30 =====
31
32 Kernel BPF
33 ----------
34
35 bpf_map_lookup_elem()
36 ~~~~~~~~~~~~~~~~~~~~~
37
38 .. code-block:: c
39
40 void *bpf_map_lookup_elem(struct bpf_map *map, const void *key)
41
42 Array elements can be retrieved using the ``bpf_map_lookup_elem()`` helper.
43 This helper returns a pointer into the array element, so to avoid data races
44 with userspace reading the value, the user must use primitives like
45 ``__sync_fetch_and_add()`` when updating the value in-place.
46
47 bpf_map_update_elem()
48 ~~~~~~~~~~~~~~~~~~~~~
49
50 .. code-block:: c
51
52 long bpf_map_update_elem(struct bpf_map *map, const void *key, const void *value, u64 flags)
53
54 Array elements can be updated using the ``bpf_map_update_elem()`` helper.
55
56 ``bpf_map_update_elem()`` returns 0 on success, or negative error in case of
57 failure.
58
59 Since the array is of constant size, ``bpf_map_delete_elem()`` is not supported.
60 To clear an array element, you may use ``bpf_map_update_elem()`` to insert a
61 zero value to that index.
62
63 Per CPU Array
64 -------------
65
66 Values stored in ``BPF_MAP_TYPE_ARRAY`` can be accessed by multiple programs
67 across different CPUs. To restrict storage to a single CPU, you may use a
68 ``BPF_MAP_TYPE_PERCPU_ARRAY``.
69
70 When using a ``BPF_MAP_TYPE_PERCPU_ARRAY`` the ``bpf_map_update_elem()`` and
71 ``bpf_map_lookup_elem()`` helpers automatically access the slot for the current
72 CPU.
73
74 bpf_map_lookup_percpu_elem()
75 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~
76
77 .. code-block:: c
78
79 void *bpf_map_lookup_percpu_elem(struct bpf_map *map, const void *key, u32 cpu)
80
81 The ``bpf_map_lookup_percpu_elem()`` helper can be used to lookup the array
82 value for a specific CPU. Returns value on success , or ``NULL`` if no entry was
83 found or ``cpu`` is invalid.
84
85 Concurrency
86 -----------
87
88 Since kernel version 5.1, the BPF infrastructure provides ``struct bpf_spin_lock``
89 to synchronize access.
90
91 Userspace
92 ---------
93
94 Access from userspace uses libbpf APIs with the same names as above, with
95 the map identified by its ``fd``.
96
97 Examples
98 ========
99
100 Please see the ``tools/testing/selftests/bpf`` directory for functional
101 examples. The code samples below demonstrate API usage.
102
103 Kernel BPF
104 ----------
105
106 This snippet shows how to declare an array in a BPF program.
107
108 .. code-block:: c
109
110 struct {
111 __uint(type, BPF_MAP_TYPE_ARRAY);
112 __type(key, u32);
113 __type(value, long);
114 __uint(max_entries, 256);
115 } my_map SEC(".maps");
116
117
118 This example BPF program shows how to access an array element.
119
120 .. code-block:: c
121
122 int bpf_prog(struct __sk_buff *skb)
123 {
124 struct iphdr ip;
125 int index;
126 long *value;
127
128 if (bpf_skb_load_bytes(skb, ETH_HLEN, &ip, sizeof(ip)) < 0)
129 return 0;
130
131 index = ip.protocol;
132 value = bpf_map_lookup_elem(&my_map, &index);
133 if (value)
134 __sync_fetch_and_add(value, skb->len);
135
136 return 0;
137 }
138
139 Userspace
140 ---------
141
142 BPF_MAP_TYPE_ARRAY
143 ~~~~~~~~~~~~~~~~~~
144
145 This snippet shows how to create an array, using ``bpf_map_create_opts`` to
146 set flags.
147
148 .. code-block:: c
149
150 #include <bpf/libbpf.h>
151 #include <bpf/bpf.h>
152
153 int create_array()
154 {
155 int fd;
156 LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_MMAPABLE);
157
158 fd = bpf_map_create(BPF_MAP_TYPE_ARRAY,
159 "example_array", /* name */
160 sizeof(__u32), /* key size */
161 sizeof(long), /* value size */
162 256, /* max entries */
163 &opts); /* create opts */
164 return fd;
165 }
166
167 This snippet shows how to initialize the elements of an array.
168
169 .. code-block:: c
170
171 int initialize_array(int fd)
172 {
173 __u32 i;
174 long value;
175 int ret;
176
177 for (i = 0; i < 256; i++) {
178 value = i;
179 ret = bpf_map_update_elem(fd, &i, &value, BPF_ANY);
180 if (ret < 0)
181 return ret;
182 }
183
184 return ret;
185 }
186
187 This snippet shows how to retrieve an element value from an array.
188
189 .. code-block:: c
190
191 int lookup(int fd)
192 {
193 __u32 index = 42;
194 long value;
195 int ret;
196
197 ret = bpf_map_lookup_elem(fd, &index, &value);
198 if (ret < 0)
199 return ret;
200
201 /* use value here */
202 assert(value == 42);
203
204 return ret;
205 }
206
207 BPF_MAP_TYPE_PERCPU_ARRAY
208 ~~~~~~~~~~~~~~~~~~~~~~~~~
209
210 This snippet shows how to initialize the elements of a per CPU array.
211
212 .. code-block:: c
213
214 int initialize_array(int fd)
215 {
216 int ncpus = libbpf_num_possible_cpus();
217 long values[ncpus];
218 __u32 i, j;
219 int ret;
220
221 for (i = 0; i < 256 ; i++) {
222 for (j = 0; j < ncpus; j++)
223 values[j] = i;
224 ret = bpf_map_update_elem(fd, &i, &values, BPF_ANY);
225 if (ret < 0)
226 return ret;
227 }
228
229 return ret;
230 }
231
232 This snippet shows how to access the per CPU elements of an array value.
233
234 .. code-block:: c
235
236 int lookup(int fd)
237 {
238 int ncpus = libbpf_num_possible_cpus();
239 __u32 index = 42, j;
240 long values[ncpus];
241 int ret;
242
243 ret = bpf_map_lookup_elem(fd, &index, &values);
244 if (ret < 0)
245 return ret;
246
247 for (j = 0; j < ncpus; j++) {
248 /* Use per CPU value here */
249 assert(values[j] == 42);
250 }
251
252 return ret;
253 }
254
255 Semantics
256 =========
257
258 As shown in the example above, when accessing a ``BPF_MAP_TYPE_PERCPU_ARRAY``
259 in userspace, each value is an array with ``ncpus`` elements.
260
261 When calling ``bpf_map_update_elem()`` the flag ``BPF_NOEXIST`` can not be used
262 for these maps.
263

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Array map 구조와 mmap 지원

1-28

`BPF_MAP_TYPE_ARRAY and BPF_MAP_TYPE_PERCPU_ARRAY` 문서는 `GPL-2.0-only` license를 따르며 `Copyright (C) 2022 Red Hat, Inc.`를 명시합니다.

`BPF_MAP_TYPE_ARRAY`는 `kernel version 3.19`에 도입되었고 `BPF_MAP_TYPE_PERCPU_ARRAY`는 `version 4.6`에 도입되었습니다.

`BPF_MAP_TYPE_ARRAY`와 `BPF_MAP_TYPE_PERCPU_ARRAY`는 generic array storage를 제공합니다. Key type은 4-byte unsigned 32-bit integer이고 map size는 고정입니다. Array size는 생성할 때 `max_entries`로 정하며 모든 element를 미리 allocate하고 zero initialize합니다.

`BPF_MAP_TYPE_PERCPU_ARRAY`는 CPU마다 서로 다른 memory region을 사용하지만 `BPF_MAP_TYPE_ARRAY`는 모든 CPU가 같은 region을 사용합니다. Value는 어떤 size든 저장할 수 있으나 모든 array element는 8 byte에 align됩니다.

Kernel 5.5부터 `BPF_F_MMAPABLE` flag를 설정해 `BPF_MAP_TYPE_ARRAY`의 memory mapping을 enable할 수 있습니다. Map definition은 page-aligned되어 첫 page에서 시작하고, 모든 array value를 담을 page-sized·page-aligned memory block은 두 번째 page부터 allocate됩니다. 경우에 따라 memory가 필요 이상으로 allocate될 수 있습니다.

Memory mapping을 사용하면 performance와 usability가 좋아집니다. Userspace program이 data에 접근하고 변경하기 위해 helper function을 호출하지 않아도 되기 때문입니다.

Kernel BPF lookup과 update helper

29-62

Kernel BPF에서 array element를 조회하는 helper prototype은 다음과 같습니다.

void *bpf_map_lookup_elem(struct bpf_map *map, const void *key)

`bpf_map_lookup_elem()`은 array element 내부를 가리키는 pointer를 반환합니다. Userspace가 value를 읽는 동안 in-place update와 data race가 발생하지 않도록 `__sync_fetch_and_add()` 같은 primitive를 사용해야 합니다.

Array element를 update하는 helper prototype은 다음과 같습니다.

long bpf_map_update_elem(struct bpf_map *map, const void *key, const void *value, u64 flags)

`bpf_map_update_elem()`은 성공하면 0을, 실패하면 negative error를 반환합니다. Array size가 고정이므로 `bpf_map_delete_elem()`은 지원하지 않습니다. Element를 clear하려면 해당 index에 zero value를 `bpf_map_update_elem()`으로 넣습니다.

Per-CPU array 접근

63-84

`BPF_MAP_TYPE_ARRAY`의 value는 서로 다른 CPU에서 실행되는 여러 program이 함께 접근할 수 있습니다. Storage를 한 CPU로 제한하려면 `BPF_MAP_TYPE_PERCPU_ARRAY`를 사용합니다.

`BPF_MAP_TYPE_PERCPU_ARRAY`에서 `bpf_map_update_elem()`과 `bpf_map_lookup_elem()`은 current CPU의 slot에 자동으로 접근합니다.

특정 CPU의 array value를 조회하는 helper prototype은 다음과 같습니다.

void *bpf_map_lookup_percpu_elem(struct bpf_map *map, const void *key, u32 cpu)

`bpf_map_lookup_percpu_elem()`은 지정한 CPU의 array value를 찾습니다. 성공하면 value를 반환하고 entry가 없거나 `cpu`가 invalid하면 `NULL`을 반환합니다.

Concurrency와 userspace API

85-96

Kernel 5.1부터 BPF infrastructure는 concurrent access를 synchronize하기 위한 `struct bpf_spin_lock`을 제공합니다.

Userspace에서는 위 helper와 같은 name의 libbpf API를 사용하며 map은 `fd`로 식별합니다.

Kernel BPF array 선언과 접근 예제

97-138

기능 예제는 `tools/testing/selftests/bpf` directory에 있습니다. 다음 code는 key `u32`, value `long`, `max_entries` 256인 `BPF_MAP_TYPE_ARRAY`를 `.maps` section에 선언합니다.

struct {
        __uint(type, BPF_MAP_TYPE_ARRAY);
        __type(key, u32);
        __type(value, long);
        __uint(max_entries, 256);
} my_map SEC(".maps");

다음 BPF program은 packet의 IP protocol을 index로 사용해 array element를 찾습니다. `bpf_skb_load_bytes()`로 IP header를 읽고 `bpf_map_lookup_elem()`으로 value pointer를 얻은 뒤 `__sync_fetch_and_add()`로 packet length를 atomic하게 더합니다.

int bpf_prog(struct __sk_buff *skb)
{
        struct iphdr ip;
        int index;
        long *value;

        if (bpf_skb_load_bytes(skb, ETH_HLEN, &ip, sizeof(ip)) < 0)
                return 0;

        index = ip.protocol;
        value = bpf_map_lookup_elem(&my_map, &index);
        if (value)
                __sync_fetch_and_add(value, skb->len);

        return 0;
}

Userspace array 생성·초기화·조회

139-206

Userspace의 첫 예제는 `bpf_map_create_opts`로 `BPF_F_MMAPABLE` flag를 설정하고 `bpf_map_create()`로 `BPF_MAP_TYPE_ARRAY`를 생성합니다. Key는 `__u32`, value는 `long`, 최대 entry 수는 256입니다.

#include <bpf/libbpf.h>
#include <bpf/bpf.h>

int create_array()
{
        int fd;
        LIBBPF_OPTS(bpf_map_create_opts, opts, .map_flags = BPF_F_MMAPABLE);

        fd = bpf_map_create(BPF_MAP_TYPE_ARRAY,
                            "example_array",       /* name */
                            sizeof(__u32),         /* key size */
                            sizeof(long),          /* value size */
                            256,                   /* max entries */
                            &opts);                /* create opts */
        return fd;
}

다음 예제는 0부터 255까지 순회하면서 각 index와 같은 `long` value를 `BPF_ANY` flag로 저장합니다. Update가 실패하면 negative return value를 즉시 반환합니다.

int initialize_array(int fd)
{
        __u32 i;
        long value;
        int ret;

        for (i = 0; i < 256; i++) {
                value = i;
                ret = bpf_map_update_elem(fd, &i, &value, BPF_ANY);
                if (ret < 0)
                        return ret;
        }

        return ret;
}

마지막 array 예제는 index 42를 `bpf_map_lookup_elem()`으로 조회해 userspace buffer `value`에 받고 결과가 42인지 확인합니다.

int lookup(int fd)
{
        __u32 index = 42;
        long value;
        int ret;

        ret = bpf_map_lookup_elem(fd, &index, &value);
        if (ret < 0)
                return ret;

        /* use value here */
        assert(value == 42);

        return ret;
}

Userspace per-CPU array 예제

207-254

Per-CPU array를 initialize할 때 `libbpf_num_possible_cpus()`로 가능한 CPU 수를 구하고 `long values[ncpus]` buffer를 준비합니다. 각 map index마다 모든 CPU slot을 같은 값으로 채운 뒤 `bpf_map_update_elem()`에 전체 buffer를 전달합니다.

int initialize_array(int fd)
{
        int ncpus = libbpf_num_possible_cpus();
        long values[ncpus];
        __u32 i, j;
        int ret;

        for (i = 0; i < 256 ; i++) {
                for (j = 0; j < ncpus; j++)
                        values[j] = i;
                ret = bpf_map_update_elem(fd, &i, &values, BPF_ANY);
                if (ret < 0)
                        return ret;
        }

        return ret;
}

Per-CPU value를 조회할 때도 `ncpus` element를 담을 buffer를 전달합니다. 조회 후 각 CPU slot을 순회하여 index 42에 저장한 값이 모두 42인지 확인합니다.

int lookup(int fd)
{
        int ncpus = libbpf_num_possible_cpus();
        __u32 index = 42, j;
        long values[ncpus];
        int ret;

        ret = bpf_map_lookup_elem(fd, &index, &values);
        if (ret < 0)
                return ret;

        for (j = 0; j < ncpus; j++) {
                /* Use per CPU value here */
                assert(values[j] == 42);
        }

        return ret;
}

Per-CPU userspace semantics

255-262

Userspace에서 `BPF_MAP_TYPE_PERCPU_ARRAY`에 접근하면 각 map value는 `ncpus` element를 가진 array로 표현됩니다.

이 map type에 `bpf_map_update_elem()`을 호출할 때는 `BPF_NOEXIST` flag를 사용할 수 없습니다. 모든 element가 생성 시점에 이미 allocate되어 존재하기 때문입니다.