← Documents Documentation/filesystems/resctrl.rst GitHub 원문 ↗

Linux 6.18.37 · Filesystems

User Interface for Resource Control feature (resctrl)

Resctrl의 CAT·MBA·MBM·group ABI, pseudo-locking과 errata를 다룬 전문 번역입니다.

Source pathDocumentation/filesystems/resctrl.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

resctrl.rst:1-1848

Resctrl은 Intel RDT와 AMD QoS의 cache allocation, memory bandwidth 제어, occupancy·bandwidth monitoring을 디렉터리·파일 ABI로 제공한다. CTRL_MON은 resource 할당과 집계를, MON은 부모 group 안 task subset의 관찰을 담당한다.

정확한 운용에는 CLOSID·RMID 한도, RMID limbo, task와 CPU 귀속 우선순위, CBM 연속성·독점 overlap, MBA 백분율과 실제 package bandwidth의 차이를 함께 고려해야 한다.

Pseudo-locking과 MBM event assignment는 일반 할당보다 강한 책임을 요구한다. 전자는 CPU affinity와 cache eviction 가능성을 application이 관리하고, 후자는 hardware counter 할당·filter 변경 뒤 `Unavailable` 상태를 userspace가 처리해야 한다.

Resctrl 관리 흐름
`info`에서 resource·CLOSID·RMID·counter capability 확인CTRL_MON·MON group 생성과 task·CPU 귀속`schemata`로 CAT·MBA·SMBA 설정`mon_data`와 MBM assignment로 관찰`last_cmd_status`, tracing, errata 보정으로 검증

기능 탐지에서 group·schemata·monitoring·검증까지 이어진다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2 .. include:: <isonum.txt>
3
4 =====================================================
5 User Interface for Resource Control feature (resctrl)
6 =====================================================
7
8 :Copyright: |copy| 2016 Intel Corporation
9 :Authors: - Fenghua Yu <fenghua.yu@intel.com>
10 - Tony Luck <tony.luck@intel.com>
11 - Vikas Shivappa <vikas.shivappa@intel.com>
12
13
14 Intel refers to this feature as Intel Resource Director Technology(Intel(R) RDT).
15 AMD refers to this feature as AMD Platform Quality of Service(AMD QoS).
16
17 This feature is enabled by the CONFIG_X86_CPU_RESCTRL and the x86 /proc/cpuinfo
18 flag bits:
19
20 =============================================== ================================
21 RDT (Resource Director Technology) Allocation "rdt_a"
22 CAT (Cache Allocation Technology) "cat_l3", "cat_l2"
23 CDP (Code and Data Prioritization) "cdp_l3", "cdp_l2"
24 CQM (Cache QoS Monitoring) "cqm_llc", "cqm_occup_llc"
25 MBM (Memory Bandwidth Monitoring) "cqm_mbm_total", "cqm_mbm_local"
26 MBA (Memory Bandwidth Allocation) "mba"
27 SMBA (Slow Memory Bandwidth Allocation) ""
28 BMEC (Bandwidth Monitoring Event Configuration) ""
29 ABMC (Assignable Bandwidth Monitoring Counters) ""
30 =============================================== ================================
31
32 Historically, new features were made visible by default in /proc/cpuinfo. This
33 resulted in the feature flags becoming hard to parse by humans. Adding a new
34 flag to /proc/cpuinfo should be avoided if user space can obtain information
35 about the feature from resctrl's info directory.
36
37 To use the feature mount the file system::
38
39 # mount -t resctrl resctrl [-o cdp[,cdpl2][,mba_MBps][,debug]] /sys/fs/resctrl
40
41 mount options are:
42
43 "cdp":
44 Enable code/data prioritization in L3 cache allocations.
45 "cdpl2":
46 Enable code/data prioritization in L2 cache allocations.
47 "mba_MBps":
48 Enable the MBA Software Controller(mba_sc) to specify MBA
49 bandwidth in MiBps
50 "debug":
51 Make debug files accessible. Available debug files are annotated with
52 "Available only with debug option".
53
54 L2 and L3 CDP are controlled separately.
55
56 RDT features are orthogonal. A particular system may support only
57 monitoring, only control, or both monitoring and control. Cache
58 pseudo-locking is a unique way of using cache control to "pin" or
59 "lock" data in the cache. Details can be found in
60 "Cache Pseudo-Locking".
61
62
63 The mount succeeds if either of allocation or monitoring is present, but
64 only those files and directories supported by the system will be created.
65 For more details on the behavior of the interface during monitoring
66 and allocation, see the "Resource alloc and monitor groups" section.
67
68 Info directory
69 ==============
70
71 The 'info' directory contains information about the enabled
72 resources. Each resource has its own subdirectory. The subdirectory
73 names reflect the resource names.
74
75 Each subdirectory contains the following files with respect to
76 allocation:
77
78 Cache resource(L3/L2) subdirectory contains the following files
79 related to allocation:
80
81 "num_closids":
82 The number of CLOSIDs which are valid for this
83 resource. The kernel uses the smallest number of
84 CLOSIDs of all enabled resources as limit.
85 "cbm_mask":
86 The bitmask which is valid for this resource.
87 This mask is equivalent to 100%.
88 "min_cbm_bits":
89 The minimum number of consecutive bits which
90 must be set when writing a mask.
91
92 "shareable_bits":
93 Bitmask of shareable resource with other executing
94 entities (e.g. I/O). User can use this when
95 setting up exclusive cache partitions. Note that
96 some platforms support devices that have their
97 own settings for cache use which can over-ride
98 these bits.
99 "bit_usage":
100 Annotated capacity bitmasks showing how all
101 instances of the resource are used. The legend is:
102
103 "0":
104 Corresponding region is unused. When the system's
105 resources have been allocated and a "0" is found
106 in "bit_usage" it is a sign that resources are
107 wasted.
108
109 "H":
110 Corresponding region is used by hardware only
111 but available for software use. If a resource
112 has bits set in "shareable_bits" but not all
113 of these bits appear in the resource groups'
114 schematas then the bits appearing in
115 "shareable_bits" but no resource group will
116 be marked as "H".
117 "X":
118 Corresponding region is available for sharing and
119 used by hardware and software. These are the
120 bits that appear in "shareable_bits" as
121 well as a resource group's allocation.
122 "S":
123 Corresponding region is used by software
124 and available for sharing.
125 "E":
126 Corresponding region is used exclusively by
127 one resource group. No sharing allowed.
128 "P":
129 Corresponding region is pseudo-locked. No
130 sharing allowed.
131 "sparse_masks":
132 Indicates if non-contiguous 1s value in CBM is supported.
133
134 "0":
135 Only contiguous 1s value in CBM is supported.
136 "1":
137 Non-contiguous 1s value in CBM is supported.
138
139 Memory bandwidth(MB) subdirectory contains the following files
140 with respect to allocation:
141
142 "min_bandwidth":
143 The minimum memory bandwidth percentage which
144 user can request.
145
146 "bandwidth_gran":
147 The granularity in which the memory bandwidth
148 percentage is allocated. The allocated
149 b/w percentage is rounded off to the next
150 control step available on the hardware. The
151 available bandwidth control steps are:
152 min_bandwidth + N * bandwidth_gran.
153
154 "delay_linear":
155 Indicates if the delay scale is linear or
156 non-linear. This field is purely informational
157 only.
158
159 "thread_throttle_mode":
160 Indicator on Intel systems of how tasks running on threads
161 of a physical core are throttled in cases where they
162 request different memory bandwidth percentages:
163
164 "max":
165 the smallest percentage is applied
166 to all threads
167 "per-thread":
168 bandwidth percentages are directly applied to
169 the threads running on the core
170
171 If RDT monitoring is available there will be an "L3_MON" directory
172 with the following files:
173
174 "num_rmids":
175 The number of RMIDs available. This is the
176 upper bound for how many "CTRL_MON" + "MON"
177 groups can be created.
178
179 "mon_features":
180 Lists the monitoring events if
181 monitoring is enabled for the resource.
182 Example::
183
184 # cat /sys/fs/resctrl/info/L3_MON/mon_features
185 llc_occupancy
186 mbm_total_bytes
187 mbm_local_bytes
188
189 If the system supports Bandwidth Monitoring Event
190 Configuration (BMEC), then the bandwidth events will
191 be configurable. The output will be::
192
193 # cat /sys/fs/resctrl/info/L3_MON/mon_features
194 llc_occupancy
195 mbm_total_bytes
196 mbm_total_bytes_config
197 mbm_local_bytes
198 mbm_local_bytes_config
199
200 "mbm_total_bytes_config", "mbm_local_bytes_config":
201 Read/write files containing the configuration for the mbm_total_bytes
202 and mbm_local_bytes events, respectively, when the Bandwidth
203 Monitoring Event Configuration (BMEC) feature is supported.
204 The event configuration settings are domain specific and affect
205 all the CPUs in the domain. When either event configuration is
206 changed, the bandwidth counters for all RMIDs of both events
207 (mbm_total_bytes as well as mbm_local_bytes) are cleared for that
208 domain. The next read for every RMID will report "Unavailable"
209 and subsequent reads will report the valid value.
210
211 Following are the types of events supported:
212
213 ==== ========================================================
214 Bits Description
215 ==== ========================================================
216 6 Dirty Victims from the QOS domain to all types of memory
217 5 Reads to slow memory in the non-local NUMA domain
218 4 Reads to slow memory in the local NUMA domain
219 3 Non-temporal writes to non-local NUMA domain
220 2 Non-temporal writes to local NUMA domain
221 1 Reads to memory in the non-local NUMA domain
222 0 Reads to memory in the local NUMA domain
223 ==== ========================================================
224
225 By default, the mbm_total_bytes configuration is set to 0x7f to count
226 all the event types and the mbm_local_bytes configuration is set to
227 0x15 to count all the local memory events.
228
229 Examples:
230
231 * To view the current configuration::
232 ::
233
234 # cat /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
235 0=0x7f;1=0x7f;2=0x7f;3=0x7f
236
237 # cat /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config
238 0=0x15;1=0x15;3=0x15;4=0x15
239
240 * To change the mbm_total_bytes to count only reads on domain 0,
241 the bits 0, 1, 4 and 5 needs to be set, which is 110011b in binary
242 (in hexadecimal 0x33):
243 ::
244
245 # echo "0=0x33" > /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
246
247 # cat /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
248 0=0x33;1=0x7f;2=0x7f;3=0x7f
249
250 * To change the mbm_local_bytes to count all the slow memory reads on
251 domain 0 and 1, the bits 4 and 5 needs to be set, which is 110000b
252 in binary (in hexadecimal 0x30):
253 ::
254
255 # echo "0=0x30;1=0x30" > /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config
256
257 # cat /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config
258 0=0x30;1=0x30;3=0x15;4=0x15
259
260 "mbm_assign_mode":
261 The supported counter assignment modes. The enclosed brackets indicate which mode
262 is enabled. The MBM events associated with counters may reset when "mbm_assign_mode"
263 is changed.
264 ::
265
266 # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
267 [mbm_event]
268 default
269
270 "mbm_event":
271
272 mbm_event mode allows users to assign a hardware counter to an RMID, event
273 pair and monitor the bandwidth usage as long as it is assigned. The hardware
274 continues to track the assigned counter until it is explicitly unassigned by
275 the user. Each event within a resctrl group can be assigned independently.
276
277 In this mode, a monitoring event can only accumulate data while it is backed
278 by a hardware counter. Use "mbm_L3_assignments" found in each CTRL_MON and MON
279 group to specify which of the events should have a counter assigned. The number
280 of counters available is described in the "num_mbm_cntrs" file. Changing the
281 mode may cause all counters on the resource to reset.
282
283 Moving to mbm_event counter assignment mode requires users to assign the counters
284 to the events. Otherwise, the MBM event counters will return 'Unassigned' when read.
285
286 The mode is beneficial for AMD platforms that support more CTRL_MON
287 and MON groups than available hardware counters. By default, this
288 feature is enabled on AMD platforms with the ABMC (Assignable Bandwidth
289 Monitoring Counters) capability, ensuring counters remain assigned even
290 when the corresponding RMID is not actively used by any processor.
291
292 "default":
293
294 In default mode, resctrl assumes there is a hardware counter for each
295 event within every CTRL_MON and MON group. On AMD platforms, it is
296 recommended to use the mbm_event mode, if supported, to prevent reset of MBM
297 events between reads resulting from hardware re-allocating counters. This can
298 result in misleading values or display "Unavailable" if no counter is assigned
299 to the event.
300
301 * To enable "mbm_event" counter assignment mode:
302 ::
303
304 # echo "mbm_event" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
305
306 * To enable "default" monitoring mode:
307 ::
308
309 # echo "default" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
310
311 "num_mbm_cntrs":
312 The maximum number of counters (total of available and assigned counters) in
313 each domain when the system supports mbm_event mode.
314
315 For example, on a system with maximum of 32 memory bandwidth monitoring
316 counters in each of its L3 domains:
317 ::
318
319 # cat /sys/fs/resctrl/info/L3_MON/num_mbm_cntrs
320 0=32;1=32
321
322 "available_mbm_cntrs":
323 The number of counters available for assignment in each domain when mbm_event
324 mode is enabled on the system.
325
326 For example, on a system with 30 available [hardware] assignable counters
327 in each of its L3 domains:
328 ::
329
330 # cat /sys/fs/resctrl/info/L3_MON/available_mbm_cntrs
331 0=30;1=30
332
333 "event_configs":
334 Directory that exists when "mbm_event" counter assignment mode is supported.
335 Contains a sub-directory for each MBM event that can be assigned to a counter.
336
337 Two MBM events are supported by default: mbm_local_bytes and mbm_total_bytes.
338 Each MBM event's sub-directory contains a file named "event_filter" that is
339 used to view and modify which memory transactions the MBM event is configured
340 with. The file is accessible only when "mbm_event" counter assignment mode is
341 enabled.
342
343 List of memory transaction types supported:
344
345 ========================== ========================================================
346 Name Description
347 ========================== ========================================================
348 dirty_victim_writes_all Dirty Victims from the QOS domain to all types of memory
349 remote_reads_slow_memory Reads to slow memory in the non-local NUMA domain
350 local_reads_slow_memory Reads to slow memory in the local NUMA domain
351 remote_non_temporal_writes Non-temporal writes to non-local NUMA domain
352 local_non_temporal_writes Non-temporal writes to local NUMA domain
353 remote_reads Reads to memory in the non-local NUMA domain
354 local_reads Reads to memory in the local NUMA domain
355 ========================== ========================================================
356
357 For example::
358
359 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
360 local_reads,remote_reads,local_non_temporal_writes,remote_non_temporal_writes,
361 local_reads_slow_memory,remote_reads_slow_memory,dirty_victim_writes_all
362
363 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
364 local_reads,local_non_temporal_writes,local_reads_slow_memory
365
366 Modify the event configuration by writing to the "event_filter" file within
367 the "event_configs" directory. The read/write "event_filter" file contains the
368 configuration of the event that reflects which memory transactions are counted by it.
369
370 For example::
371
372 # echo "local_reads, local_non_temporal_writes" >
373 /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
374
375 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
376 local_reads,local_non_temporal_writes
377
378 "mbm_assign_on_mkdir":
379 Exists when "mbm_event" counter assignment mode is supported. Accessible
380 only when "mbm_event" counter assignment mode is enabled.
381
382 Determines if a counter will automatically be assigned to an RMID, MBM event
383 pair when its associated monitor group is created via mkdir. Enabled by default
384 on boot, also when switched from "default" mode to "mbm_event" counter assignment
385 mode. Users can disable this capability by writing to the interface.
386
387 "0":
388 Auto assignment is disabled.
389 "1":
390 Auto assignment is enabled.
391
392 Example::
393
394 # echo 0 > /sys/fs/resctrl/info/L3_MON/mbm_assign_on_mkdir
395 # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_on_mkdir
396 0
397
398 "max_threshold_occupancy":
399 Read/write file provides the largest value (in
400 bytes) at which a previously used LLC_occupancy
401 counter can be considered for re-use.
402
403 Finally, in the top level of the "info" directory there is a file
404 named "last_cmd_status". This is reset with every "command" issued
405 via the file system (making new directories or writing to any of the
406 control files). If the command was successful, it will read as "ok".
407 If the command failed, it will provide more information that can be
408 conveyed in the error returns from file operations. E.g.
409 ::
410
411 # echo L3:0=f7 > schemata
412 bash: echo: write error: Invalid argument
413 # cat info/last_cmd_status
414 mask f7 has non-consecutive 1-bits
415
416 Resource alloc and monitor groups
417 =================================
418
419 Resource groups are represented as directories in the resctrl file
420 system. The default group is the root directory which, immediately
421 after mounting, owns all the tasks and cpus in the system and can make
422 full use of all resources.
423
424 On a system with RDT control features additional directories can be
425 created in the root directory that specify different amounts of each
426 resource (see "schemata" below). The root and these additional top level
427 directories are referred to as "CTRL_MON" groups below.
428
429 On a system with RDT monitoring the root directory and other top level
430 directories contain a directory named "mon_groups" in which additional
431 directories can be created to monitor subsets of tasks in the CTRL_MON
432 group that is their ancestor. These are called "MON" groups in the rest
433 of this document.
434
435 Removing a directory will move all tasks and cpus owned by the group it
436 represents to the parent. Removing one of the created CTRL_MON groups
437 will automatically remove all MON groups below it.
438
439 Moving MON group directories to a new parent CTRL_MON group is supported
440 for the purpose of changing the resource allocations of a MON group
441 without impacting its monitoring data or assigned tasks. This operation
442 is not allowed for MON groups which monitor CPUs. No other move
443 operation is currently allowed other than simply renaming a CTRL_MON or
444 MON group.
445
446 All groups contain the following files:
447
448 "tasks":
449 Reading this file shows the list of all tasks that belong to
450 this group. Writing a task id to the file will add a task to the
451 group. Multiple tasks can be added by separating the task ids
452 with commas. Tasks will be assigned sequentially. Multiple
453 failures are not supported. A single failure encountered while
454 attempting to assign a task will cause the operation to abort and
455 already added tasks before the failure will remain in the group.
456 Failures will be logged to /sys/fs/resctrl/info/last_cmd_status.
457
458 If the group is a CTRL_MON group the task is removed from
459 whichever previous CTRL_MON group owned the task and also from
460 any MON group that owned the task. If the group is a MON group,
461 then the task must already belong to the CTRL_MON parent of this
462 group. The task is removed from any previous MON group.
463
464
465 "cpus":
466 Reading this file shows a bitmask of the logical CPUs owned by
467 this group. Writing a mask to this file will add and remove
468 CPUs to/from this group. As with the tasks file a hierarchy is
469 maintained where MON groups may only include CPUs owned by the
470 parent CTRL_MON group.
471 When the resource group is in pseudo-locked mode this file will
472 only be readable, reflecting the CPUs associated with the
473 pseudo-locked region.
474
475
476 "cpus_list":
477 Just like "cpus", only using ranges of CPUs instead of bitmasks.
478
479
480 When control is enabled all CTRL_MON groups will also contain:
481
482 "schemata":
483 A list of all the resources available to this group.
484 Each resource has its own line and format - see below for details.
485
486 "size":
487 Mirrors the display of the "schemata" file to display the size in
488 bytes of each allocation instead of the bits representing the
489 allocation.
490
491 "mode":
492 The "mode" of the resource group dictates the sharing of its
493 allocations. A "shareable" resource group allows sharing of its
494 allocations while an "exclusive" resource group does not. A
495 cache pseudo-locked region is created by first writing
496 "pseudo-locksetup" to the "mode" file before writing the cache
497 pseudo-locked region's schemata to the resource group's "schemata"
498 file. On successful pseudo-locked region creation the mode will
499 automatically change to "pseudo-locked".
500
501 "ctrl_hw_id":
502 Available only with debug option. The identifier used by hardware
503 for the control group. On x86 this is the CLOSID.
504
505 When monitoring is enabled all MON groups will also contain:
506
507 "mon_data":
508 This contains a set of files organized by L3 domain and by
509 RDT event. E.g. on a system with two L3 domains there will
510 be subdirectories "mon_L3_00" and "mon_L3_01". Each of these
511 directories have one file per event (e.g. "llc_occupancy",
512 "mbm_total_bytes", and "mbm_local_bytes"). In a MON group these
513 files provide a read out of the current value of the event for
514 all tasks in the group. In CTRL_MON groups these files provide
515 the sum for all tasks in the CTRL_MON group and all tasks in
516 MON groups. Please see example section for more details on usage.
517 On systems with Sub-NUMA Cluster (SNC) enabled there are extra
518 directories for each node (located within the "mon_L3_XX" directory
519 for the L3 cache they occupy). These are named "mon_sub_L3_YY"
520 where "YY" is the node number.
521
522 When the 'mbm_event' counter assignment mode is enabled, reading
523 an MBM event of a MON group returns 'Unassigned' if no hardware
524 counter is assigned to it. For CTRL_MON groups, 'Unassigned' is
525 returned if the MBM event does not have an assigned counter in the
526 CTRL_MON group nor in any of its associated MON groups.
527
528 "mon_hw_id":
529 Available only with debug option. The identifier used by hardware
530 for the monitor group. On x86 this is the RMID.
531
532 When monitoring is enabled all MON groups may also contain:
533
534 "mbm_L3_assignments":
535 Exists when "mbm_event" counter assignment mode is supported and lists the
536 counter assignment states of the group.
537
538 The assignment list is displayed in the following format:
539
540 <Event>:<Domain ID>=<Assignment state>;<Domain ID>=<Assignment state>
541
542 Event: A valid MBM event in the
543 /sys/fs/resctrl/info/L3_MON/event_configs directory.
544
545 Domain ID: A valid domain ID. When writing, '*' applies the changes
546 to all the domains.
547
548 Assignment states:
549
550 _ : No counter assigned.
551
552 e : Counter assigned exclusively.
553
554 Example:
555
556 To display the counter assignment states for the default group.
557 ::
558
559 # cd /sys/fs/resctrl
560 # cat /sys/fs/resctrl/mbm_L3_assignments
561 mbm_total_bytes:0=e;1=e
562 mbm_local_bytes:0=e;1=e
563
564 Assignments can be modified by writing to the interface.
565
566 Examples:
567
568 To unassign the counter associated with the mbm_total_bytes event on domain 0:
569 ::
570
571 # echo "mbm_total_bytes:0=_" > /sys/fs/resctrl/mbm_L3_assignments
572 # cat /sys/fs/resctrl/mbm_L3_assignments
573 mbm_total_bytes:0=_;1=e
574 mbm_local_bytes:0=e;1=e
575
576 To unassign the counter associated with the mbm_total_bytes event on all the domains:
577 ::
578
579 # echo "mbm_total_bytes:*=_" > /sys/fs/resctrl/mbm_L3_assignments
580 # cat /sys/fs/resctrl/mbm_L3_assignments
581 mbm_total_bytes:0=_;1=_
582 mbm_local_bytes:0=e;1=e
583
584 To assign a counter associated with the mbm_total_bytes event on all domains in
585 exclusive mode:
586 ::
587
588 # echo "mbm_total_bytes:*=e" > /sys/fs/resctrl/mbm_L3_assignments
589 # cat /sys/fs/resctrl/mbm_L3_assignments
590 mbm_total_bytes:0=e;1=e
591 mbm_local_bytes:0=e;1=e
592
593 When the "mba_MBps" mount option is used all CTRL_MON groups will also contain:
594
595 "mba_MBps_event":
596 Reading this file shows which memory bandwidth event is used
597 as input to the software feedback loop that keeps memory bandwidth
598 below the value specified in the schemata file. Writing the
599 name of one of the supported memory bandwidth events found in
600 /sys/fs/resctrl/info/L3_MON/mon_features changes the input
601 event.
602
603 Resource allocation rules
604 -------------------------
605
606 When a task is running the following rules define which resources are
607 available to it:
608
609 1) If the task is a member of a non-default group, then the schemata
610 for that group is used.
611
612 2) Else if the task belongs to the default group, but is running on a
613 CPU that is assigned to some specific group, then the schemata for the
614 CPU's group is used.
615
616 3) Otherwise the schemata for the default group is used.
617
618 Resource monitoring rules
619 -------------------------
620 1) If a task is a member of a MON group, or non-default CTRL_MON group
621 then RDT events for the task will be reported in that group.
622
623 2) If a task is a member of the default CTRL_MON group, but is running
624 on a CPU that is assigned to some specific group, then the RDT events
625 for the task will be reported in that group.
626
627 3) Otherwise RDT events for the task will be reported in the root level
628 "mon_data" group.
629
630
631 Notes on cache occupancy monitoring and control
632 ===============================================
633 When moving a task from one group to another you should remember that
634 this only affects *new* cache allocations by the task. E.g. you may have
635 a task in a monitor group showing 3 MB of cache occupancy. If you move
636 to a new group and immediately check the occupancy of the old and new
637 groups you will likely see that the old group is still showing 3 MB and
638 the new group zero. When the task accesses locations still in cache from
639 before the move, the h/w does not update any counters. On a busy system
640 you will likely see the occupancy in the old group go down as cache lines
641 are evicted and re-used while the occupancy in the new group rises as
642 the task accesses memory and loads into the cache are counted based on
643 membership in the new group.
644
645 The same applies to cache allocation control. Moving a task to a group
646 with a smaller cache partition will not evict any cache lines. The
647 process may continue to use them from the old partition.
648
649 Hardware uses CLOSid(Class of service ID) and an RMID(Resource monitoring ID)
650 to identify a control group and a monitoring group respectively. Each of
651 the resource groups are mapped to these IDs based on the kind of group. The
652 number of CLOSid and RMID are limited by the hardware and hence the creation of
653 a "CTRL_MON" directory may fail if we run out of either CLOSID or RMID
654 and creation of "MON" group may fail if we run out of RMIDs.
655
656 max_threshold_occupancy - generic concepts
657 ------------------------------------------
658
659 Note that an RMID once freed may not be immediately available for use as
660 the RMID is still tagged the cache lines of the previous user of RMID.
661 Hence such RMIDs are placed on limbo list and checked back if the cache
662 occupancy has gone down. If there is a time when system has a lot of
663 limbo RMIDs but which are not ready to be used, user may see an -EBUSY
664 during mkdir.
665
666 max_threshold_occupancy is a user configurable value to determine the
667 occupancy at which an RMID can be freed.
668
669 The mon_llc_occupancy_limbo tracepoint gives the precise occupancy in bytes
670 for a subset of RMID that are not immediately available for allocation.
671 This can't be relied on to produce output every second, it may be necessary
672 to attempt to create an empty monitor group to force an update. Output may
673 only be produced if creation of a control or monitor group fails.
674
675 Schemata files - general concepts
676 ---------------------------------
677 Each line in the file describes one resource. The line starts with
678 the name of the resource, followed by specific values to be applied
679 in each of the instances of that resource on the system.
680
681 Cache IDs
682 ---------
683 On current generation systems there is one L3 cache per socket and L2
684 caches are generally just shared by the hyperthreads on a core, but this
685 isn't an architectural requirement. We could have multiple separate L3
686 caches on a socket, multiple cores could share an L2 cache. So instead
687 of using "socket" or "core" to define the set of logical cpus sharing
688 a resource we use a "Cache ID". At a given cache level this will be a
689 unique number across the whole system (but it isn't guaranteed to be a
690 contiguous sequence, there may be gaps). To find the ID for each logical
691 CPU look in /sys/devices/system/cpu/cpu*/cache/index*/id
692
693 Cache Bit Masks (CBM)
694 ---------------------
695 For cache resources we describe the portion of the cache that is available
696 for allocation using a bitmask. The maximum value of the mask is defined
697 by each cpu model (and may be different for different cache levels). It
698 is found using CPUID, but is also provided in the "info" directory of
699 the resctrl file system in "info/{resource}/cbm_mask". Some Intel hardware
700 requires that these masks have all the '1' bits in a contiguous block. So
701 0x3, 0x6 and 0xC are legal 4-bit masks with two bits set, but 0x5, 0x9
702 and 0xA are not. Check /sys/fs/resctrl/info/{resource}/sparse_masks
703 if non-contiguous 1s value is supported. On a system with a 20-bit mask
704 each bit represents 5% of the capacity of the cache. You could partition
705 the cache into four equal parts with masks: 0x1f, 0x3e0, 0x7c00, 0xf8000.
706
707 Notes on Sub-NUMA Cluster mode
708 ==============================
709 When SNC mode is enabled, Linux may load balance tasks between Sub-NUMA
710 nodes much more readily than between regular NUMA nodes since the CPUs
711 on Sub-NUMA nodes share the same L3 cache and the system may report
712 the NUMA distance between Sub-NUMA nodes with a lower value than used
713 for regular NUMA nodes.
714
715 The top-level monitoring files in each "mon_L3_XX" directory provide
716 the sum of data across all SNC nodes sharing an L3 cache instance.
717 Users who bind tasks to the CPUs of a specific Sub-NUMA node can read
718 the "llc_occupancy", "mbm_total_bytes", and "mbm_local_bytes" in the
719 "mon_sub_L3_YY" directories to get node local data.
720
721 Memory bandwidth allocation is still performed at the L3 cache
722 level. I.e. throttling controls are applied to all SNC nodes.
723
724 L3 cache allocation bitmaps also apply to all SNC nodes. But note that
725 the amount of L3 cache represented by each bit is divided by the number
726 of SNC nodes per L3 cache. E.g. with a 100MB cache on a system with 10-bit
727 allocation masks each bit normally represents 10MB. With SNC mode enabled
728 with two SNC nodes per L3 cache, each bit only represents 5MB.
729
730 Memory bandwidth Allocation and monitoring
731 ==========================================
732
733 For Memory bandwidth resource, by default the user controls the resource
734 by indicating the percentage of total memory bandwidth.
735
736 The minimum bandwidth percentage value for each cpu model is predefined
737 and can be looked up through "info/MB/min_bandwidth". The bandwidth
738 granularity that is allocated is also dependent on the cpu model and can
739 be looked up at "info/MB/bandwidth_gran". The available bandwidth
740 control steps are: min_bw + N * bw_gran. Intermediate values are rounded
741 to the next control step available on the hardware.
742
743 The bandwidth throttling is a core specific mechanism on some of Intel
744 SKUs. Using a high bandwidth and a low bandwidth setting on two threads
745 sharing a core may result in both threads being throttled to use the
746 low bandwidth (see "thread_throttle_mode").
747
748 The fact that Memory bandwidth allocation(MBA) may be a core
749 specific mechanism where as memory bandwidth monitoring(MBM) is done at
750 the package level may lead to confusion when users try to apply control
751 via the MBA and then monitor the bandwidth to see if the controls are
752 effective. Below are such scenarios:
753
754 1. User may *not* see increase in actual bandwidth when percentage
755 values are increased:
756
757 This can occur when aggregate L2 external bandwidth is more than L3
758 external bandwidth. Consider an SKL SKU with 24 cores on a package and
759 where L2 external is 10GBps (hence aggregate L2 external bandwidth is
760 240GBps) and L3 external bandwidth is 100GBps. Now a workload with '20
761 threads, having 50% bandwidth, each consuming 5GBps' consumes the max L3
762 bandwidth of 100GBps although the percentage value specified is only 50%
763 << 100%. Hence increasing the bandwidth percentage will not yield any
764 more bandwidth. This is because although the L2 external bandwidth still
765 has capacity, the L3 external bandwidth is fully used. Also note that
766 this would be dependent on number of cores the benchmark is run on.
767
768 2. Same bandwidth percentage may mean different actual bandwidth
769 depending on # of threads:
770
771 For the same SKU in #1, a 'single thread, with 10% bandwidth' and '4
772 thread, with 10% bandwidth' can consume up to 10GBps and 40GBps although
773 they have same percentage bandwidth of 10%. This is simply because as
774 threads start using more cores in an rdtgroup, the actual bandwidth may
775 increase or vary although user specified bandwidth percentage is same.
776
777 In order to mitigate this and make the interface more user friendly,
778 resctrl added support for specifying the bandwidth in MiBps as well. The
779 kernel underneath would use a software feedback mechanism or a "Software
780 Controller(mba_sc)" which reads the actual bandwidth using MBM counters
781 and adjust the memory bandwidth percentages to ensure::
782
783 "actual bandwidth < user specified bandwidth".
784
785 By default, the schemata would take the bandwidth percentage values
786 where as user can switch to the "MBA software controller" mode using
787 a mount option 'mba_MBps'. The schemata format is specified in the below
788 sections.
789
790 L3 schemata file details (code and data prioritization disabled)
791 ----------------------------------------------------------------
792 With CDP disabled the L3 schemata format is::
793
794 L3:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
795
796 L3 schemata file details (CDP enabled via mount option to resctrl)
797 ------------------------------------------------------------------
798 When CDP is enabled L3 control is split into two separate resources
799 so you can specify independent masks for code and data like this::
800
801 L3DATA:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
802 L3CODE:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
803
804 L2 schemata file details
805 ------------------------
806 CDP is supported at L2 using the 'cdpl2' mount option. The schemata
807 format is either::
808
809 L2:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
810
811 or
812
813 L2DATA:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
814 L2CODE:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
815
816
817 Memory bandwidth Allocation (default mode)
818 ------------------------------------------
819
820 Memory b/w domain is L3 cache.
821 ::
822
823 MB:<cache_id0>=bandwidth0;<cache_id1>=bandwidth1;...
824
825 Memory bandwidth Allocation specified in MiBps
826 ----------------------------------------------
827
828 Memory bandwidth domain is L3 cache.
829 ::
830
831 MB:<cache_id0>=bw_MiBps0;<cache_id1>=bw_MiBps1;...
832
833 Slow Memory Bandwidth Allocation (SMBA)
834 ---------------------------------------
835 AMD hardware supports Slow Memory Bandwidth Allocation (SMBA).
836 CXL.memory is the only supported "slow" memory device. With the
837 support of SMBA, the hardware enables bandwidth allocation on
838 the slow memory devices. If there are multiple such devices in
839 the system, the throttling logic groups all the slow sources
840 together and applies the limit on them as a whole.
841
842 The presence of SMBA (with CXL.memory) is independent of slow memory
843 devices presence. If there are no such devices on the system, then
844 configuring SMBA will have no impact on the performance of the system.
845
846 The bandwidth domain for slow memory is L3 cache. Its schemata file
847 is formatted as:
848 ::
849
850 SMBA:<cache_id0>=bandwidth0;<cache_id1>=bandwidth1;...
851
852 Reading/writing the schemata file
853 ---------------------------------
854 Reading the schemata file will show the state of all resources
855 on all domains. When writing you only need to specify those values
856 which you wish to change. E.g.
857 ::
858
859 # cat schemata
860 L3DATA:0=fffff;1=fffff;2=fffff;3=fffff
861 L3CODE:0=fffff;1=fffff;2=fffff;3=fffff
862 # echo "L3DATA:2=3c0;" > schemata
863 # cat schemata
864 L3DATA:0=fffff;1=fffff;2=3c0;3=fffff
865 L3CODE:0=fffff;1=fffff;2=fffff;3=fffff
866
867 Reading/writing the schemata file (on AMD systems)
868 --------------------------------------------------
869 Reading the schemata file will show the current bandwidth limit on all
870 domains. The allocated resources are in multiples of one eighth GB/s.
871 When writing to the file, you need to specify what cache id you wish to
872 configure the bandwidth limit.
873
874 For example, to allocate 2GB/s limit on the first cache id:
875
876 ::
877
878 # cat schemata
879 MB:0=2048;1=2048;2=2048;3=2048
880 L3:0=ffff;1=ffff;2=ffff;3=ffff
881
882 # echo "MB:1=16" > schemata
883 # cat schemata
884 MB:0=2048;1= 16;2=2048;3=2048
885 L3:0=ffff;1=ffff;2=ffff;3=ffff
886
887 Reading/writing the schemata file (on AMD systems) with SMBA feature
888 --------------------------------------------------------------------
889 Reading and writing the schemata file is the same as without SMBA in
890 above section.
891
892 For example, to allocate 8GB/s limit on the first cache id:
893
894 ::
895
896 # cat schemata
897 SMBA:0=2048;1=2048;2=2048;3=2048
898 MB:0=2048;1=2048;2=2048;3=2048
899 L3:0=ffff;1=ffff;2=ffff;3=ffff
900
901 # echo "SMBA:1=64" > schemata
902 # cat schemata
903 SMBA:0=2048;1= 64;2=2048;3=2048
904 MB:0=2048;1=2048;2=2048;3=2048
905 L3:0=ffff;1=ffff;2=ffff;3=ffff
906
907 Cache Pseudo-Locking
908 ====================
909 CAT enables a user to specify the amount of cache space that an
910 application can fill. Cache pseudo-locking builds on the fact that a
911 CPU can still read and write data pre-allocated outside its current
912 allocated area on a cache hit. With cache pseudo-locking, data can be
913 preloaded into a reserved portion of cache that no application can
914 fill, and from that point on will only serve cache hits. The cache
915 pseudo-locked memory is made accessible to user space where an
916 application can map it into its virtual address space and thus have
917 a region of memory with reduced average read latency.
918
919 The creation of a cache pseudo-locked region is triggered by a request
920 from the user to do so that is accompanied by a schemata of the region
921 to be pseudo-locked. The cache pseudo-locked region is created as follows:
922
923 - Create a CAT allocation CLOSNEW with a CBM matching the schemata
924 from the user of the cache region that will contain the pseudo-locked
925 memory. This region must not overlap with any current CAT allocation/CLOS
926 on the system and no future overlap with this cache region is allowed
927 while the pseudo-locked region exists.
928 - Create a contiguous region of memory of the same size as the cache
929 region.
930 - Flush the cache, disable hardware prefetchers, disable preemption.
931 - Make CLOSNEW the active CLOS and touch the allocated memory to load
932 it into the cache.
933 - Set the previous CLOS as active.
934 - At this point the closid CLOSNEW can be released - the cache
935 pseudo-locked region is protected as long as its CBM does not appear in
936 any CAT allocation. Even though the cache pseudo-locked region will from
937 this point on not appear in any CBM of any CLOS an application running with
938 any CLOS will be able to access the memory in the pseudo-locked region since
939 the region continues to serve cache hits.
940 - The contiguous region of memory loaded into the cache is exposed to
941 user-space as a character device.
942
943 Cache pseudo-locking increases the probability that data will remain
944 in the cache via carefully configuring the CAT feature and controlling
945 application behavior. There is no guarantee that data is placed in
946 cache. Instructions like INVD, WBINVD, CLFLUSH, etc. can still evict
947 “locked” data from cache. Power management C-states may shrink or
948 power off cache. Deeper C-states will automatically be restricted on
949 pseudo-locked region creation.
950
951 It is required that an application using a pseudo-locked region runs
952 with affinity to the cores (or a subset of the cores) associated
953 with the cache on which the pseudo-locked region resides. A sanity check
954 within the code will not allow an application to map pseudo-locked memory
955 unless it runs with affinity to cores associated with the cache on which the
956 pseudo-locked region resides. The sanity check is only done during the
957 initial mmap() handling, there is no enforcement afterwards and the
958 application self needs to ensure it remains affine to the correct cores.
959
960 Pseudo-locking is accomplished in two stages:
961
962 1) During the first stage the system administrator allocates a portion
963 of cache that should be dedicated to pseudo-locking. At this time an
964 equivalent portion of memory is allocated, loaded into allocated
965 cache portion, and exposed as a character device.
966 2) During the second stage a user-space application maps (mmap()) the
967 pseudo-locked memory into its address space.
968
969 Cache Pseudo-Locking Interface
970 ------------------------------
971 A pseudo-locked region is created using the resctrl interface as follows:
972
973 1) Create a new resource group by creating a new directory in /sys/fs/resctrl.
974 2) Change the new resource group's mode to "pseudo-locksetup" by writing
975 "pseudo-locksetup" to the "mode" file.
976 3) Write the schemata of the pseudo-locked region to the "schemata" file. All
977 bits within the schemata should be "unused" according to the "bit_usage"
978 file.
979
980 On successful pseudo-locked region creation the "mode" file will contain
981 "pseudo-locked" and a new character device with the same name as the resource
982 group will exist in /dev/pseudo_lock. This character device can be mmap()'ed
983 by user space in order to obtain access to the pseudo-locked memory region.
984
985 An example of cache pseudo-locked region creation and usage can be found below.
986
987 Cache Pseudo-Locking Debugging Interface
988 ----------------------------------------
989 The pseudo-locking debugging interface is enabled by default (if
990 CONFIG_DEBUG_FS is enabled) and can be found in /sys/kernel/debug/resctrl.
991
992 There is no explicit way for the kernel to test if a provided memory
993 location is present in the cache. The pseudo-locking debugging interface uses
994 the tracing infrastructure to provide two ways to measure cache residency of
995 the pseudo-locked region:
996
997 1) Memory access latency using the pseudo_lock_mem_latency tracepoint. Data
998 from these measurements are best visualized using a hist trigger (see
999 example below). In this test the pseudo-locked region is traversed at
1000 a stride of 32 bytes while hardware prefetchers and preemption
1001 are disabled. This also provides a substitute visualization of cache
1002 hits and misses.
1003 2) Cache hit and miss measurements using model specific precision counters if
1004 available. Depending on the levels of cache on the system the pseudo_lock_l2
1005 and pseudo_lock_l3 tracepoints are available.
1007 When a pseudo-locked region is created a new debugfs directory is created for
1008 it in debugfs as /sys/kernel/debug/resctrl/<newdir>. A single
1009 write-only file, pseudo_lock_measure, is present in this directory. The
1010 measurement of the pseudo-locked region depends on the number written to this
1011 debugfs file:
1013 1:
1014 writing "1" to the pseudo_lock_measure file will trigger the latency
1015 measurement captured in the pseudo_lock_mem_latency tracepoint. See
1016 example below.
1017 2:
1018 writing "2" to the pseudo_lock_measure file will trigger the L2 cache
1019 residency (cache hits and misses) measurement captured in the
1020 pseudo_lock_l2 tracepoint. See example below.
1021 3:
1022 writing "3" to the pseudo_lock_measure file will trigger the L3 cache
1023 residency (cache hits and misses) measurement captured in the
1024 pseudo_lock_l3 tracepoint.
1026 All measurements are recorded with the tracing infrastructure. This requires
1027 the relevant tracepoints to be enabled before the measurement is triggered.
1029 Example of latency debugging interface
1030 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1031 In this example a pseudo-locked region named "newlock" was created. Here is
1032 how we can measure the latency in cycles of reading from this region and
1033 visualize this data with a histogram that is available if CONFIG_HIST_TRIGGERS
1034 is set::
1036 # :> /sys/kernel/tracing/trace
1037 # echo 'hist:keys=latency' > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/trigger
1038 # echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
1039 # echo 1 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
1040 # echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
1041 # cat /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/hist
1043 # event histogram
1044 #
1045 # trigger info: hist:keys=latency:vals=hitcount:sort=hitcount:size=2048 [active]
1046 #
1048 { latency: 456 } hitcount: 1
1049 { latency: 50 } hitcount: 83
1050 { latency: 36 } hitcount: 96
1051 { latency: 44 } hitcount: 174
1052 { latency: 48 } hitcount: 195
1053 { latency: 46 } hitcount: 262
1054 { latency: 42 } hitcount: 693
1055 { latency: 40 } hitcount: 3204
1056 { latency: 38 } hitcount: 3484
1058 Totals:
1059 Hits: 8192
1060 Entries: 9
1061 Dropped: 0
1063 Example of cache hits/misses debugging
1064 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1065 In this example a pseudo-locked region named "newlock" was created on the L2
1066 cache of a platform. Here is how we can obtain details of the cache hits
1067 and misses using the platform's precision counters.
1068 ::
1070 # :> /sys/kernel/tracing/trace
1071 # echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
1072 # echo 2 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
1073 # echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
1074 # cat /sys/kernel/tracing/trace
1076 # tracer: nop
1077 #
1078 # _-----=> irqs-off
1079 # / _----=> need-resched
1080 # | / _---=> hardirq/softirq
1081 # || / _--=> preempt-depth
1082 # ||| / delay
1083 # TASK-PID CPU# |||| TIMESTAMP FUNCTION
1084 # | | | |||| | |
1085 pseudo_lock_mea-1672 [002] .... 3132.860500: pseudo_lock_l2: hits=4097 miss=0
1088 Examples for RDT allocation usage
1089 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1091 1) Example 1
1093 On a two socket machine (one L3 cache per socket) with just four bits
1094 for cache bit masks, minimum b/w of 10% with a memory bandwidth
1095 granularity of 10%.
1096 ::
1098 # mount -t resctrl resctrl /sys/fs/resctrl
1099 # cd /sys/fs/resctrl
1100 # mkdir p0 p1
1101 # echo "L3:0=3;1=c\nMB:0=50;1=50" > /sys/fs/resctrl/p0/schemata
1102 # echo "L3:0=3;1=3\nMB:0=50;1=50" > /sys/fs/resctrl/p1/schemata
1104 The default resource group is unmodified, so we have access to all parts
1105 of all caches (its schemata file reads "L3:0=f;1=f").
1107 Tasks that are under the control of group "p0" may only allocate from the
1108 "lower" 50% on cache ID 0, and the "upper" 50% of cache ID 1.
1109 Tasks in group "p1" use the "lower" 50% of cache on both sockets.
1111 Similarly, tasks that are under the control of group "p0" may use a
1112 maximum memory b/w of 50% on socket0 and 50% on socket 1.
1113 Tasks in group "p1" may also use 50% memory b/w on both sockets.
1114 Note that unlike cache masks, memory b/w cannot specify whether these
1115 allocations can overlap or not. The allocations specifies the maximum
1116 b/w that the group may be able to use and the system admin can configure
1117 the b/w accordingly.
1119 If resctrl is using the software controller (mba_sc) then user can enter the
1120 max b/w in MB rather than the percentage values.
1121 ::
1123 # echo "L3:0=3;1=c\nMB:0=1024;1=500" > /sys/fs/resctrl/p0/schemata
1124 # echo "L3:0=3;1=3\nMB:0=1024;1=500" > /sys/fs/resctrl/p1/schemata
1126 In the above example the tasks in "p1" and "p0" on socket 0 would use a max b/w
1127 of 1024MB where as on socket 1 they would use 500MB.
1129 2) Example 2
1131 Again two sockets, but this time with a more realistic 20-bit mask.
1133 Two real time tasks pid=1234 running on processor 0 and pid=5678 running on
1134 processor 1 on socket 0 on a 2-socket and dual core machine. To avoid noisy
1135 neighbors, each of the two real-time tasks exclusively occupies one quarter
1136 of L3 cache on socket 0.
1137 ::
1139 # mount -t resctrl resctrl /sys/fs/resctrl
1140 # cd /sys/fs/resctrl
1142 First we reset the schemata for the default group so that the "upper"
1143 50% of the L3 cache on socket 0 and 50% of memory b/w cannot be used by
1144 ordinary tasks::
1146 # echo "L3:0=3ff;1=fffff\nMB:0=50;1=100" > schemata
1148 Next we make a resource group for our first real time task and give
1149 it access to the "top" 25% of the cache on socket 0.
1150 ::
1152 # mkdir p0
1153 # echo "L3:0=f8000;1=fffff" > p0/schemata
1155 Finally we move our first real time task into this resource group. We
1156 also use taskset(1) to ensure the task always runs on a dedicated CPU
1157 on socket 0. Most uses of resource groups will also constrain which
1158 processors tasks run on.
1159 ::
1161 # echo 1234 > p0/tasks
1162 # taskset -cp 1 1234
1164 Ditto for the second real time task (with the remaining 25% of cache)::
1166 # mkdir p1
1167 # echo "L3:0=7c00;1=fffff" > p1/schemata
1168 # echo 5678 > p1/tasks
1169 # taskset -cp 2 5678
1171 For the same 2 socket system with memory b/w resource and CAT L3 the
1172 schemata would look like(Assume min_bandwidth 10 and bandwidth_gran is
1173 10):
1175 For our first real time task this would request 20% memory b/w on socket 0.
1176 ::
1178 # echo -e "L3:0=f8000;1=fffff\nMB:0=20;1=100" > p0/schemata
1180 For our second real time task this would request an other 20% memory b/w
1181 on socket 0.
1182 ::
1184 # echo -e "L3:0=f8000;1=fffff\nMB:0=20;1=100" > p0/schemata
1186 3) Example 3
1188 A single socket system which has real-time tasks running on core 4-7 and
1189 non real-time workload assigned to core 0-3. The real-time tasks share text
1190 and data, so a per task association is not required and due to interaction
1191 with the kernel it's desired that the kernel on these cores shares L3 with
1192 the tasks.
1193 ::
1195 # mount -t resctrl resctrl /sys/fs/resctrl
1196 # cd /sys/fs/resctrl
1198 First we reset the schemata for the default group so that the "upper"
1199 50% of the L3 cache on socket 0, and 50% of memory bandwidth on socket 0
1200 cannot be used by ordinary tasks::
1202 # echo "L3:0=3ff\nMB:0=50" > schemata
1204 Next we make a resource group for our real time cores and give it access
1205 to the "top" 50% of the cache on socket 0 and 50% of memory bandwidth on
1206 socket 0.
1207 ::
1209 # mkdir p0
1210 # echo "L3:0=ffc00\nMB:0=50" > p0/schemata
1212 Finally we move core 4-7 over to the new group and make sure that the
1213 kernel and the tasks running there get 50% of the cache. They should
1214 also get 50% of memory bandwidth assuming that the cores 4-7 are SMT
1215 siblings and only the real time threads are scheduled on the cores 4-7.
1216 ::
1218 # echo F0 > p0/cpus
1220 4) Example 4
1222 The resource groups in previous examples were all in the default "shareable"
1223 mode allowing sharing of their cache allocations. If one resource group
1224 configures a cache allocation then nothing prevents another resource group
1225 to overlap with that allocation.
1227 In this example a new exclusive resource group will be created on a L2 CAT
1228 system with two L2 cache instances that can be configured with an 8-bit
1229 capacity bitmask. The new exclusive resource group will be configured to use
1230 25% of each cache instance.
1231 ::
1233 # mount -t resctrl resctrl /sys/fs/resctrl/
1234 # cd /sys/fs/resctrl
1236 First, we observe that the default group is configured to allocate to all L2
1237 cache::
1239 # cat schemata
1240 L2:0=ff;1=ff
1242 We could attempt to create the new resource group at this point, but it will
1243 fail because of the overlap with the schemata of the default group::
1245 # mkdir p0
1246 # echo 'L2:0=0x3;1=0x3' > p0/schemata
1247 # cat p0/mode
1248 shareable
1249 # echo exclusive > p0/mode
1250 -sh: echo: write error: Invalid argument
1251 # cat info/last_cmd_status
1252 schemata overlaps
1254 To ensure that there is no overlap with another resource group the default
1255 resource group's schemata has to change, making it possible for the new
1256 resource group to become exclusive.
1257 ::
1259 # echo 'L2:0=0xfc;1=0xfc' > schemata
1260 # echo exclusive > p0/mode
1261 # grep . p0/*
1262 p0/cpus:0
1263 p0/mode:exclusive
1264 p0/schemata:L2:0=03;1=03
1265 p0/size:L2:0=262144;1=262144
1267 A new resource group will on creation not overlap with an exclusive resource
1268 group::
1270 # mkdir p1
1271 # grep . p1/*
1272 p1/cpus:0
1273 p1/mode:shareable
1274 p1/schemata:L2:0=fc;1=fc
1275 p1/size:L2:0=786432;1=786432
1277 The bit_usage will reflect how the cache is used::
1279 # cat info/L2/bit_usage
1280 0=SSSSSSEE;1=SSSSSSEE
1282 A resource group cannot be forced to overlap with an exclusive resource group::
1284 # echo 'L2:0=0x1;1=0x1' > p1/schemata
1285 -sh: echo: write error: Invalid argument
1286 # cat info/last_cmd_status
1287 overlaps with exclusive group
1289 Example of Cache Pseudo-Locking
1290 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1291 Lock portion of L2 cache from cache id 1 using CBM 0x3. Pseudo-locked
1292 region is exposed at /dev/pseudo_lock/newlock that can be provided to
1293 application for argument to mmap().
1294 ::
1296 # mount -t resctrl resctrl /sys/fs/resctrl/
1297 # cd /sys/fs/resctrl
1299 Ensure that there are bits available that can be pseudo-locked, since only
1300 unused bits can be pseudo-locked the bits to be pseudo-locked needs to be
1301 removed from the default resource group's schemata::
1303 # cat info/L2/bit_usage
1304 0=SSSSSSSS;1=SSSSSSSS
1305 # echo 'L2:1=0xfc' > schemata
1306 # cat info/L2/bit_usage
1307 0=SSSSSSSS;1=SSSSSS00
1309 Create a new resource group that will be associated with the pseudo-locked
1310 region, indicate that it will be used for a pseudo-locked region, and
1311 configure the requested pseudo-locked region capacity bitmask::
1313 # mkdir newlock
1314 # echo pseudo-locksetup > newlock/mode
1315 # echo 'L2:1=0x3' > newlock/schemata
1317 On success the resource group's mode will change to pseudo-locked, the
1318 bit_usage will reflect the pseudo-locked region, and the character device
1319 exposing the pseudo-locked region will exist::
1321 # cat newlock/mode
1322 pseudo-locked
1323 # cat info/L2/bit_usage
1324 0=SSSSSSSS;1=SSSSSSPP
1325 # ls -l /dev/pseudo_lock/newlock
1326 crw------- 1 root root 243, 0 Apr 3 05:01 /dev/pseudo_lock/newlock
1328 ::
1330 /*
1331 * Example code to access one page of pseudo-locked cache region
1332 * from user space.
1333 */
1334 #define _GNU_SOURCE
1335 #include <fcntl.h>
1336 #include <sched.h>
1337 #include <stdio.h>
1338 #include <stdlib.h>
1339 #include <unistd.h>
1340 #include <sys/mman.h>
1342 /*
1343 * It is required that the application runs with affinity to only
1344 * cores associated with the pseudo-locked region. Here the cpu
1345 * is hardcoded for convenience of example.
1346 */
1347 static int cpuid = 2;
1349 int main(int argc, char *argv[])
1350 {
1351 cpu_set_t cpuset;
1352 long page_size;
1353 void *mapping;
1354 int dev_fd;
1355 int ret;
1357 page_size = sysconf(_SC_PAGESIZE);
1359 CPU_ZERO(&cpuset);
1360 CPU_SET(cpuid, &cpuset);
1361 ret = sched_setaffinity(0, sizeof(cpuset), &cpuset);
1362 if (ret < 0) {
1363 perror("sched_setaffinity");
1364 exit(EXIT_FAILURE);
1365 }
1367 dev_fd = open("/dev/pseudo_lock/newlock", O_RDWR);
1368 if (dev_fd < 0) {
1369 perror("open");
1370 exit(EXIT_FAILURE);
1371 }
1373 mapping = mmap(0, page_size, PROT_READ | PROT_WRITE, MAP_SHARED,
1374 dev_fd, 0);
1375 if (mapping == MAP_FAILED) {
1376 perror("mmap");
1377 close(dev_fd);
1378 exit(EXIT_FAILURE);
1379 }
1381 /* Application interacts with pseudo-locked memory @mapping */
1383 ret = munmap(mapping, page_size);
1384 if (ret < 0) {
1385 perror("munmap");
1386 close(dev_fd);
1387 exit(EXIT_FAILURE);
1388 }
1390 close(dev_fd);
1391 exit(EXIT_SUCCESS);
1392 }
1394 Locking between applications
1395 ----------------------------
1397 Certain operations on the resctrl filesystem, composed of read/writes
1398 to/from multiple files, must be atomic.
1400 As an example, the allocation of an exclusive reservation of L3 cache
1401 involves:
1403 1. Read the cbmmasks from each directory or the per-resource "bit_usage"
1404 2. Find a contiguous set of bits in the global CBM bitmask that is clear
1405 in any of the directory cbmmasks
1406 3. Create a new directory
1407 4. Set the bits found in step 2 to the new directory "schemata" file
1409 If two applications attempt to allocate space concurrently then they can
1410 end up allocating the same bits so the reservations are shared instead of
1411 exclusive.
1413 To coordinate atomic operations on the resctrlfs and to avoid the problem
1414 above, the following locking procedure is recommended:
1416 Locking is based on flock, which is available in libc and also as a shell
1417 script command
1419 Write lock:
1421 A) Take flock(LOCK_EX) on /sys/fs/resctrl
1422 B) Read/write the directory structure.
1423 C) funlock
1425 Read lock:
1427 A) Take flock(LOCK_SH) on /sys/fs/resctrl
1428 B) If success read the directory structure.
1429 C) funlock
1431 Example with bash::
1433 # Atomically read directory structure
1434 $ flock -s /sys/fs/resctrl/ find /sys/fs/resctrl
1436 # Read directory contents and create new subdirectory
1438 $ cat create-dir.sh
1439 find /sys/fs/resctrl/ > output.txt
1440 mask = function-of(output.txt)
1441 mkdir /sys/fs/resctrl/newres/
1442 echo mask > /sys/fs/resctrl/newres/schemata
1444 $ flock /sys/fs/resctrl/ ./create-dir.sh
1446 Example with C::
1448 /*
1449 * Example code do take advisory locks
1450 * before accessing resctrl filesystem
1451 */
1452 #include <sys/file.h>
1453 #include <stdlib.h>
1455 void resctrl_take_shared_lock(int fd)
1456 {
1457 int ret;
1459 /* take shared lock on resctrl filesystem */
1460 ret = flock(fd, LOCK_SH);
1461 if (ret) {
1462 perror("flock");
1463 exit(-1);
1464 }
1465 }
1467 void resctrl_take_exclusive_lock(int fd)
1468 {
1469 int ret;
1471 /* release lock on resctrl filesystem */
1472 ret = flock(fd, LOCK_EX);
1473 if (ret) {
1474 perror("flock");
1475 exit(-1);
1476 }
1477 }
1479 void resctrl_release_lock(int fd)
1480 {
1481 int ret;
1483 /* take shared lock on resctrl filesystem */
1484 ret = flock(fd, LOCK_UN);
1485 if (ret) {
1486 perror("flock");
1487 exit(-1);
1488 }
1489 }
1491 void main(void)
1492 {
1493 int fd, ret;
1495 fd = open("/sys/fs/resctrl", O_DIRECTORY);
1496 if (fd == -1) {
1497 perror("open");
1498 exit(-1);
1499 }
1500 resctrl_take_shared_lock(fd);
1501 /* code to read directory contents */
1502 resctrl_release_lock(fd);
1504 resctrl_take_exclusive_lock(fd);
1505 /* code to read and write directory contents */
1506 resctrl_release_lock(fd);
1507 }
1509 Examples for RDT Monitoring along with allocation usage
1510 =======================================================
1511 Reading monitored data
1512 ----------------------
1513 Reading an event file (for ex: mon_data/mon_L3_00/llc_occupancy) would
1514 show the current snapshot of LLC occupancy of the corresponding MON
1515 group or CTRL_MON group.
1518 Example 1 (Monitor CTRL_MON group and subset of tasks in CTRL_MON group)
1519 ------------------------------------------------------------------------
1520 On a two socket machine (one L3 cache per socket) with just four bits
1521 for cache bit masks::
1523 # mount -t resctrl resctrl /sys/fs/resctrl
1524 # cd /sys/fs/resctrl
1525 # mkdir p0 p1
1526 # echo "L3:0=3;1=c" > /sys/fs/resctrl/p0/schemata
1527 # echo "L3:0=3;1=3" > /sys/fs/resctrl/p1/schemata
1528 # echo 5678 > p1/tasks
1529 # echo 5679 > p1/tasks
1531 The default resource group is unmodified, so we have access to all parts
1532 of all caches (its schemata file reads "L3:0=f;1=f").
1534 Tasks that are under the control of group "p0" may only allocate from the
1535 "lower" 50% on cache ID 0, and the "upper" 50% of cache ID 1.
1536 Tasks in group "p1" use the "lower" 50% of cache on both sockets.
1538 Create monitor groups and assign a subset of tasks to each monitor group.
1539 ::
1541 # cd /sys/fs/resctrl/p1/mon_groups
1542 # mkdir m11 m12
1543 # echo 5678 > m11/tasks
1544 # echo 5679 > m12/tasks
1546 fetch data (data shown in bytes)
1547 ::
1549 # cat m11/mon_data/mon_L3_00/llc_occupancy
1550 16234000
1551 # cat m11/mon_data/mon_L3_01/llc_occupancy
1552 14789000
1553 # cat m12/mon_data/mon_L3_00/llc_occupancy
1554 16789000
1556 The parent ctrl_mon group shows the aggregated data.
1557 ::
1559 # cat /sys/fs/resctrl/p1/mon_data/mon_l3_00/llc_occupancy
1560 31234000
1562 Example 2 (Monitor a task from its creation)
1563 --------------------------------------------
1564 On a two socket machine (one L3 cache per socket)::
1566 # mount -t resctrl resctrl /sys/fs/resctrl
1567 # cd /sys/fs/resctrl
1568 # mkdir p0 p1
1570 An RMID is allocated to the group once its created and hence the <cmd>
1571 below is monitored from its creation.
1572 ::
1574 # echo $$ > /sys/fs/resctrl/p1/tasks
1575 # <cmd>
1577 Fetch the data::
1579 # cat /sys/fs/resctrl/p1/mon_data/mon_l3_00/llc_occupancy
1580 31789000
1582 Example 3 (Monitor without CAT support or before creating CAT groups)
1583 ---------------------------------------------------------------------
1585 Assume a system like HSW has only CQM and no CAT support. In this case
1586 the resctrl will still mount but cannot create CTRL_MON directories.
1587 But user can create different MON groups within the root group thereby
1588 able to monitor all tasks including kernel threads.
1590 This can also be used to profile jobs cache size footprint before being
1591 able to allocate them to different allocation groups.
1592 ::
1594 # mount -t resctrl resctrl /sys/fs/resctrl
1595 # cd /sys/fs/resctrl
1596 # mkdir mon_groups/m01
1597 # mkdir mon_groups/m02
1599 # echo 3478 > /sys/fs/resctrl/mon_groups/m01/tasks
1600 # echo 2467 > /sys/fs/resctrl/mon_groups/m02/tasks
1602 Monitor the groups separately and also get per domain data. From the
1603 below its apparent that the tasks are mostly doing work on
1604 domain(socket) 0.
1605 ::
1607 # cat /sys/fs/resctrl/mon_groups/m01/mon_L3_00/llc_occupancy
1608 31234000
1609 # cat /sys/fs/resctrl/mon_groups/m01/mon_L3_01/llc_occupancy
1610 34555
1611 # cat /sys/fs/resctrl/mon_groups/m02/mon_L3_00/llc_occupancy
1612 31234000
1613 # cat /sys/fs/resctrl/mon_groups/m02/mon_L3_01/llc_occupancy
1614 32789
1617 Example 4 (Monitor real time tasks)
1618 -----------------------------------
1620 A single socket system which has real time tasks running on cores 4-7
1621 and non real time tasks on other cpus. We want to monitor the cache
1622 occupancy of the real time threads on these cores.
1623 ::
1625 # mount -t resctrl resctrl /sys/fs/resctrl
1626 # cd /sys/fs/resctrl
1627 # mkdir p1
1629 Move the cpus 4-7 over to p1::
1631 # echo f0 > p1/cpus
1633 View the llc occupancy snapshot::
1635 # cat /sys/fs/resctrl/p1/mon_data/mon_L3_00/llc_occupancy
1636 11234000
1639 Examples on working with mbm_assign_mode
1640 ========================================
1642 a. Check if MBM counter assignment mode is supported.
1643 ::
1645 # mount -t resctrl resctrl /sys/fs/resctrl/
1647 # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
1648 [mbm_event]
1649 default
1651 The "mbm_event" mode is detected and enabled.
1653 b. Check how many assignable counters are supported.
1654 ::
1656 # cat /sys/fs/resctrl/info/L3_MON/num_mbm_cntrs
1657 0=32;1=32
1659 c. Check how many assignable counters are available for assignment in each domain.
1660 ::
1662 # cat /sys/fs/resctrl/info/L3_MON/available_mbm_cntrs
1663 0=30;1=30
1665 d. To list the default group's assign states.
1666 ::
1668 # cat /sys/fs/resctrl/mbm_L3_assignments
1669 mbm_total_bytes:0=e;1=e
1670 mbm_local_bytes:0=e;1=e
1672 e. To unassign the counter associated with the mbm_total_bytes event on domain 0.
1673 ::
1675 # echo "mbm_total_bytes:0=_" > /sys/fs/resctrl/mbm_L3_assignments
1676 # cat /sys/fs/resctrl/mbm_L3_assignments
1677 mbm_total_bytes:0=_;1=e
1678 mbm_local_bytes:0=e;1=e
1680 f. To unassign the counter associated with the mbm_total_bytes event on all domains.
1681 ::
1683 # echo "mbm_total_bytes:*=_" > /sys/fs/resctrl/mbm_L3_assignments
1684 # cat /sys/fs/resctrl/mbm_L3_assignment
1685 mbm_total_bytes:0=_;1=_
1686 mbm_local_bytes:0=e;1=e
1688 g. To assign a counter associated with the mbm_total_bytes event on all domains in
1689 exclusive mode.
1690 ::
1692 # echo "mbm_total_bytes:*=e" > /sys/fs/resctrl/mbm_L3_assignments
1693 # cat /sys/fs/resctrl/mbm_L3_assignments
1694 mbm_total_bytes:0=e;1=e
1695 mbm_local_bytes:0=e;1=e
1697 h. Read the events mbm_total_bytes and mbm_local_bytes of the default group. There is
1698 no change in reading the events with the assignment.
1699 ::
1701 # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_total_bytes
1702 779247936
1703 # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_total_bytes
1704 562324232
1705 # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
1706 212122123
1707 # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
1708 121212144
1710 i. Check the event configurations.
1711 ::
1713 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
1714 local_reads,remote_reads,local_non_temporal_writes,remote_non_temporal_writes,
1715 local_reads_slow_memory,remote_reads_slow_memory,dirty_victim_writes_all
1717 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
1718 local_reads,local_non_temporal_writes,local_reads_slow_memory
1720 j. Change the event configuration for mbm_local_bytes.
1721 ::
1723 # echo "local_reads, local_non_temporal_writes, local_reads_slow_memory, remote_reads" >
1724 /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
1726 # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
1727 local_reads,local_non_temporal_writes,local_reads_slow_memory,remote_reads
1729 k. Now read the local events again. The first read may come back with "Unavailable"
1730 status. The subsequent read of mbm_local_bytes will display the current value.
1731 ::
1733 # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
1734 Unavailable
1735 # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
1736 2252323
1737 # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
1738 Unavailable
1739 # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
1740 1566565
1742 l. Users have the option to go back to 'default' mbm_assign_mode if required. This can be
1743 done using the following command. Note that switching the mbm_assign_mode may reset all
1744 the MBM counters (and thus all MBM events) of all the resctrl groups.
1745 ::
1747 # echo "default" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
1748 # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
1749 mbm_event
1750 [default]
1752 m. Unmount the resctrl filesystem.
1753 ::
1755 # umount /sys/fs/resctrl/
1757 Intel RDT Errata
1758 ================
1760 Intel MBM Counters May Report System Memory Bandwidth Incorrectly
1761 -----------------------------------------------------------------
1763 Errata SKX99 for Skylake server and BDF102 for Broadwell server.
1765 Problem: Intel Memory Bandwidth Monitoring (MBM) counters track metrics
1766 according to the assigned Resource Monitor ID (RMID) for that logical
1767 core. The IA32_QM_CTR register (MSR 0xC8E), used to report these
1768 metrics, may report incorrect system bandwidth for certain RMID values.
1770 Implication: Due to the errata, system memory bandwidth may not match
1771 what is reported.
1773 Workaround: MBM total and local readings are corrected according to the
1774 following correction factor table:
1776 +---------------+---------------+---------------+-----------------+
1777 |core count |rmid count |rmid threshold |correction factor|
1778 +---------------+---------------+---------------+-----------------+
1779 |1 |8 |0 |1.000000 |
1780 +---------------+---------------+---------------+-----------------+
1781 |2 |16 |0 |1.000000 |
1782 +---------------+---------------+---------------+-----------------+
1783 |3 |24 |15 |0.969650 |
1784 +---------------+---------------+---------------+-----------------+
1785 |4 |32 |0 |1.000000 |
1786 +---------------+---------------+---------------+-----------------+
1787 |6 |48 |31 |0.969650 |
1788 +---------------+---------------+---------------+-----------------+
1789 |7 |56 |47 |1.142857 |
1790 +---------------+---------------+---------------+-----------------+
1791 |8 |64 |0 |1.000000 |
1792 +---------------+---------------+---------------+-----------------+
1793 |9 |72 |63 |1.185115 |
1794 +---------------+---------------+---------------+-----------------+
1795 |10 |80 |63 |1.066553 |
1796 +---------------+---------------+---------------+-----------------+
1797 |11 |88 |79 |1.454545 |
1798 +---------------+---------------+---------------+-----------------+
1799 |12 |96 |0 |1.000000 |
1800 +---------------+---------------+---------------+-----------------+
1801 |13 |104 |95 |1.230769 |
1802 +---------------+---------------+---------------+-----------------+
1803 |14 |112 |95 |1.142857 |
1804 +---------------+---------------+---------------+-----------------+
1805 |15 |120 |95 |1.066667 |
1806 +---------------+---------------+---------------+-----------------+
1807 |16 |128 |0 |1.000000 |
1808 +---------------+---------------+---------------+-----------------+
1809 |17 |136 |127 |1.254863 |
1810 +---------------+---------------+---------------+-----------------+
1811 |18 |144 |127 |1.185255 |
1812 +---------------+---------------+---------------+-----------------+
1813 |19 |152 |0 |1.000000 |
1814 +---------------+---------------+---------------+-----------------+
1815 |20 |160 |127 |1.066667 |
1816 +---------------+---------------+---------------+-----------------+
1817 |21 |168 |0 |1.000000 |
1818 +---------------+---------------+---------------+-----------------+
1819 |22 |176 |159 |1.454334 |
1820 +---------------+---------------+---------------+-----------------+
1821 |23 |184 |0 |1.000000 |
1822 +---------------+---------------+---------------+-----------------+
1823 |24 |192 |127 |0.969744 |
1824 +---------------+---------------+---------------+-----------------+
1825 |25 |200 |191 |1.280246 |
1826 +---------------+---------------+---------------+-----------------+
1827 |26 |208 |191 |1.230921 |
1828 +---------------+---------------+---------------+-----------------+
1829 |27 |216 |0 |1.000000 |
1830 +---------------+---------------+---------------+-----------------+
1831 |28 |224 |191 |1.143118 |
1832 +---------------+---------------+---------------+-----------------+
1834 If rmid > rmid threshold, MBM total and local values should be multiplied
1835 by the correction factor.
1837 See:
1839 1. Erratum SKX99 in Intel Xeon Processor Scalable Family Specification Update:
1840 http://web.archive.org/web/20200716124958/https://www.intel.com/content/www/us/en/processors/xeon/scalable/xeon-scalable-spec-update.html
1842 2. Erratum BDF102 in Intel Xeon E5-2600 v4 Processor Product Family Specification Update:
1843 http://web.archive.org/web/20191125200531/https://www.intel.com/content/dam/www/public/us/en/documents/specification-updates/xeon-e5-v4-spec-update.pdf
1845 3. The errata in Intel Resource Director Technology (Intel RDT) on 2nd Generation Intel Xeon Scalable Processors Reference Manual:
1846 https://software.intel.com/content/www/us/en/develop/articles/intel-resource-director-technology-rdt-reference-manual.html
1848 for further information.

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

Resctrl 기능과 마운트

1-67

Resctrl은 CPU의 공유 cache와 memory bandwidth를 할당하고 관찰하는 Resource Control 사용자 인터페이스다. Intel은 이를 Intel Resource Director Technology(Intel RDT), AMD는 AMD Platform Quality of Service(AMD QoS)라고 부른다.

기능은 `CONFIG_X86_CPU_RESCTRL`과 x86 `/proc/cpuinfo` flag로 활성 여부를 확인한다. 그러나 새 기능을 모두 cpuinfo flag로 추가하면 사람이 읽기 어려우므로, userspace가 resctrl의 `info` 디렉터리에서 확인할 수 있는 기능은 새 flag 추가를 피해야 한다.

Resctrl 기능과 CPU flag
기능Flag
RDT Allocation`rdt_a`
L3/L2 CAT`cat_l3`, `cat_l2`
L3/L2 CDP`cdp_l3`, `cdp_l2`
CQM`cqm_llc`, `cqm_occup_llc`
MBM`cqm_mbm_total`, `cqm_mbm_local`
MBA`mba`
SMBA, BMEC, ABMCcpuinfo flag 없이 `info`에서 탐지

할당·monitoring·우선순위 기능의 전통적인 cpuinfo 표시다.

mount -t resctrl resctrl [-o cdp[,cdpl2][,mba_MBps][,debug]] /sys/fs/resctrl
마운트 옵션
옵션효과
`cdp`L3 cache allocation에서 code/data prioritization 활성
`cdpl2`L2 cache allocation에서 code/data prioritization 활성
`mba_MBps`MBA software controller(`mba_sc`)를 켜 MiBps 단위로 지정
`debug`debug 전용 파일 노출

L2·L3 CDP는 서로 독립적으로 제어한다.

RDT 기능은 서로 직교하므로 시스템은 monitoring만, control만, 또는 둘 다 지원할 수 있다. Cache pseudo-locking은 cache control을 이용해 데이터를 cache에 pin하는 독특한 사용법이다.

할당 또는 monitoring 중 하나만 있어도 마운트는 성공하지만, 실제 시스템이 지원하는 파일과 디렉터리만 생성된다.

.. SPDX-License-Identifier: GPL-2.0
.. include:: <isonum.txt>

=====================================================
User Interface for Resource Control feature (resctrl)
=====================================================

:Copyright: |copy| 2016 Intel Corporation
:Authors: - Fenghua Yu <fenghua.yu@intel.com>
          - Tony Luck <tony.luck@intel.com>
          - Vikas Shivappa <vikas.shivappa@intel.com>


Intel refers to this feature as Intel Resource Director Technology(Intel(R) RDT).
AMD refers to this feature as AMD Platform Quality of Service(AMD QoS).

This feature is enabled by the CONFIG_X86_CPU_RESCTRL and the x86 /proc/cpuinfo
flag bits:

===============================================        ================================
RDT (Resource Director Technology) Allocation        "rdt_a"
CAT (Cache Allocation Technology)                "cat_l3", "cat_l2"
CDP (Code and Data Prioritization)                "cdp_l3", "cdp_l2"
CQM (Cache QoS Monitoring)                        "cqm_llc", "cqm_occup_llc"
MBM (Memory Bandwidth Monitoring)                "cqm_mbm_total", "cqm_mbm_local"
MBA (Memory Bandwidth Allocation)                "mba"
SMBA (Slow Memory Bandwidth Allocation)         ""
BMEC (Bandwidth Monitoring Event Configuration) ""
ABMC (Assignable Bandwidth Monitoring Counters) ""
===============================================        ================================

Historically, new features were made visible by default in /proc/cpuinfo. This
resulted in the feature flags becoming hard to parse by humans. Adding a new
flag to /proc/cpuinfo should be avoided if user space can obtain information
about the feature from resctrl's info directory.

To use the feature mount the file system::

 # mount -t resctrl resctrl [-o cdp[,cdpl2][,mba_MBps][,debug]] /sys/fs/resctrl

mount options are:

"cdp":
        Enable code/data prioritization in L3 cache allocations.
"cdpl2":
        Enable code/data prioritization in L2 cache allocations.
"mba_MBps":
        Enable the MBA Software Controller(mba_sc) to specify MBA
        bandwidth in MiBps
"debug":
        Make debug files accessible. Available debug files are annotated with
        "Available only with debug option".

L2 and L3 CDP are controlled separately.

RDT features are orthogonal. A particular system may support only
monitoring, only control, or both monitoring and control.  Cache
pseudo-locking is a unique way of using cache control to "pin" or
"lock" data in the cache. Details can be found in
"Cache Pseudo-Locking".


The mount succeeds if either of allocation or monitoring is present, but
only those files and directories supported by the system will be created.
For more details on the behavior of the interface during monitoring
and allocation, see the "Resource alloc and monitor groups" section.

`info` 할당 리소스 파일

68-150

`info` 디렉터리는 활성화된 resource 정보를 담고 resource마다 이름이 같은 하위 디렉터리를 만든다.

Cache resource 정보
파일의미
`num_closids`resource에서 유효한 CLOSID 수; kernel은 활성 resource 중 최소값을 한도로 사용
`cbm_mask`resource의 100%를 나타내는 유효 bitmask
`min_cbm_bits`mask에 반드시 연속해서 설정해야 하는 최소 bit 수
`shareable_bits`I/O 같은 다른 실행 주체와 공유할 수 있는 bit; 장치 설정이 덮어쓸 수 있음
`sparse_masks`CBM의 비연속 1 지원 여부: `0` 연속만, `1` 비연속 허용

L3/L2 디렉터리의 allocation 관련 파일이다.

`bit_usage`는 resource instance마다 capacity bit 사용 상태를 문자로 주석 처리해 보여 준다. 모든 할당 뒤에도 `0`이 있으면 resource가 낭비되고 있다는 신호다.

`bit_usage` 범례
표시상태
`0`미사용
`H`hardware만 사용하지만 software 사용 가능
`X`공유 가능하며 hardware와 software가 모두 사용
`S`software가 사용하며 공유 가능
`E`resource group 하나가 독점, 공유 금지
`P`pseudo-locked, 공유 금지

각 cache way가 hardware·software·독점·pseudo-lock 중 어디에 쓰이는지 나타낸다.

Memory bandwidth 정보
파일의미
`min_bandwidth`요청 가능한 최소 bandwidth 백분율
`bandwidth_gran`hardware 제어 step; `min_bandwidth + N * bandwidth_gran`
`delay_linear`delay scale의 선형 여부를 알리는 정보 전용 값
`thread_throttle_mode=max`physical core thread들의 요청 중 가장 작은 백분율을 모두에 적용
`thread_throttle_mode=per-thread`각 thread에 요청 백분율을 직접 적용

`info/MB`의 allocation 제어 특성이다.

요청 값이 hardware step 사이에 있으면 다음 사용 가능한 제어 step으로 반올림된다.

Info directory
==============

The 'info' directory contains information about the enabled
resources. Each resource has its own subdirectory. The subdirectory
names reflect the resource names.

Each subdirectory contains the following files with respect to
allocation:

Cache resource(L3/L2)  subdirectory contains the following files
related to allocation:

"num_closids":
                The number of CLOSIDs which are valid for this
                resource. The kernel uses the smallest number of
                CLOSIDs of all enabled resources as limit.
"cbm_mask":
                The bitmask which is valid for this resource.
                This mask is equivalent to 100%.
"min_cbm_bits":
                The minimum number of consecutive bits which
                must be set when writing a mask.

"shareable_bits":
                Bitmask of shareable resource with other executing
                entities (e.g. I/O). User can use this when
                setting up exclusive cache partitions. Note that
                some platforms support devices that have their
                own settings for cache use which can over-ride
                these bits.
"bit_usage":
                Annotated capacity bitmasks showing how all
                instances of the resource are used. The legend is:

                        "0":
                              Corresponding region is unused. When the system's
                              resources have been allocated and a "0" is found
                              in "bit_usage" it is a sign that resources are
                              wasted.

                        "H":
                              Corresponding region is used by hardware only
                              but available for software use. If a resource
                              has bits set in "shareable_bits" but not all
                              of these bits appear in the resource groups'
                              schematas then the bits appearing in
                              "shareable_bits" but no resource group will
                              be marked as "H".
                        "X":
                              Corresponding region is available for sharing and
                              used by hardware and software. These are the
                              bits that appear in "shareable_bits" as
                              well as a resource group's allocation.
                        "S":
                              Corresponding region is used by software
                              and available for sharing.
                        "E":
                              Corresponding region is used exclusively by
                              one resource group. No sharing allowed.
                        "P":
                              Corresponding region is pseudo-locked. No
                              sharing allowed.
"sparse_masks":
                Indicates if non-contiguous 1s value in CBM is supported.

                        "0":
                              Only contiguous 1s value in CBM is supported.
                        "1":
                              Non-contiguous 1s value in CBM is supported.

Memory bandwidth(MB) subdirectory contains the following files
with respect to allocation:

"min_bandwidth":
                The minimum memory bandwidth percentage which
                user can request.

"bandwidth_gran":
                The granularity in which the memory bandwidth
                percentage is allocated. The allocated
                b/w percentage is rounded off to the next
                control step available on the hardware. The

`L3_MON`과 BMEC event 구성

151-241

RDT monitoring이 있으면 `info/L3_MON` 디렉터리가 생긴다. `num_rmids`는 생성 가능한 `CTRL_MON + MON` group 수의 상한인 RMID 개수다. `mon_features`는 `llc_occupancy`, `mbm_total_bytes`, `mbm_local_bytes` 같은 활성 monitoring event를 나열한다.

BMEC를 지원하면 `mbm_total_bytes_config`와 `mbm_local_bytes_config`가 추가된다. 이 read/write 파일은 domain별 event mask를 저장하며 같은 domain의 모든 CPU에 영향을 준다.

어느 하나의 event 구성을 바꾸면 해당 domain에서 두 event의 모든 RMID bandwidth counter가 초기화된다. 각 RMID의 다음 첫 read는 `Unavailable`, 이후 read는 유효값을 반환한다.

BMEC mask bit
BitTransaction
6QoS domain에서 모든 memory로 나가는 dirty victim
5non-local NUMA domain의 slow memory read
4local NUMA domain의 slow memory read
3non-local NUMA domain non-temporal write
2local NUMA domain non-temporal write
1non-local NUMA domain memory read
0local NUMA domain memory read

Bandwidth counter에 포함할 memory transaction 종류다.

기본 `mbm_total_bytes_config`는 모든 event를 세는 `0x7f`, `mbm_local_bytes_config`는 local memory event를 세는 `0x15`다. 출력 형식은 `domain=mask`를 세미콜론으로 연결한다.

# 현재 구성
cat /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
cat /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config

# domain 0에서 read만 집계: bits 0,1,4,5 = 0x33
echo "0=0x33" > /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config

# domain 0,1에서 slow memory read 집계: bits 4,5 = 0x30
echo "0=0x30;1=0x30" > /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config

부분 write는 지정한 domain만 바꾸며 나머지 domain mask는 기존 값을 유지한다.

                available bandwidth control steps are:
                min_bandwidth + N * bandwidth_gran.

"delay_linear":
                Indicates if the delay scale is linear or
                non-linear. This field is purely informational
                only.

"thread_throttle_mode":
                Indicator on Intel systems of how tasks running on threads
                of a physical core are throttled in cases where they
                request different memory bandwidth percentages:

                "max":
                        the smallest percentage is applied
                        to all threads
                "per-thread":
                        bandwidth percentages are directly applied to
                        the threads running on the core

If RDT monitoring is available there will be an "L3_MON" directory
with the following files:

"num_rmids":
                The number of RMIDs available. This is the
                upper bound for how many "CTRL_MON" + "MON"
                groups can be created.

"mon_features":
                Lists the monitoring events if
                monitoring is enabled for the resource.
                Example::

                        # cat /sys/fs/resctrl/info/L3_MON/mon_features
                        llc_occupancy
                        mbm_total_bytes
                        mbm_local_bytes

                If the system supports Bandwidth Monitoring Event
                Configuration (BMEC), then the bandwidth events will
                be configurable. The output will be::

                        # cat /sys/fs/resctrl/info/L3_MON/mon_features
                        llc_occupancy
                        mbm_total_bytes
                        mbm_total_bytes_config
                        mbm_local_bytes
                        mbm_local_bytes_config

"mbm_total_bytes_config", "mbm_local_bytes_config":
        Read/write files containing the configuration for the mbm_total_bytes
        and mbm_local_bytes events, respectively, when the Bandwidth
        Monitoring Event Configuration (BMEC) feature is supported.
        The event configuration settings are domain specific and affect
        all the CPUs in the domain. When either event configuration is
        changed, the bandwidth counters for all RMIDs of both events
        (mbm_total_bytes as well as mbm_local_bytes) are cleared for that
        domain. The next read for every RMID will report "Unavailable"
        and subsequent reads will report the valid value.

        Following are the types of events supported:

        ====    ========================================================
        Bits    Description
        ====    ========================================================
        6       Dirty Victims from the QOS domain to all types of memory
        5       Reads to slow memory in the non-local NUMA domain
        4       Reads to slow memory in the local NUMA domain
        3       Non-temporal writes to non-local NUMA domain
        2       Non-temporal writes to local NUMA domain
        1       Reads to memory in the non-local NUMA domain
        0       Reads to memory in the local NUMA domain
        ====    ========================================================

        By default, the mbm_total_bytes configuration is set to 0x7f to count
        all the event types and the mbm_local_bytes configuration is set to
        0x15 to count all the local memory events.

        Examples:

        * To view the current configuration::
          ::

            # cat /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
            0=0x7f;1=0x7f;2=0x7f;3=0x7f

            # cat /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config
            0=0x15;1=0x15;3=0x15;4=0x15

        * To change the mbm_total_bytes to count only reads on domain 0,
          the bits 0, 1, 4 and 5 needs to be set, which is 110011b in binary

MBM counter assignment와 event filter

242-415

`mbm_assign_mode`는 지원되는 counter 할당 mode를 보여 주며 대괄호가 현재 mode를 표시한다. Mode 전환 때 MBM event와 연결된 counter가 reset될 수 있다.

`mbm_assign_mode`
Mode동작
`mbm_event`사용자가 RMID·event 쌍에 counter를 명시적으로 할당하고 해제할 때까지 유지
`default`모든 CTRL_MON/MON group의 각 event에 hardware counter가 있다고 가정

Hardware counter가 RMID event에 언제 연결되는지 정의한다.

`mbm_event`에서는 hardware counter가 붙어 있을 때만 event가 누적된다. 각 group의 `mbm_L3_assignments`로 event별 할당을 정하고 `num_mbm_cntrs`에서 최대 counter 수를 확인한다. 할당하지 않으면 read 결과는 `Unassigned`다.

이 mode는 group 수가 hardware counter보다 많은 AMD 플랫폼에 유리하다. ABMC capability가 있는 AMD에서는 기본 활성화돼 RMID가 CPU에서 사용 중이 아니어도 counter 연결을 유지한다. AMD에서 지원된다면 hardware 재할당에 따른 reset·오해 가능한 값·`Unavailable`을 피하려고 `mbm_event`를 권장한다.

echo "mbm_event" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
echo "default" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
cat /sys/fs/resctrl/info/L3_MON/num_mbm_cntrs
cat /sys/fs/resctrl/info/L3_MON/available_mbm_cntrs

`num_mbm_cntrs`는 domain별 전체 counter 수, `available_mbm_cntrs`는 현재 할당 가능한 counter 수다. 예시는 각 L3 domain에 최대 32개, 그중 30개가 사용 가능함을 보여 준다.

`event_configs`는 할당 가능한 MBM event마다 하위 디렉터리를 가진다. 기본 event는 `mbm_local_bytes`, `mbm_total_bytes`이며, `mbm_event` mode에서만 접근 가능한 `event_filter`가 집계할 transaction을 정한다.

`event_filter` transaction 이름
이름설명
`dirty_victim_writes_all`모든 memory로 나가는 dirty victim
`remote_reads_slow_memory`remote NUMA slow memory read
`local_reads_slow_memory`local NUMA slow memory read
`remote_non_temporal_writes`remote non-temporal write
`local_non_temporal_writes`local non-temporal write
`remote_reads`remote memory read
`local_reads`local memory read

BMEC bit에 대응하는 사람이 읽을 수 있는 이름이다.

`event_filter`를 읽으면 쉼표로 연결한 현재 구성, 쓰면 새 transaction 집합을 지정한다.

`mbm_assign_on_mkdir`는 monitor group을 `mkdir`할 때 RMID·MBM event 쌍에 counter를 자동 할당할지 정한다. Boot와 `default`에서 `mbm_event`로 전환할 때 기본값은 `1`; `0`은 자동 할당 비활성이다.

`max_threshold_occupancy`는 사용했던 LLC occupancy counter를 재사용 가능하다고 볼 최대 byte 값이다.

`info/last_cmd_status`는 디렉터리 생성이나 control 파일 write마다 reset된다. 성공하면 `ok`, 실패하면 file operation의 errno보다 자세한 이유를 제공한다. 예를 들어 비연속 CBM `f7` write는 `mask f7 has non-consecutive 1-bits`를 남긴다.

          (in hexadecimal 0x33):
          ::

            # echo  "0=0x33" > /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config

            # cat /sys/fs/resctrl/info/L3_MON/mbm_total_bytes_config
            0=0x33;1=0x7f;2=0x7f;3=0x7f

        * To change the mbm_local_bytes to count all the slow memory reads on
          domain 0 and 1, the bits 4 and 5 needs to be set, which is 110000b
          in binary (in hexadecimal 0x30):
          ::

            # echo  "0=0x30;1=0x30" > /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config

            # cat /sys/fs/resctrl/info/L3_MON/mbm_local_bytes_config
            0=0x30;1=0x30;3=0x15;4=0x15

"mbm_assign_mode":
        The supported counter assignment modes. The enclosed brackets indicate which mode
        is enabled. The MBM events associated with counters may reset when "mbm_assign_mode"
        is changed.
        ::

          # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
          [mbm_event]
          default

        "mbm_event":

        mbm_event mode allows users to assign a hardware counter to an RMID, event
        pair and monitor the bandwidth usage as long as it is assigned. The hardware
        continues to track the assigned counter until it is explicitly unassigned by
        the user. Each event within a resctrl group can be assigned independently.

        In this mode, a monitoring event can only accumulate data while it is backed
        by a hardware counter. Use "mbm_L3_assignments" found in each CTRL_MON and MON
        group to specify which of the events should have a counter assigned. The number
        of counters available is described in the "num_mbm_cntrs" file. Changing the
        mode may cause all counters on the resource to reset.

        Moving to mbm_event counter assignment mode requires users to assign the counters
        to the events. Otherwise, the MBM event counters will return 'Unassigned' when read.

        The mode is beneficial for AMD platforms that support more CTRL_MON
        and MON groups than available hardware counters. By default, this
        feature is enabled on AMD platforms with the ABMC (Assignable Bandwidth
        Monitoring Counters) capability, ensuring counters remain assigned even
        when the corresponding RMID is not actively used by any processor.

        "default":

        In default mode, resctrl assumes there is a hardware counter for each
        event within every CTRL_MON and MON group. On AMD platforms, it is
        recommended to use the mbm_event mode, if supported, to prevent reset of MBM
        events between reads resulting from hardware re-allocating counters. This can
        result in misleading values or display "Unavailable" if no counter is assigned
        to the event.

        * To enable "mbm_event" counter assignment mode:
          ::

            # echo "mbm_event" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode

        * To enable "default" monitoring mode:
          ::

            # echo "default" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode

"num_mbm_cntrs":
        The maximum number of counters (total of available and assigned counters) in
        each domain when the system supports mbm_event mode.

        For example, on a system with maximum of 32 memory bandwidth monitoring
        counters in each of its L3 domains:
        ::

          # cat /sys/fs/resctrl/info/L3_MON/num_mbm_cntrs
          0=32;1=32

"available_mbm_cntrs":
        The number of counters available for assignment in each domain when mbm_event
        mode is enabled on the system.

        For example, on a system with 30 available [hardware] assignable counters
        in each of its L3 domains:
        ::

          # cat /sys/fs/resctrl/info/L3_MON/available_mbm_cntrs
          0=30;1=30

"event_configs":
        Directory that exists when "mbm_event" counter assignment mode is supported.
        Contains a sub-directory for each MBM event that can be assigned to a counter.

        Two MBM events are supported by default: mbm_local_bytes and mbm_total_bytes.
        Each MBM event's sub-directory contains a file named "event_filter" that is
        used to view and modify which memory transactions the MBM event is configured
        with. The file is accessible only when "mbm_event" counter assignment mode is
        enabled.

        List of memory transaction types supported:

        ==========================  ========================================================
        Name                            Description
        ==========================  ========================================================
        dirty_victim_writes_all     Dirty Victims from the QOS domain to all types of memory
        remote_reads_slow_memory    Reads to slow memory in the non-local NUMA domain
        local_reads_slow_memory     Reads to slow memory in the local NUMA domain
        remote_non_temporal_writes  Non-temporal writes to non-local NUMA domain
        local_non_temporal_writes   Non-temporal writes to local NUMA domain
        remote_reads                Reads to memory in the non-local NUMA domain
        local_reads                 Reads to memory in the local NUMA domain
        ==========================  ========================================================

        For example::

          # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
          local_reads,remote_reads,local_non_temporal_writes,remote_non_temporal_writes,
          local_reads_slow_memory,remote_reads_slow_memory,dirty_victim_writes_all

          # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
          local_reads,local_non_temporal_writes,local_reads_slow_memory

        Modify the event configuration by writing to the "event_filter" file within
        the "event_configs" directory. The read/write "event_filter" file contains the
        configuration of the event that reflects which memory transactions are counted by it.

        For example::

          # echo "local_reads, local_non_temporal_writes" >
            /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter

          # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
           local_reads,local_non_temporal_writes

"mbm_assign_on_mkdir":
        Exists when "mbm_event" counter assignment mode is supported. Accessible
        only when "mbm_event" counter assignment mode is enabled.

        Determines if a counter will automatically be assigned to an RMID, MBM event
        pair when its associated monitor group is created via mkdir. Enabled by default
        on boot, also when switched from "default" mode to "mbm_event" counter assignment
        mode. Users can disable this capability by writing to the interface.

        "0":
                Auto assignment is disabled.
        "1":
                Auto assignment is enabled.

        Example::

          # echo 0 > /sys/fs/resctrl/info/L3_MON/mbm_assign_on_mkdir
          # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_on_mkdir
          0

"max_threshold_occupancy":
                Read/write file provides the largest value (in
                bytes) at which a previously used LLC_occupancy
                counter can be considered for re-use.

Finally, in the top level of the "info" directory there is a file
named "last_cmd_status". This is reset with every "command" issued
via the file system (making new directories or writing to any of the
control files). If the command was successful, it will read as "ok".
If the command failed, it will provide more information that can be
conveyed in the error returns from file operations. E.g.
::

        # echo L3:0=f7 > schemata
        bash: echo: write error: Invalid argument
        # cat info/last_cmd_status
        mask f7 has non-consecutive 1-bits

CTRL_MON·MON group과 파일

416-602

Resource group은 resctrl 디렉터리로 표현한다. 마운트 직후 root인 default group은 시스템의 모든 task와 CPU를 소유하고 모든 resource를 전부 사용할 수 있다.

Control 기능이 있으면 root 아래에 resource 양을 달리하는 디렉터리를 만들 수 있다. Root와 이 top-level 디렉터리를 `CTRL_MON` group이라 한다. Monitoring이 있으면 각 CTRL_MON의 `mon_groups` 아래에 그 조상 group의 task subset을 관찰하는 `MON` group을 만든다.

Group 디렉터리를 제거하면 소유 task와 CPU가 부모로 이동한다. CTRL_MON을 제거하면 아래 MON도 모두 자동 제거된다. CPU를 monitor하는 MON은 이동할 수 없지만 task만 monitor하는 MON은 monitoring data와 task를 유지한 채 새 CTRL_MON 부모로 옮겨 allocation을 바꿀 수 있다. 그 밖에는 단순 rename만 허용한다.

모든 group의 기본 파일
파일동작
`tasks`task ID 목록; 쉼표 입력을 순차 할당, 첫 실패에서 중단하며 앞선 성공은 유지
`cpus`group 소유 logical CPU bitmask; pseudo-locked mode에서는 read-only
`cpus_list``cpus`와 같지만 CPU range 형식

Task와 CPU 소유권은 CTRL_MON 계층을 따른다.

Task를 CTRL_MON에 쓰면 이전 CTRL_MON과 모든 MON에서 제거된다. MON에 쓰려면 이미 그 MON의 부모 CTRL_MON에 속해야 하며 이전 MON에서는 제거된다. 여러 task 할당 중 실패는 `last_cmd_status`에 기록된다.

Control group 파일
파일의미
`schemata`group이 사용할 모든 resource의 domain별 값
`size`schemata bit 대신 allocation byte 크기 표시
`mode``shareable`, `exclusive`, `pseudo-locksetup`, `pseudo-locked`
`ctrl_hw_id`debug 전용 hardware control ID; x86에서는 CLOSID

Control이 활성화된 모든 CTRL_MON에 존재한다.

Pseudo-locked region은 `mode`에 `pseudo-locksetup`을 쓴 뒤 `schemata`에 cache 영역을 쓰면 만들며, 성공하면 mode가 자동으로 `pseudo-locked`가 된다.

Monitoring이 활성화되면 `mon_data`가 L3 domain과 event별 파일을 가진다. 두 domain이면 `mon_L3_00`, `mon_L3_01` 아래에 `llc_occupancy`, `mbm_total_bytes`, `mbm_local_bytes` 등이 생긴다. MON에서는 해당 group task 값, CTRL_MON에서는 자체 task와 모든 하위 MON task의 합을 보여 준다.

SNC에서는 `mon_L3_XX` 아래 node별 `mon_sub_L3_YY`가 추가된다. `mbm_event` mode에서 MON event에 counter가 없으면 `Unassigned`; CTRL_MON은 자체와 모든 관련 MON 어디에도 할당이 없을 때 `Unassigned`다. `mon_hw_id`는 debug 전용 RMID다.

`mbm_L3_assignments`는 `mbm_event` 지원 시 group의 counter 상태를 `<Event>:<Domain ID>=<state>` 형식으로 표시한다. Domain `*`는 write 때 모든 domain, `_`는 미할당, `e`는 exclusive 할당이다.

cat /sys/fs/resctrl/mbm_L3_assignments
echo "mbm_total_bytes:0=_" > /sys/fs/resctrl/mbm_L3_assignments
echo "mbm_total_bytes:*=_" > /sys/fs/resctrl/mbm_L3_assignments
echo "mbm_total_bytes:*=e" > /sys/fs/resctrl/mbm_L3_assignments

`mba_MBps` 마운트 시 CTRL_MON의 `mba_MBps_event`는 software feedback loop 입력 event를 보여 준다. `mon_features`가 지원하는 bandwidth event 이름을 쓰면 입력을 바꾼다.

Resource alloc and monitor groups
=================================

Resource groups are represented as directories in the resctrl file
system.  The default group is the root directory which, immediately
after mounting, owns all the tasks and cpus in the system and can make
full use of all resources.

On a system with RDT control features additional directories can be
created in the root directory that specify different amounts of each
resource (see "schemata" below). The root and these additional top level
directories are referred to as "CTRL_MON" groups below.

On a system with RDT monitoring the root directory and other top level
directories contain a directory named "mon_groups" in which additional
directories can be created to monitor subsets of tasks in the CTRL_MON
group that is their ancestor. These are called "MON" groups in the rest
of this document.

Removing a directory will move all tasks and cpus owned by the group it
represents to the parent. Removing one of the created CTRL_MON groups
will automatically remove all MON groups below it.

Moving MON group directories to a new parent CTRL_MON group is supported
for the purpose of changing the resource allocations of a MON group
without impacting its monitoring data or assigned tasks. This operation
is not allowed for MON groups which monitor CPUs. No other move
operation is currently allowed other than simply renaming a CTRL_MON or
MON group.

All groups contain the following files:

"tasks":
        Reading this file shows the list of all tasks that belong to
        this group. Writing a task id to the file will add a task to the
        group. Multiple tasks can be added by separating the task ids
        with commas. Tasks will be assigned sequentially. Multiple
        failures are not supported. A single failure encountered while
        attempting to assign a task will cause the operation to abort and
        already added tasks before the failure will remain in the group.
        Failures will be logged to /sys/fs/resctrl/info/last_cmd_status.

        If the group is a CTRL_MON group the task is removed from
        whichever previous CTRL_MON group owned the task and also from
        any MON group that owned the task. If the group is a MON group,
        then the task must already belong to the CTRL_MON parent of this
        group. The task is removed from any previous MON group.


"cpus":
        Reading this file shows a bitmask of the logical CPUs owned by
        this group. Writing a mask to this file will add and remove
        CPUs to/from this group. As with the tasks file a hierarchy is
        maintained where MON groups may only include CPUs owned by the
        parent CTRL_MON group.
        When the resource group is in pseudo-locked mode this file will
        only be readable, reflecting the CPUs associated with the
        pseudo-locked region.


"cpus_list":
        Just like "cpus", only using ranges of CPUs instead of bitmasks.


When control is enabled all CTRL_MON groups will also contain:

"schemata":
        A list of all the resources available to this group.
        Each resource has its own line and format - see below for details.

"size":
        Mirrors the display of the "schemata" file to display the size in
        bytes of each allocation instead of the bits representing the
        allocation.

"mode":
        The "mode" of the resource group dictates the sharing of its
        allocations. A "shareable" resource group allows sharing of its
        allocations while an "exclusive" resource group does not. A
        cache pseudo-locked region is created by first writing
        "pseudo-locksetup" to the "mode" file before writing the cache
        pseudo-locked region's schemata to the resource group's "schemata"
        file. On successful pseudo-locked region creation the mode will
        automatically change to "pseudo-locked".

"ctrl_hw_id":
        Available only with debug option. The identifier used by hardware
        for the control group. On x86 this is the CLOSID.

When monitoring is enabled all MON groups will also contain:

"mon_data":
        This contains a set of files organized by L3 domain and by
        RDT event. E.g. on a system with two L3 domains there will
        be subdirectories "mon_L3_00" and "mon_L3_01".        Each of these
        directories have one file per event (e.g. "llc_occupancy",
        "mbm_total_bytes", and "mbm_local_bytes"). In a MON group these
        files provide a read out of the current value of the event for
        all tasks in the group. In CTRL_MON groups these files provide
        the sum for all tasks in the CTRL_MON group and all tasks in
        MON groups. Please see example section for more details on usage.
        On systems with Sub-NUMA Cluster (SNC) enabled there are extra
        directories for each node (located within the "mon_L3_XX" directory
        for the L3 cache they occupy). These are named "mon_sub_L3_YY"
        where "YY" is the node number.

        When the 'mbm_event' counter assignment mode is enabled, reading
        an MBM event of a MON group returns 'Unassigned' if no hardware
        counter is assigned to it. For CTRL_MON groups, 'Unassigned' is
        returned if the MBM event does not have an assigned counter in the
        CTRL_MON group nor in any of its associated MON groups.

"mon_hw_id":
        Available only with debug option. The identifier used by hardware
        for the monitor group. On x86 this is the RMID.

When monitoring is enabled all MON groups may also contain:

"mbm_L3_assignments":
        Exists when "mbm_event" counter assignment mode is supported and lists the
        counter assignment states of the group.

        The assignment list is displayed in the following format:

        <Event>:<Domain ID>=<Assignment state>;<Domain ID>=<Assignment state>

        Event: A valid MBM event in the
               /sys/fs/resctrl/info/L3_MON/event_configs directory.

        Domain ID: A valid domain ID. When writing, '*' applies the changes
                   to all the domains.

        Assignment states:

        _ : No counter assigned.

        e : Counter assigned exclusively.

        Example:

        To display the counter assignment states for the default group.
        ::

         # cd /sys/fs/resctrl
         # cat /sys/fs/resctrl/mbm_L3_assignments
           mbm_total_bytes:0=e;1=e
           mbm_local_bytes:0=e;1=e

        Assignments can be modified by writing to the interface.

        Examples:

        To unassign the counter associated with the mbm_total_bytes event on domain 0:
        ::

         # echo "mbm_total_bytes:0=_" > /sys/fs/resctrl/mbm_L3_assignments
         # cat /sys/fs/resctrl/mbm_L3_assignments
           mbm_total_bytes:0=_;1=e
           mbm_local_bytes:0=e;1=e

        To unassign the counter associated with the mbm_total_bytes event on all the domains:
        ::

         # echo "mbm_total_bytes:*=_" > /sys/fs/resctrl/mbm_L3_assignments
         # cat /sys/fs/resctrl/mbm_L3_assignments
           mbm_total_bytes:0=_;1=_
           mbm_local_bytes:0=e;1=e

        To assign a counter associated with the mbm_total_bytes event on all domains in
        exclusive mode:
        ::

         # echo "mbm_total_bytes:*=e" > /sys/fs/resctrl/mbm_L3_assignments
         # cat /sys/fs/resctrl/mbm_L3_assignments
           mbm_total_bytes:0=e;1=e
           mbm_local_bytes:0=e;1=e

When the "mba_MBps" mount option is used all CTRL_MON groups will also contain:

"mba_MBps_event":
        Reading this file shows which memory bandwidth event is used
        as input to the software feedback loop that keeps memory bandwidth
        below the value specified in the schemata file. Writing the
        name of one of the supported memory bandwidth events found in
        /sys/fs/resctrl/info/L3_MON/mon_features changes the input
        event.

Allocation·monitoring 선택 규칙

603-630
Task resource allocation 우선순위
순위조건사용 schemata
1Task가 non-default group 소속Task group
2Task는 default지만 실행 CPU가 특정 group 소속CPU group
3그 밖default group

Task 소속이 CPU 소속보다 우선한다.

Task monitoring 귀속 우선순위
순위조건보고 위치
1Task가 MON 또는 non-default CTRL_MON 소속그 group
2Task는 default CTRL_MON이지만 CPU가 특정 group 소속CPU group
3그 밖root `mon_data`

Event가 어느 mon_data에 집계되는지 정한다.

두 규칙 모두 명시적인 task membership을 먼저 보고, default task에 한해 현재 실행 CPU의 group을 보며, 마지막에 root default로 돌아간다.

Resource allocation rules
-------------------------

When a task is running the following rules define which resources are
available to it:

1) If the task is a member of a non-default group, then the schemata
   for that group is used.

2) Else if the task belongs to the default group, but is running on a
   CPU that is assigned to some specific group, then the schemata for the
   CPU's group is used.

3) Otherwise the schemata for the default group is used.

Resource monitoring rules
-------------------------
1) If a task is a member of a MON group, or non-default CTRL_MON group
   then RDT events for the task will be reported in that group.

2) If a task is a member of the default CTRL_MON group, but is running
   on a CPU that is assigned to some specific group, then the RDT events
   for the task will be reported in that group.

3) Otherwise RDT events for the task will be reported in the root level
   "mon_data" group.

Occupancy·RMID limbo·Cache ID·CBM

631-706

Task를 다른 group으로 옮겨도 새 cache allocation에만 영향을 준다. 옛 group에서 3MB occupancy를 보이던 task를 옮긴 직후에는 옛 group이 여전히 3MB, 새 group은 0일 수 있다. 기존 cache line을 다시 접근해도 hardware counter는 갱신되지 않으며, eviction과 새 load가 진행되면서 옛 값은 내려가고 새 값은 올라간다.

Cache allocation control도 마찬가지다. 더 작은 partition으로 옮겨도 기존 line을 강제 퇴출하지 않으므로 process가 옛 partition의 line을 계속 사용할 수 있다.

Hardware는 control group을 CLOSID, monitoring group을 RMID로 식별한다. 개수가 제한돼 CTRL_MON 생성은 CLOSID 또는 RMID 부족으로, MON 생성은 RMID 부족으로 실패할 수 있다.

해제한 RMID가 이전 사용자의 cache line에 아직 tag돼 있으면 즉시 재사용할 수 없다. Limbo 목록에서 occupancy가 내려가는지 확인하며, 재사용 불가능한 limbo RMID가 많으면 `mkdir`가 `-EBUSY`를 반환할 수 있다. `max_threshold_occupancy`가 재사용 가능한 occupancy byte 한계를 정한다.

`mon_llc_occupancy_limbo` tracepoint는 즉시 할당할 수 없는 RMID subset의 정확한 byte occupancy를 제공한다. 매초 출력된다고 보장할 수 없고 빈 monitor group 생성을 시도해 update를 강제해야 할 수 있으며 group 생성 실패 때만 출력될 수도 있다.

Schemata의 각 행은 resource 이름과 시스템의 각 instance에 적용할 값을 담는다.

Cache 공유 단위는 socket·core라는 가정 대신 `Cache ID`로 식별한다. 같은 cache level에서 시스템 전체에 유일하지만 연속 번호라고 보장하지 않는다. Logical CPU별 ID는 `/sys/devices/system/cpu/cpu*/cache/index*/id`에서 찾는다.

Cache allocation은 CBM bitmask로 표현한다. 최대 mask는 CPU model과 cache level마다 다르며 CPUID와 `info/{resource}/cbm_mask`에서 확인한다. 일부 Intel hardware는 1 bit가 연속해야 하므로 4-bit mask에서 `0x3`, `0x6`, `0xC`는 유효하지만 `0x5`, `0x9`, `0xA`는 유효하지 않다. `sparse_masks`로 비연속 지원을 확인한다.

20-bit mask에서는 bit 하나가 cache 용량 5%다. 네 등분 mask는 `0x1f`, `0x3e0`, `0x7c00`, `0xf8000`이다.

RMID 재사용
Group 삭제와 RMID 해제이전 cache line에 RMID tag 잔존RMID를 limbo 목록에 보관LLC occupancy가 threshold 아래인지 재검사조건 충족 뒤 새 group에 재할당

Group 삭제가 곧바로 hardware ID 재사용을 뜻하지 않는다.

Notes on cache occupancy monitoring and control
===============================================
When moving a task from one group to another you should remember that
this only affects *new* cache allocations by the task. E.g. you may have
a task in a monitor group showing 3 MB of cache occupancy. If you move
to a new group and immediately check the occupancy of the old and new
groups you will likely see that the old group is still showing 3 MB and
the new group zero. When the task accesses locations still in cache from
before the move, the h/w does not update any counters. On a busy system
you will likely see the occupancy in the old group go down as cache lines
are evicted and re-used while the occupancy in the new group rises as
the task accesses memory and loads into the cache are counted based on
membership in the new group.

The same applies to cache allocation control. Moving a task to a group
with a smaller cache partition will not evict any cache lines. The
process may continue to use them from the old partition.

Hardware uses CLOSid(Class of service ID) and an RMID(Resource monitoring ID)
to identify a control group and a monitoring group respectively. Each of
the resource groups are mapped to these IDs based on the kind of group. The
number of CLOSid and RMID are limited by the hardware and hence the creation of
a "CTRL_MON" directory may fail if we run out of either CLOSID or RMID
and creation of "MON" group may fail if we run out of RMIDs.

max_threshold_occupancy - generic concepts
------------------------------------------

Note that an RMID once freed may not be immediately available for use as
the RMID is still tagged the cache lines of the previous user of RMID.
Hence such RMIDs are placed on limbo list and checked back if the cache
occupancy has gone down. If there is a time when system has a lot of
limbo RMIDs but which are not ready to be used, user may see an -EBUSY
during mkdir.

max_threshold_occupancy is a user configurable value to determine the
occupancy at which an RMID can be freed.

The mon_llc_occupancy_limbo tracepoint gives the precise occupancy in bytes
for a subset of RMID that are not immediately available for allocation.
This can't be relied on to produce output every second, it may be necessary
to attempt to create an empty monitor group to force an update. Output may
only be produced if creation of a control or monitor group fails.

Schemata files - general concepts
---------------------------------
Each line in the file describes one resource. The line starts with
the name of the resource, followed by specific values to be applied
in each of the instances of that resource on the system.

Cache IDs
---------
On current generation systems there is one L3 cache per socket and L2
caches are generally just shared by the hyperthreads on a core, but this
isn't an architectural requirement. We could have multiple separate L3
caches on a socket, multiple cores could share an L2 cache. So instead
of using "socket" or "core" to define the set of logical cpus sharing
a resource we use a "Cache ID". At a given cache level this will be a
unique number across the whole system (but it isn't guaranteed to be a
contiguous sequence, there may be gaps).  To find the ID for each logical
CPU look in /sys/devices/system/cpu/cpu*/cache/index*/id

Cache Bit Masks (CBM)
---------------------
For cache resources we describe the portion of the cache that is available
for allocation using a bitmask. The maximum value of the mask is defined
by each cpu model (and may be different for different cache levels). It
is found using CPUID, but is also provided in the "info" directory of
the resctrl file system in "info/{resource}/cbm_mask". Some Intel hardware
requires that these masks have all the '1' bits in a contiguous block. So
0x3, 0x6 and 0xC are legal 4-bit masks with two bits set, but 0x5, 0x9
and 0xA are not. Check /sys/fs/resctrl/info/{resource}/sparse_masks
if non-contiguous 1s value is supported. On a system with a 20-bit mask
each bit represents 5% of the capacity of the cache. You could partition
the cache into four equal parts with masks: 0x1f, 0x3e0, 0x7c00, 0xf8000.

SNC와 Memory bandwidth 의미

707-789

SNC mode에서는 같은 L3 cache를 공유하고 NUMA distance도 더 작게 보고될 수 있어 Linux가 일반 NUMA node 사이보다 Sub-NUMA node 사이 task를 더 적극적으로 load balance할 수 있다.

각 `mon_L3_XX` top-level monitoring 파일은 같은 L3를 공유하는 모든 SNC node의 합이다. 특정 node CPU에 task를 bind했다면 `mon_sub_L3_YY`의 `llc_occupancy`, `mbm_total_bytes`, `mbm_local_bytes`에서 node-local 값을 읽는다.

Memory bandwidth allocation과 L3 CBM은 여전히 L3 cache level의 모든 SNC node에 적용된다. 다만 bit 하나가 나타내는 L3 용량은 L3당 SNC node 수로 나뉜다. 100MB·10-bit cache는 보통 bit당 10MB지만 SNC node 둘이면 bit당 5MB다.

MBA 기본 인터페이스는 전체 memory bandwidth의 백분율을 지정한다. 최소값은 `info/MB/min_bandwidth`, granularity는 `info/MB/bandwidth_gran`이며 step은 `min_bw + N * bw_gran`이다.

일부 Intel SKU의 throttling은 core 단위다. 같은 core thread가 서로 다른 값을 요청하면 `thread_throttle_mode`에 따라 낮은 값이 둘 모두에 적용될 수 있다.

MBA는 core별인데 MBM은 package level일 수 있어 제어 효과 해석이 헷갈릴 수 있다. 24-core package에서 core당 L2 외부 bandwidth가 10GBps, L3 외부가 100GBps이면 20 thread가 각각 50%로 5GBps를 써 이미 L3 100GBps를 채운다. 백분율을 높여도 실제 bandwidth는 늘지 않는다.

같은 10%라도 thread 하나는 최대 10GBps, 네 thread는 최대 40GBps를 쓸 수 있다. Group이 더 많은 core를 사용하면 지정 백분율이 같아도 실제 bandwidth가 달라진다.

이를 완화하려고 `mba_MBps`와 software controller `mba_sc`가 MBM counter의 실제 값을 읽어 percentage를 조정해 `actual bandwidth < user specified bandwidth`를 유지한다.

MBA 해석 주의
요인영향
Core별 throttling같은 값도 사용하는 core 수에 따라 총 bandwidth 변화
L3 외부 병목L2 여유가 있어도 percentage 증가 효과 없음
SMT thread낮은 요청이 core 전체에 적용될 수 있음
`mba_sc`MBM feedback으로 MiBps 목표를 추종

백분율은 package 전체 절대 대역폭 한도가 아니다.

Notes on Sub-NUMA Cluster mode
==============================
When SNC mode is enabled, Linux may load balance tasks between Sub-NUMA
nodes much more readily than between regular NUMA nodes since the CPUs
on Sub-NUMA nodes share the same L3 cache and the system may report
the NUMA distance between Sub-NUMA nodes with a lower value than used
for regular NUMA nodes.

The top-level monitoring files in each "mon_L3_XX" directory provide
the sum of data across all SNC nodes sharing an L3 cache instance.
Users who bind tasks to the CPUs of a specific Sub-NUMA node can read
the "llc_occupancy", "mbm_total_bytes", and "mbm_local_bytes" in the
"mon_sub_L3_YY" directories to get node local data.

Memory bandwidth allocation is still performed at the L3 cache
level. I.e. throttling controls are applied to all SNC nodes.

L3 cache allocation bitmaps also apply to all SNC nodes. But note that
the amount of L3 cache represented by each bit is divided by the number
of SNC nodes per L3 cache. E.g. with a 100MB cache on a system with 10-bit
allocation masks each bit normally represents 10MB. With SNC mode enabled
with two SNC nodes per L3 cache, each bit only represents 5MB.

Memory bandwidth Allocation and monitoring
==========================================

For Memory bandwidth resource, by default the user controls the resource
by indicating the percentage of total memory bandwidth.

The minimum bandwidth percentage value for each cpu model is predefined
and can be looked up through "info/MB/min_bandwidth". The bandwidth
granularity that is allocated is also dependent on the cpu model and can
be looked up at "info/MB/bandwidth_gran". The available bandwidth
control steps are: min_bw + N * bw_gran. Intermediate values are rounded
to the next control step available on the hardware.

The bandwidth throttling is a core specific mechanism on some of Intel
SKUs. Using a high bandwidth and a low bandwidth setting on two threads
sharing a core may result in both threads being throttled to use the
low bandwidth (see "thread_throttle_mode").

The fact that Memory bandwidth allocation(MBA) may be a core
specific mechanism where as memory bandwidth monitoring(MBM) is done at
the package level may lead to confusion when users try to apply control
via the MBA and then monitor the bandwidth to see if the controls are
effective. Below are such scenarios:

1. User may *not* see increase in actual bandwidth when percentage
   values are increased:

This can occur when aggregate L2 external bandwidth is more than L3
external bandwidth. Consider an SKL SKU with 24 cores on a package and
where L2 external  is 10GBps (hence aggregate L2 external bandwidth is
240GBps) and L3 external bandwidth is 100GBps. Now a workload with '20
threads, having 50% bandwidth, each consuming 5GBps' consumes the max L3
bandwidth of 100GBps although the percentage value specified is only 50%
<< 100%. Hence increasing the bandwidth percentage will not yield any
more bandwidth. This is because although the L2 external bandwidth still
has capacity, the L3 external bandwidth is fully used. Also note that
this would be dependent on number of cores the benchmark is run on.

2. Same bandwidth percentage may mean different actual bandwidth
   depending on # of threads:

For the same SKU in #1, a 'single thread, with 10% bandwidth' and '4
thread, with 10% bandwidth' can consume up to 10GBps and 40GBps although
they have same percentage bandwidth of 10%. This is simply because as
threads start using more cores in an rdtgroup, the actual bandwidth may
increase or vary although user specified bandwidth percentage is same.

In order to mitigate this and make the interface more user friendly,
resctrl added support for specifying the bandwidth in MiBps as well.  The
kernel underneath would use a software feedback mechanism or a "Software
Controller(mba_sc)" which reads the actual bandwidth using MBM counters
and adjust the memory bandwidth percentages to ensure::

        "actual bandwidth < user specified bandwidth".

By default, the schemata would take the bandwidth percentage values
where as user can switch to the "MBA software controller" mode using
a mount option 'mba_MBps'. The schemata format is specified in the below
sections.

Schemata 형식과 AMD SMBA

790-906
Schemata resource 행
기능형식
L3, CDP off`L3:<cache_id>=<cbm>;...`
L3, CDP on`L3DATA:<cache_id>=<cbm>;...`와 `L3CODE:...`
L2, CDP off`L2:<cache_id>=<cbm>;...`
L2, `cdpl2``L2DATA:...`와 `L2CODE:...`
MBA percentage`MB:<cache_id>=bandwidth;...`
MBA MiBps`MB:<cache_id>=bw_MiBps;...`
AMD SMBA`SMBA:<cache_id>=bandwidth;...`

CDP와 mount mode에 따라 이름과 값 의미가 달라진다.

Memory bandwidth domain은 L3 cache다. AMD SMBA는 CXL.memory만 slow memory 장치로 지원하며 여러 장치가 있으면 모든 slow source를 묶어 전체에 한도를 적용한다. SMBA capability와 실제 slow memory 장치 존재는 독립적이므로 장치가 없으면 설정해도 성능에 영향이 없다.

Schemata를 읽으면 모든 resource의 모든 domain 상태가 나오고, 쓸 때는 바꿀 값만 지정하면 된다. 예를 들어 `L3DATA:2=3c0`만 쓰면 domain 2 data mask만 바뀌고 나머지 L3DATA와 L3CODE는 유지된다.

cat schemata
echo "L3DATA:2=3c0;" > schemata

# AMD: 1/8 GB/s 단위, cache id 1에 2GB/s = 16
echo "MB:1=16" > schemata

# AMD SMBA: cache id 1에 8GB/s = 64
echo "SMBA:1=64" > schemata

AMD schemata의 bandwidth는 1/8GB/s 배수다. 따라서 `MB:1=16`은 cache ID 1에 2GB/s, `SMBA:1=64`는 slow memory에 8GB/s 한도를 뜻한다.

부분 schemata 갱신
현재 schemata 전체 읽기바꿀 resource 이름 선택대상 cache ID와 새 값만 작성Hardware granularity와 mode에 맞춰 적용다시 읽어 전체 domain 상태 확인

지정한 resource·domain만 바꾸고 나머지 상태는 유지한다.

L3 schemata file details (code and data prioritization disabled)
----------------------------------------------------------------
With CDP disabled the L3 schemata format is::

        L3:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...

L3 schemata file details (CDP enabled via mount option to resctrl)
------------------------------------------------------------------
When CDP is enabled L3 control is split into two separate resources
so you can specify independent masks for code and data like this::

        L3DATA:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
        L3CODE:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...

L2 schemata file details
------------------------
CDP is supported at L2 using the 'cdpl2' mount option. The schemata
format is either::

        L2:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...

or

        L2DATA:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...
        L2CODE:<cache_id0>=<cbm>;<cache_id1>=<cbm>;...


Memory bandwidth Allocation (default mode)
------------------------------------------

Memory b/w domain is L3 cache.
::

        MB:<cache_id0>=bandwidth0;<cache_id1>=bandwidth1;...

Memory bandwidth Allocation specified in MiBps
----------------------------------------------

Memory bandwidth domain is L3 cache.
::

        MB:<cache_id0>=bw_MiBps0;<cache_id1>=bw_MiBps1;...

Slow Memory Bandwidth Allocation (SMBA)
---------------------------------------
AMD hardware supports Slow Memory Bandwidth Allocation (SMBA).
CXL.memory is the only supported "slow" memory device. With the
support of SMBA, the hardware enables bandwidth allocation on
the slow memory devices. If there are multiple such devices in
the system, the throttling logic groups all the slow sources
together and applies the limit on them as a whole.

The presence of SMBA (with CXL.memory) is independent of slow memory
devices presence. If there are no such devices on the system, then
configuring SMBA will have no impact on the performance of the system.

The bandwidth domain for slow memory is L3 cache. Its schemata file
is formatted as:
::

        SMBA:<cache_id0>=bandwidth0;<cache_id1>=bandwidth1;...

Reading/writing the schemata file
---------------------------------
Reading the schemata file will show the state of all resources
on all domains. When writing you only need to specify those values
which you wish to change.  E.g.
::

  # cat schemata
  L3DATA:0=fffff;1=fffff;2=fffff;3=fffff
  L3CODE:0=fffff;1=fffff;2=fffff;3=fffff
  # echo "L3DATA:2=3c0;" > schemata
  # cat schemata
  L3DATA:0=fffff;1=fffff;2=3c0;3=fffff
  L3CODE:0=fffff;1=fffff;2=fffff;3=fffff

Reading/writing the schemata file (on AMD systems)
--------------------------------------------------
Reading the schemata file will show the current bandwidth limit on all
domains. The allocated resources are in multiples of one eighth GB/s.
When writing to the file, you need to specify what cache id you wish to
configure the bandwidth limit.

For example, to allocate 2GB/s limit on the first cache id:

::

  # cat schemata
    MB:0=2048;1=2048;2=2048;3=2048
    L3:0=ffff;1=ffff;2=ffff;3=ffff

  # echo "MB:1=16" > schemata
  # cat schemata
    MB:0=2048;1=  16;2=2048;3=2048
    L3:0=ffff;1=ffff;2=ffff;3=ffff

Reading/writing the schemata file (on AMD systems) with SMBA feature
--------------------------------------------------------------------
Reading and writing the schemata file is the same as without SMBA in
above section.

For example, to allocate 8GB/s limit on the first cache id:

::

  # cat schemata
    SMBA:0=2048;1=2048;2=2048;3=2048
      MB:0=2048;1=2048;2=2048;3=2048
      L3:0=ffff;1=ffff;2=ffff;3=ffff

  # echo "SMBA:1=64" > schemata
  # cat schemata
    SMBA:0=2048;1=  64;2=2048;3=2048
      MB:0=2048;1=2048;2=2048;3=2048
      L3:0=ffff;1=ffff;2=ffff;3=ffff

Cache pseudo-locking 원리

907-968

CAT는 application이 채울 수 있는 cache 공간을 제한한다. CPU는 현재 할당 영역 밖에 미리 들어간 데이터도 cache hit이면 읽고 쓸 수 있다. Pseudo-locking은 어느 application도 채울 수 없는 cache 부분에 데이터를 preload하고 이후 cache hit로만 제공해 평균 read latency가 낮은 memory 영역을 userspace에 mapping한다.

사용자가 pseudo-lock할 영역의 schemata와 함께 요청하면 먼저 일치하는 CBM의 새 CAT allocation `CLOSNEW`를 만든다. 이 영역은 현재 어느 CLOS와도 겹치면 안 되고 region이 존재하는 동안 미래에도 겹침을 허용하지 않는다.

Pseudo-locked region 생성
미사용 CBM으로 `CLOSNEW` 생성Cache 영역과 같은 크기의 contiguous memory 할당Cache flush, hardware prefetcher·preemption 비활성`CLOSNEW`를 활성화하고 memory를 touch해 cache 적재이전 CLOS 복구 후 `CLOSNEW` CLOSID 반환Contiguous memory를 character device로 userspace에 노출

Exclusive CBM과 연속 memory를 준비해 cache를 채운 뒤 CLOSID를 반환한다.

Pseudo-locked CBM은 이후 어떤 CAT allocation에도 나타나지 않아 보호된다. 어느 CLOS에서 실행하는 application도 cache hit를 통해 이 memory에 접근할 수 있다.

Pseudo-locking은 cache 잔류 확률을 높일 뿐 배치를 보장하지 않는다. `INVD`, `WBINVD`, `CLFLUSH` 같은 명령은 데이터를 evict할 수 있고 C-state는 cache를 축소하거나 끌 수 있다. Region 생성 때 더 깊은 C-state를 자동 제한한다.

Application은 region이 있는 cache와 연결된 core 또는 그 subset에 affinity를 둬야 한다. 초기 `mmap()` 때 sanity check로 잘못된 affinity를 거부하지만 이후에는 강제하지 않으므로 application이 올바른 core에 계속 머물 책임이 있다.

두 단계 중 관리자는 cache 일부를 전용으로 할당하고 같은 크기 memory를 적재해 character device로 노출한다. Userspace application은 둘째 단계에서 그 device를 `mmap()`한다.

Cache Pseudo-Locking
====================
CAT enables a user to specify the amount of cache space that an
application can fill. Cache pseudo-locking builds on the fact that a
CPU can still read and write data pre-allocated outside its current
allocated area on a cache hit. With cache pseudo-locking, data can be
preloaded into a reserved portion of cache that no application can
fill, and from that point on will only serve cache hits. The cache
pseudo-locked memory is made accessible to user space where an
application can map it into its virtual address space and thus have
a region of memory with reduced average read latency.

The creation of a cache pseudo-locked region is triggered by a request
from the user to do so that is accompanied by a schemata of the region
to be pseudo-locked. The cache pseudo-locked region is created as follows:

- Create a CAT allocation CLOSNEW with a CBM matching the schemata
  from the user of the cache region that will contain the pseudo-locked
  memory. This region must not overlap with any current CAT allocation/CLOS
  on the system and no future overlap with this cache region is allowed
  while the pseudo-locked region exists.
- Create a contiguous region of memory of the same size as the cache
  region.
- Flush the cache, disable hardware prefetchers, disable preemption.
- Make CLOSNEW the active CLOS and touch the allocated memory to load
  it into the cache.
- Set the previous CLOS as active.
- At this point the closid CLOSNEW can be released - the cache
  pseudo-locked region is protected as long as its CBM does not appear in
  any CAT allocation. Even though the cache pseudo-locked region will from
  this point on not appear in any CBM of any CLOS an application running with
  any CLOS will be able to access the memory in the pseudo-locked region since
  the region continues to serve cache hits.
- The contiguous region of memory loaded into the cache is exposed to
  user-space as a character device.

Cache pseudo-locking increases the probability that data will remain
in the cache via carefully configuring the CAT feature and controlling
application behavior. There is no guarantee that data is placed in
cache. Instructions like INVD, WBINVD, CLFLUSH, etc. can still evict
“locked” data from cache. Power management C-states may shrink or
power off cache. Deeper C-states will automatically be restricted on
pseudo-locked region creation.

It is required that an application using a pseudo-locked region runs
with affinity to the cores (or a subset of the cores) associated
with the cache on which the pseudo-locked region resides. A sanity check
within the code will not allow an application to map pseudo-locked memory
unless it runs with affinity to cores associated with the cache on which the
pseudo-locked region resides. The sanity check is only done during the
initial mmap() handling, there is no enforcement afterwards and the
application self needs to ensure it remains affine to the correct cores.

Pseudo-locking is accomplished in two stages:

1) During the first stage the system administrator allocates a portion
   of cache that should be dedicated to pseudo-locking. At this time an
   equivalent portion of memory is allocated, loaded into allocated
   cache portion, and exposed as a character device.
2) During the second stage a user-space application maps (mmap()) the
   pseudo-locked memory into its address space.

Pseudo-locking 생성·debug 인터페이스

969-1028

Pseudo-locked region 생성은 `/sys/fs/resctrl`에 resource group 디렉터리를 만들고, `mode`에 `pseudo-locksetup`을 쓴 다음, `bit_usage`에서 모두 미사용인 bit로 `schemata`를 쓰는 세 단계다.

성공하면 mode가 `pseudo-locked`로 바뀌고 `/dev/pseudo_lock`에 group과 같은 이름의 character device가 생긴다. Userspace는 이 device를 `mmap()`해 region에 접근한다.

`CONFIG_DEBUG_FS`가 활성화되면 `/sys/kernel/debug/resctrl`에 pseudo-lock debug 인터페이스가 기본 제공된다. Kernel은 임의 memory 위치가 cache에 있는지 직접 검사할 방법이 없어 tracing으로 residency를 측정한다.

`pseudo_lock_measure` 값
측정Tracepoint
`1`32-byte stride memory access latency`pseudo_lock_mem_latency`
`2`L2 cache hit/miss`pseudo_lock_l2`
`3`L3 cache hit/miss`pseudo_lock_l3`

Region별 debugfs write-only 파일에 쓴 번호가 측정 종류를 정한다.

Latency test는 hardware prefetcher와 preemption을 끄고 32-byte stride로 region을 순회하며 cache hit/miss의 대체 시각화도 제공한다. L2/L3 측정은 platform의 model-specific precision counter가 있을 때 사용한다.

Region 생성 시 `/sys/kernel/debug/resctrl/<newdir>/pseudo_lock_measure`가 생긴다. 측정 전에 관련 tracepoint를 활성화해야 모든 결과가 tracing infrastructure에 기록된다.

Cache Pseudo-Locking Interface
------------------------------
A pseudo-locked region is created using the resctrl interface as follows:

1) Create a new resource group by creating a new directory in /sys/fs/resctrl.
2) Change the new resource group's mode to "pseudo-locksetup" by writing
   "pseudo-locksetup" to the "mode" file.
3) Write the schemata of the pseudo-locked region to the "schemata" file. All
   bits within the schemata should be "unused" according to the "bit_usage"
   file.

On successful pseudo-locked region creation the "mode" file will contain
"pseudo-locked" and a new character device with the same name as the resource
group will exist in /dev/pseudo_lock. This character device can be mmap()'ed
by user space in order to obtain access to the pseudo-locked memory region.

An example of cache pseudo-locked region creation and usage can be found below.

Cache Pseudo-Locking Debugging Interface
----------------------------------------
The pseudo-locking debugging interface is enabled by default (if
CONFIG_DEBUG_FS is enabled) and can be found in /sys/kernel/debug/resctrl.

There is no explicit way for the kernel to test if a provided memory
location is present in the cache. The pseudo-locking debugging interface uses
the tracing infrastructure to provide two ways to measure cache residency of
the pseudo-locked region:

1) Memory access latency using the pseudo_lock_mem_latency tracepoint. Data
   from these measurements are best visualized using a hist trigger (see
   example below). In this test the pseudo-locked region is traversed at
   a stride of 32 bytes while hardware prefetchers and preemption
   are disabled. This also provides a substitute visualization of cache
   hits and misses.
2) Cache hit and miss measurements using model specific precision counters if
   available. Depending on the levels of cache on the system the pseudo_lock_l2
   and pseudo_lock_l3 tracepoints are available.

When a pseudo-locked region is created a new debugfs directory is created for
it in debugfs as /sys/kernel/debug/resctrl/<newdir>. A single
write-only file, pseudo_lock_measure, is present in this directory. The
measurement of the pseudo-locked region depends on the number written to this
debugfs file:

1:
     writing "1" to the pseudo_lock_measure file will trigger the latency
     measurement captured in the pseudo_lock_mem_latency tracepoint. See
     example below.
2:
     writing "2" to the pseudo_lock_measure file will trigger the L2 cache
     residency (cache hits and misses) measurement captured in the
     pseudo_lock_l2 tracepoint. See example below.
3:
     writing "3" to the pseudo_lock_measure file will trigger the L3 cache
     residency (cache hits and misses) measurement captured in the
     pseudo_lock_l3 tracepoint.

All measurements are recorded with the tracing infrastructure. This requires
the relevant tracepoints to be enabled before the measurement is triggered.

Pseudo-lock latency·hit/miss 예제

1029-1087

`newlock` region의 latency는 trace를 비우고 `pseudo_lock_mem_latency`에 `hist:keys=latency` trigger를 설정한 뒤 event를 활성화하고 `pseudo_lock_measure`에 1을 쓰는 순서로 측정한다. 이후 event를 끄고 histogram을 읽는다.

:> /sys/kernel/tracing/trace
echo 'hist:keys=latency' > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/trigger
echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
echo 1 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
cat /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/hist

예시 histogram은 총 8192 hit를 9개 latency bucket으로 나누며 38 cycle 3484회, 40 cycle 3204회가 대부분이고 456 cycle 1회 같은 긴 지연도 보여 준다.

L2 cache hit/miss는 `pseudo_lock_l2` tracepoint를 활성화하고 measure 파일에 2를 쓴 뒤 trace를 읽는다. 예시는 `hits=4097 miss=0`을 기록한다. L3에서는 대응 tracepoint와 값 3을 사용한다.

:> /sys/kernel/tracing/trace
echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
echo 2 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
cat /sys/kernel/tracing/trace
Pseudo-lock 측정 순서
Trace buffer 초기화필요하면 histogram trigger 설정대상 tracepoint 활성화`pseudo_lock_measure`에 1·2·3 writeTracepoint 비활성화 후 hist 또는 trace 읽기

Tracepoint를 먼저 켜고 측정을 trigger한 뒤 즉시 꺼 결과를 고립한다.

Example of latency debugging interface
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
In this example a pseudo-locked region named "newlock" was created. Here is
how we can measure the latency in cycles of reading from this region and
visualize this data with a histogram that is available if CONFIG_HIST_TRIGGERS
is set::

  # :> /sys/kernel/tracing/trace
  # echo 'hist:keys=latency' > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/trigger
  # echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
  # echo 1 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
  # echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/enable
  # cat /sys/kernel/tracing/events/resctrl/pseudo_lock_mem_latency/hist

  # event histogram
  #
  # trigger info: hist:keys=latency:vals=hitcount:sort=hitcount:size=2048 [active]
  #

  { latency:        456 } hitcount:          1
  { latency:         50 } hitcount:         83
  { latency:         36 } hitcount:         96
  { latency:         44 } hitcount:        174
  { latency:         48 } hitcount:        195
  { latency:         46 } hitcount:        262
  { latency:         42 } hitcount:        693
  { latency:         40 } hitcount:       3204
  { latency:         38 } hitcount:       3484

  Totals:
      Hits: 8192
      Entries: 9
    Dropped: 0

Example of cache hits/misses debugging
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
In this example a pseudo-locked region named "newlock" was created on the L2
cache of a platform. Here is how we can obtain details of the cache hits
and misses using the platform's precision counters.
::

  # :> /sys/kernel/tracing/trace
  # echo 1 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
  # echo 2 > /sys/kernel/debug/resctrl/newlock/pseudo_lock_measure
  # echo 0 > /sys/kernel/tracing/events/resctrl/pseudo_lock_l2/enable
  # cat /sys/kernel/tracing/trace

  # tracer: nop
  #
  #                              _-----=> irqs-off
  #                             / _----=> need-resched
  #                            | / _---=> hardirq/softirq
  #                            || / _--=> preempt-depth
  #                            ||| /     delay
  #           TASK-PID   CPU#  ||||    TIMESTAMP  FUNCTION
  #              | |       |   ||||       |         |
  pseudo_lock_mea-1672  [002] ....  3132.860500: pseudo_lock_l2: hits=4097 miss=0

RDT allocation 예제

1088-1288

예제 1은 L3 mask가 4-bit인 2-socket 시스템에서 `p0`와 `p1`을 만든다. `p0`은 cache ID 0의 하위 50% mask `3`, ID 1의 상위 50% mask `c`; `p1`은 둘 다 하위 50% `3`을 사용한다. 두 group의 MBA는 socket마다 50%다.

mkdir p0 p1
echo "L3:0=3;1=c\nMB:0=50;1=50" > p0/schemata
echo "L3:0=3;1=3\nMB:0=50;1=50" > p1/schemata

Default group은 `L3:0=f;1=f`로 전체 cache를 유지한다. Memory bandwidth mask는 cache mask처럼 overlap 위치를 지정하지 않고 group이 사용할 수 있는 최대치만 정한다. `mba_sc`에서는 percentage 대신 socket 0에 1024MB, socket 1에 500MB처럼 절대값을 쓴다.

예제 2는 20-bit mask인 2-socket 시스템에서 PID 1234와 5678 실시간 task에 socket 0 L3의 각 25%를 독점적으로 준다. 먼저 default를 `L3:0=3ff;1=fffff`, `MB:0=50;1=100`으로 줄인다.

`p0`에는 `f8000`, `p1`에는 `7c00`을 주고 task ID를 `tasks`에 쓴 뒤 `taskset -cp`로 전용 CPU에 고정한다. MBA를 함께 쓰면 각 task group에 socket 0의 20%를 요청한다.

예제 3은 single-socket에서 core 4-7의 실시간 task와 core 0-3의 일반 workload를 분리한다. Task별 연결 대신 `p0/cpus`에 mask `F0`을 써 kernel과 해당 CPU의 task 모두 L3 상위 50% `ffc00`과 bandwidth 50%를 공유하게 한다.

예제 4는 8-bit L2 instance 둘에서 각 25%를 쓰는 exclusive group을 만든다. Default가 `ff` 전체를 쓰는 상태에서 `p0` mask `03`을 exclusive로 바꾸면 `schemata overlaps`로 실패한다.

echo 'L2:0=0xfc;1=0xfc' > schemata
echo exclusive > p0/mode
cat info/L2/bit_usage
# 0=SSSSSSEE;1=SSSSSSEE

Default를 `fc`로 줄인 뒤 `p0`을 exclusive로 만들면 성공하고 size는 instance마다 262144 byte다. 새 shareable `p1`은 exclusive `03`과 겹치지 않는 `fc`를 상속해 786432 byte를 얻는다. `bit_usage`는 `SSSSSSEE`로 공유·독점 구역을 표시한다. `p1`을 `01`로 바꿔 exclusive 영역과 겹치게 하면 `overlaps with exclusive group`으로 거부한다.

Allocation 예제 핵심
예제핵심
12-socket 4-bit CBM과 50% MBA group
2실시간 task별 20-bit L3 25%와 CPU affinity
3CPU mask `F0`로 kernel 포함 core group 할당
4Default와 겹침을 제거한 뒤 L2 exclusive group 생성

Task·CPU·exclusive mode로 resource 귀속을 구성하는 네 방식이다.

Examples for RDT allocation usage
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

1) Example 1

On a two socket machine (one L3 cache per socket) with just four bits
for cache bit masks, minimum b/w of 10% with a memory bandwidth
granularity of 10%.
::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl
  # mkdir p0 p1
  # echo "L3:0=3;1=c\nMB:0=50;1=50" > /sys/fs/resctrl/p0/schemata
  # echo "L3:0=3;1=3\nMB:0=50;1=50" > /sys/fs/resctrl/p1/schemata

The default resource group is unmodified, so we have access to all parts
of all caches (its schemata file reads "L3:0=f;1=f").

Tasks that are under the control of group "p0" may only allocate from the
"lower" 50% on cache ID 0, and the "upper" 50% of cache ID 1.
Tasks in group "p1" use the "lower" 50% of cache on both sockets.

Similarly, tasks that are under the control of group "p0" may use a
maximum memory b/w of 50% on socket0 and 50% on socket 1.
Tasks in group "p1" may also use 50% memory b/w on both sockets.
Note that unlike cache masks, memory b/w cannot specify whether these
allocations can overlap or not. The allocations specifies the maximum
b/w that the group may be able to use and the system admin can configure
the b/w accordingly.

If resctrl is using the software controller (mba_sc) then user can enter the
max b/w in MB rather than the percentage values.
::

  # echo "L3:0=3;1=c\nMB:0=1024;1=500" > /sys/fs/resctrl/p0/schemata
  # echo "L3:0=3;1=3\nMB:0=1024;1=500" > /sys/fs/resctrl/p1/schemata

In the above example the tasks in "p1" and "p0" on socket 0 would use a max b/w
of 1024MB where as on socket 1 they would use 500MB.

2) Example 2

Again two sockets, but this time with a more realistic 20-bit mask.

Two real time tasks pid=1234 running on processor 0 and pid=5678 running on
processor 1 on socket 0 on a 2-socket and dual core machine. To avoid noisy
neighbors, each of the two real-time tasks exclusively occupies one quarter
of L3 cache on socket 0.
::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl

First we reset the schemata for the default group so that the "upper"
50% of the L3 cache on socket 0 and 50% of memory b/w cannot be used by
ordinary tasks::

  # echo "L3:0=3ff;1=fffff\nMB:0=50;1=100" > schemata

Next we make a resource group for our first real time task and give
it access to the "top" 25% of the cache on socket 0.
::

  # mkdir p0
  # echo "L3:0=f8000;1=fffff" > p0/schemata

Finally we move our first real time task into this resource group. We
also use taskset(1) to ensure the task always runs on a dedicated CPU
on socket 0. Most uses of resource groups will also constrain which
processors tasks run on.
::

  # echo 1234 > p0/tasks
  # taskset -cp 1 1234

Ditto for the second real time task (with the remaining 25% of cache)::

  # mkdir p1
  # echo "L3:0=7c00;1=fffff" > p1/schemata
  # echo 5678 > p1/tasks
  # taskset -cp 2 5678

For the same 2 socket system with memory b/w resource and CAT L3 the
schemata would look like(Assume min_bandwidth 10 and bandwidth_gran is
10):

For our first real time task this would request 20% memory b/w on socket 0.
::

  # echo -e "L3:0=f8000;1=fffff\nMB:0=20;1=100" > p0/schemata

For our second real time task this would request an other 20% memory b/w
on socket 0.
::

  # echo -e "L3:0=f8000;1=fffff\nMB:0=20;1=100" > p0/schemata

3) Example 3

A single socket system which has real-time tasks running on core 4-7 and
non real-time workload assigned to core 0-3. The real-time tasks share text
and data, so a per task association is not required and due to interaction
with the kernel it's desired that the kernel on these cores shares L3 with
the tasks.
::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl

First we reset the schemata for the default group so that the "upper"
50% of the L3 cache on socket 0, and 50% of memory bandwidth on socket 0
cannot be used by ordinary tasks::

  # echo "L3:0=3ff\nMB:0=50" > schemata

Next we make a resource group for our real time cores and give it access
to the "top" 50% of the cache on socket 0 and 50% of memory bandwidth on
socket 0.
::

  # mkdir p0
  # echo "L3:0=ffc00\nMB:0=50" > p0/schemata

Finally we move core 4-7 over to the new group and make sure that the
kernel and the tasks running there get 50% of the cache. They should
also get 50% of memory bandwidth assuming that the cores 4-7 are SMT
siblings and only the real time threads are scheduled on the cores 4-7.
::

  # echo F0 > p0/cpus

4) Example 4

The resource groups in previous examples were all in the default "shareable"
mode allowing sharing of their cache allocations. If one resource group
configures a cache allocation then nothing prevents another resource group
to overlap with that allocation.

In this example a new exclusive resource group will be created on a L2 CAT
system with two L2 cache instances that can be configured with an 8-bit
capacity bitmask. The new exclusive resource group will be configured to use
25% of each cache instance.
::

  # mount -t resctrl resctrl /sys/fs/resctrl/
  # cd /sys/fs/resctrl

First, we observe that the default group is configured to allocate to all L2
cache::

  # cat schemata
  L2:0=ff;1=ff

We could attempt to create the new resource group at this point, but it will
fail because of the overlap with the schemata of the default group::

  # mkdir p0
  # echo 'L2:0=0x3;1=0x3' > p0/schemata
  # cat p0/mode
  shareable
  # echo exclusive > p0/mode
  -sh: echo: write error: Invalid argument
  # cat info/last_cmd_status
  schemata overlaps

To ensure that there is no overlap with another resource group the default
resource group's schemata has to change, making it possible for the new
resource group to become exclusive.
::

  # echo 'L2:0=0xfc;1=0xfc' > schemata
  # echo exclusive > p0/mode
  # grep . p0/*
  p0/cpus:0
  p0/mode:exclusive
  p0/schemata:L2:0=03;1=03
  p0/size:L2:0=262144;1=262144

A new resource group will on creation not overlap with an exclusive resource
group::

  # mkdir p1
  # grep . p1/*
  p1/cpus:0
  p1/mode:shareable
  p1/schemata:L2:0=fc;1=fc
  p1/size:L2:0=786432;1=786432

The bit_usage will reflect how the cache is used::

  # cat info/L2/bit_usage
  0=SSSSSSEE;1=SSSSSSEE

A resource group cannot be forced to overlap with an exclusive resource group::

  # echo 'L2:0=0x1;1=0x1' > p1/schemata
  -sh: echo: write error: Invalid argument
  # cat info/last_cmd_status
  overlaps with exclusive group

Cache pseudo-locking 전체 예제

1289-1393

예제는 cache ID 1의 L2에서 CBM `0x3`을 잠그고 `/dev/pseudo_lock/newlock`으로 노출한다. 먼저 `bit_usage`가 모두 `S`인지 확인하고 default schemata를 `L2:1=0xfc`로 줄여 하위 두 bit를 `0`으로 비운다.

cat info/L2/bit_usage
echo 'L2:1=0xfc' > schemata
mkdir newlock
echo pseudo-locksetup > newlock/mode
echo 'L2:1=0x3' > newlock/schemata

성공하면 `newlock/mode`는 `pseudo-locked`, cache ID 1 bit usage는 `SSSSSSPP`가 되며 character device가 생긴다.

Userspace 예제는 `_GNU_SOURCE`와 `sched.h`, `sys/mman.h` 등을 사용한다. Hard-coded CPU 2를 `CPU_SET`과 `sched_setaffinity()`로 지정해 region cache에 연결된 core에 고정한다.

그 뒤 `/dev/pseudo_lock/newlock`을 `O_RDWR`로 열고 system page size 한 page를 `PROT_READ | PROT_WRITE`, `MAP_SHARED`로 mmap한다. Application은 반환된 `mapping`으로 pseudo-locked memory를 사용하고 `munmap()`과 `close()`로 정리한다.

각 system call 실패는 `perror()` 뒤 `EXIT_FAILURE`로 종료하며, 정상 종료는 mapping 해제와 device close 뒤 `EXIT_SUCCESS`다.

Pseudo-lock 예제 수명
Default CBM에서 `0x3` bit 제거`newlock` group과 `pseudo-locksetup` mode 생성`L2:1=0x3` schemata로 region 완성`SSSSSSPP`와 character device 확인Application CPU affinity 설정Device open·한 page mmap·사용·munmap

Cache way 예약과 application mapping을 순서대로 검증한다.

Example of Cache Pseudo-Locking
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Lock portion of L2 cache from cache id 1 using CBM 0x3. Pseudo-locked
region is exposed at /dev/pseudo_lock/newlock that can be provided to
application for argument to mmap().
::

  # mount -t resctrl resctrl /sys/fs/resctrl/
  # cd /sys/fs/resctrl

Ensure that there are bits available that can be pseudo-locked, since only
unused bits can be pseudo-locked the bits to be pseudo-locked needs to be
removed from the default resource group's schemata::

  # cat info/L2/bit_usage
  0=SSSSSSSS;1=SSSSSSSS
  # echo 'L2:1=0xfc' > schemata
  # cat info/L2/bit_usage
  0=SSSSSSSS;1=SSSSSS00

Create a new resource group that will be associated with the pseudo-locked
region, indicate that it will be used for a pseudo-locked region, and
configure the requested pseudo-locked region capacity bitmask::

  # mkdir newlock
  # echo pseudo-locksetup > newlock/mode
  # echo 'L2:1=0x3' > newlock/schemata

On success the resource group's mode will change to pseudo-locked, the
bit_usage will reflect the pseudo-locked region, and the character device
exposing the pseudo-locked region will exist::

  # cat newlock/mode
  pseudo-locked
  # cat info/L2/bit_usage
  0=SSSSSSSS;1=SSSSSSPP
  # ls -l /dev/pseudo_lock/newlock
  crw------- 1 root root 243, 0 Apr  3 05:01 /dev/pseudo_lock/newlock

::

  /*
  * Example code to access one page of pseudo-locked cache region
  * from user space.
  */
  #define _GNU_SOURCE
  #include <fcntl.h>
  #include <sched.h>
  #include <stdio.h>
  #include <stdlib.h>
  #include <unistd.h>
  #include <sys/mman.h>

  /*
  * It is required that the application runs with affinity to only
  * cores associated with the pseudo-locked region. Here the cpu
  * is hardcoded for convenience of example.
  */
  static int cpuid = 2;

  int main(int argc, char *argv[])
  {
    cpu_set_t cpuset;
    long page_size;
    void *mapping;
    int dev_fd;
    int ret;

    page_size = sysconf(_SC_PAGESIZE);

    CPU_ZERO(&cpuset);
    CPU_SET(cpuid, &cpuset);
    ret = sched_setaffinity(0, sizeof(cpuset), &cpuset);
    if (ret < 0) {
      perror("sched_setaffinity");
      exit(EXIT_FAILURE);
    }

    dev_fd = open("/dev/pseudo_lock/newlock", O_RDWR);
    if (dev_fd < 0) {
      perror("open");
      exit(EXIT_FAILURE);
    }

    mapping = mmap(0, page_size, PROT_READ | PROT_WRITE, MAP_SHARED,
            dev_fd, 0);
    if (mapping == MAP_FAILED) {
      perror("mmap");
      close(dev_fd);
      exit(EXIT_FAILURE);
    }

    /* Application interacts with pseudo-locked memory @mapping */

    ret = munmap(mapping, page_size);
    if (ret < 0) {
      perror("munmap");
      close(dev_fd);
      exit(EXIT_FAILURE);
    }

    close(dev_fd);
    exit(EXIT_SUCCESS);
  }

애플리케이션 사이 advisory locking

1394-1508

여러 resctrl 파일의 read/write로 구성된 작업은 원자적이어야 한다. 예를 들어 exclusive L3 예약은 모든 CBM 또는 `bit_usage`를 읽고, 어디에도 쓰이지 않는 연속 bit를 찾고, 새 디렉터리를 만들고, 그 bit를 새 `schemata`에 쓰는 네 단계다.

두 application이 동시에 실행하면 같은 bit를 골라 exclusive가 아니라 공유 예약을 만들 수 있다. 이를 막으려면 libc와 shell에서 사용할 수 있는 `flock`으로 `/sys/fs/resctrl` 자체를 잠근다.

권장 resctrl lock 절차
작업Lock순서
Write`flock(LOCK_EX)`exclusive lock → 구조 read/write → `LOCK_UN`
Read`flock(LOCK_SH)`shared lock → 성공 시 구조 read → `LOCK_UN`

디렉터리 구조 전체를 advisory lock 대상으로 사용한다.

flock -s /sys/fs/resctrl/ find /sys/fs/resctrl
flock /sys/fs/resctrl/ ./create-dir.sh

Shell 예제는 shared lock으로 원자적 구조 snapshot을 읽고, exclusive lock 아래 script에서 mask 계산·directory 생성·schemata write를 수행한다.

C 예제는 `/sys/fs/resctrl`을 `O_DIRECTORY`로 열고 `flock(fd, LOCK_SH)`, `flock(fd, LOCK_EX)`, `flock(fd, LOCK_UN)`을 각각 shared 획득, exclusive 획득, 해제 helper로 감싼다. Read-only 작업과 read/write 작업에 맞는 lock을 선택한다.

이는 kernel이 강제하는 transaction lock이 아니라 협력 application 사이의 advisory protocol이므로 모든 관리 도구가 같은 규약을 지켜야 한다.

Exclusive cache 예약 transaction
`LOCK_EX` 획득모든 group CBM 또는 `bit_usage` 읽기전역 미사용 연속 bit 계산새 group 생성새 `schemata`에 mask write`LOCK_UN` 해제

한 lock 범위 안에서 관찰과 갱신을 묶어 TOCTOU 충돌을 막는다.

Locking between applications
----------------------------

Certain operations on the resctrl filesystem, composed of read/writes
to/from multiple files, must be atomic.

As an example, the allocation of an exclusive reservation of L3 cache
involves:

  1. Read the cbmmasks from each directory or the per-resource "bit_usage"
  2. Find a contiguous set of bits in the global CBM bitmask that is clear
     in any of the directory cbmmasks
  3. Create a new directory
  4. Set the bits found in step 2 to the new directory "schemata" file

If two applications attempt to allocate space concurrently then they can
end up allocating the same bits so the reservations are shared instead of
exclusive.

To coordinate atomic operations on the resctrlfs and to avoid the problem
above, the following locking procedure is recommended:

Locking is based on flock, which is available in libc and also as a shell
script command

Write lock:

 A) Take flock(LOCK_EX) on /sys/fs/resctrl
 B) Read/write the directory structure.
 C) funlock

Read lock:

 A) Take flock(LOCK_SH) on /sys/fs/resctrl
 B) If success read the directory structure.
 C) funlock

Example with bash::

  # Atomically read directory structure
  $ flock -s /sys/fs/resctrl/ find /sys/fs/resctrl

  # Read directory contents and create new subdirectory

  $ cat create-dir.sh
  find /sys/fs/resctrl/ > output.txt
  mask = function-of(output.txt)
  mkdir /sys/fs/resctrl/newres/
  echo mask > /sys/fs/resctrl/newres/schemata

  $ flock /sys/fs/resctrl/ ./create-dir.sh

Example with C::

  /*
  * Example code do take advisory locks
  * before accessing resctrl filesystem
  */
  #include <sys/file.h>
  #include <stdlib.h>

  void resctrl_take_shared_lock(int fd)
  {
    int ret;

    /* take shared lock on resctrl filesystem */
    ret = flock(fd, LOCK_SH);
    if (ret) {
      perror("flock");
      exit(-1);
    }
  }

  void resctrl_take_exclusive_lock(int fd)
  {
    int ret;

    /* release lock on resctrl filesystem */
    ret = flock(fd, LOCK_EX);
    if (ret) {
      perror("flock");
      exit(-1);
    }
  }

  void resctrl_release_lock(int fd)
  {
    int ret;

    /* take shared lock on resctrl filesystem */
    ret = flock(fd, LOCK_UN);
    if (ret) {
      perror("flock");
      exit(-1);
    }
  }

  void main(void)
  {
    int fd, ret;

    fd = open("/sys/fs/resctrl", O_DIRECTORY);
    if (fd == -1) {
      perror("open");
      exit(-1);
    }
    resctrl_take_shared_lock(fd);
    /* code to read directory contents */
    resctrl_release_lock(fd);

    resctrl_take_exclusive_lock(fd);
    /* code to read and write directory contents */
    resctrl_release_lock(fd);
  }

RDT monitoring 예제

1509-1638

Event 파일, 예를 들어 `mon_data/mon_L3_00/llc_occupancy`를 읽으면 해당 MON 또는 CTRL_MON group의 현재 LLC occupancy snapshot을 byte로 얻는다.

예제 1은 2-socket 4-bit CBM에서 `p0`과 `p1`을 만들고 PID 5678·5679를 `p1`에 넣는다. `p1/mon_groups` 아래 `m11`, `m12`를 만들어 각 task를 분리하면 domain별 occupancy를 따로 읽을 수 있다. 부모 `p1/mon_data`는 두 MON을 포함한 합계 31234000을 보여 준다.

mkdir p1/mon_groups/m11 p1/mon_groups/m12
echo 5678 > p1/mon_groups/m11/tasks
echo 5679 > p1/mon_groups/m12/tasks
cat p1/mon_groups/m11/mon_data/mon_L3_00/llc_occupancy
cat p1/mon_data/mon_L3_00/llc_occupancy

예제 2는 group 생성 시 RMID가 할당된다는 점을 이용한다. 현재 shell PID `$$`를 `p1/tasks`에 넣은 뒤 `<cmd>`를 실행하면 command가 생성되는 순간부터 관찰된다.

예제 3은 CAT 없이 CQM만 있는 HSW 같은 시스템에서도 resctrl이 마운트됨을 보여 준다. CTRL_MON은 만들 수 없지만 root `mon_groups` 아래 `m01`, `m02`를 만들어 kernel thread를 포함한 task를 관찰하고 allocation 전 cache footprint를 profiling할 수 있다. 예시 수치는 workload가 주로 domain 0에서 동작함을 보여 준다.

예제 4는 single-socket의 실시간 task가 실행되는 CPU 4-7을 `p1/cpus`에 mask `f0`으로 넣고 `p1/mon_data/mon_L3_00/llc_occupancy`에서 이 CPU group의 occupancy를 읽는다.

Monitoring group 사용 패턴
예제패턴
1CTRL_MON 아래 MON으로 task subset 분리, 부모는 합계
2Shell을 group에 넣고 자식 command를 생성부터 관찰
3CAT 없이 root MON만으로 footprint profiling
4CPU mask 기반 실시간 task monitoring

Task subset, 생성 시점, monitoring-only 시스템, CPU group을 다룬다.

Examples for RDT Monitoring along with allocation usage
=======================================================
Reading monitored data
----------------------
Reading an event file (for ex: mon_data/mon_L3_00/llc_occupancy) would
show the current snapshot of LLC occupancy of the corresponding MON
group or CTRL_MON group.


Example 1 (Monitor CTRL_MON group and subset of tasks in CTRL_MON group)
------------------------------------------------------------------------
On a two socket machine (one L3 cache per socket) with just four bits
for cache bit masks::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl
  # mkdir p0 p1
  # echo "L3:0=3;1=c" > /sys/fs/resctrl/p0/schemata
  # echo "L3:0=3;1=3" > /sys/fs/resctrl/p1/schemata
  # echo 5678 > p1/tasks
  # echo 5679 > p1/tasks

The default resource group is unmodified, so we have access to all parts
of all caches (its schemata file reads "L3:0=f;1=f").

Tasks that are under the control of group "p0" may only allocate from the
"lower" 50% on cache ID 0, and the "upper" 50% of cache ID 1.
Tasks in group "p1" use the "lower" 50% of cache on both sockets.

Create monitor groups and assign a subset of tasks to each monitor group.
::

  # cd /sys/fs/resctrl/p1/mon_groups
  # mkdir m11 m12
  # echo 5678 > m11/tasks
  # echo 5679 > m12/tasks

fetch data (data shown in bytes)
::

  # cat m11/mon_data/mon_L3_00/llc_occupancy
  16234000
  # cat m11/mon_data/mon_L3_01/llc_occupancy
  14789000
  # cat m12/mon_data/mon_L3_00/llc_occupancy
  16789000

The parent ctrl_mon group shows the aggregated data.
::

  # cat /sys/fs/resctrl/p1/mon_data/mon_l3_00/llc_occupancy
  31234000

Example 2 (Monitor a task from its creation)
--------------------------------------------
On a two socket machine (one L3 cache per socket)::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl
  # mkdir p0 p1

An RMID is allocated to the group once its created and hence the <cmd>
below is monitored from its creation.
::

  # echo $$ > /sys/fs/resctrl/p1/tasks
  # <cmd>

Fetch the data::

  # cat /sys/fs/resctrl/p1/mon_data/mon_l3_00/llc_occupancy
  31789000

Example 3 (Monitor without CAT support or before creating CAT groups)
---------------------------------------------------------------------

Assume a system like HSW has only CQM and no CAT support. In this case
the resctrl will still mount but cannot create CTRL_MON directories.
But user can create different MON groups within the root group thereby
able to monitor all tasks including kernel threads.

This can also be used to profile jobs cache size footprint before being
able to allocate them to different allocation groups.
::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl
  # mkdir mon_groups/m01
  # mkdir mon_groups/m02

  # echo 3478 > /sys/fs/resctrl/mon_groups/m01/tasks
  # echo 2467 > /sys/fs/resctrl/mon_groups/m02/tasks

Monitor the groups separately and also get per domain data. From the
below its apparent that the tasks are mostly doing work on
domain(socket) 0.
::

  # cat /sys/fs/resctrl/mon_groups/m01/mon_L3_00/llc_occupancy
  31234000
  # cat /sys/fs/resctrl/mon_groups/m01/mon_L3_01/llc_occupancy
  34555
  # cat /sys/fs/resctrl/mon_groups/m02/mon_L3_00/llc_occupancy
  31234000
  # cat /sys/fs/resctrl/mon_groups/m02/mon_L3_01/llc_occupancy
  32789


Example 4 (Monitor real time tasks)
-----------------------------------

A single socket system which has real time tasks running on cores 4-7
and non real time tasks on other cpus. We want to monitor the cache
occupancy of the real time threads on these cores.
::

  # mount -t resctrl resctrl /sys/fs/resctrl
  # cd /sys/fs/resctrl
  # mkdir p1

Move the cpus 4-7 over to p1::

  # echo f0 > p1/cpus

View the llc occupancy snapshot::

  # cat /sys/fs/resctrl/p1/mon_data/mon_L3_00/llc_occupancy
  11234000

`mbm_assign_mode` 운용 예제

1639-1756

먼저 `info/L3_MON/mbm_assign_mode`에서 `[mbm_event]`가 표시되는지 확인하고, `num_mbm_cntrs`와 `available_mbm_cntrs`에서 domain별 최대·가용 counter를 읽는다. 예시는 32개 중 30개가 할당 가능하다.

Root group의 `mbm_L3_assignments`는 `mbm_total_bytes:0=e;1=e`처럼 event와 domain별 exclusive 할당을 보여 준다. `0=_`는 domain 0 해제, `*=_`는 모든 domain 해제, `*=e`는 모든 domain exclusive 할당이다.

할당 여부와 관계없이 event file 읽는 방법은 같다. 예제는 두 domain의 total과 local byte 값을 읽는다.

각 event의 `event_filter`에서 transaction 구성을 확인하고 `mbm_local_bytes` filter에 `remote_reads` 같은 항목을 추가할 수 있다. 구성을 바꾼 직후 첫 read는 counter reset 때문에 `Unavailable`, 다음 read는 현재 값을 반환할 수 있다.

필요하면 `mbm_assign_mode`에 `default`를 써 돌아간다. Mode 전환은 모든 resctrl group의 MBM counter와 event를 reset할 수 있다. 마지막에는 resctrl을 언마운트한다.

cat info/L3_MON/mbm_assign_mode
cat info/L3_MON/num_mbm_cntrs
cat info/L3_MON/available_mbm_cntrs
echo "mbm_total_bytes:0=_" > mbm_L3_assignments
echo "mbm_total_bytes:*=e" > mbm_L3_assignments
echo "default" > info/L3_MON/mbm_assign_mode
umount /sys/fs/resctrl/
MBM counter 관리
`mbm_event` 지원·활성 확인최대·가용 counter 수 확인Event·domain별 `_` 또는 `e` writeEvent filter 확인·변경첫 `Unavailable` 뒤 재읽기필요 시 `default` 복귀 후 unmount

Mode 확인부터 filter 변경·reset 처리·복귀까지의 운용 순서다.

Examples on working with mbm_assign_mode
========================================

a. Check if MBM counter assignment mode is supported.
::

  # mount -t resctrl resctrl /sys/fs/resctrl/

  # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
  [mbm_event]
  default

The "mbm_event" mode is detected and enabled.

b. Check how many assignable counters are supported.
::

  # cat /sys/fs/resctrl/info/L3_MON/num_mbm_cntrs
  0=32;1=32

c. Check how many assignable counters are available for assignment in each domain.
::

  # cat /sys/fs/resctrl/info/L3_MON/available_mbm_cntrs
  0=30;1=30

d. To list the default group's assign states.
::

  # cat /sys/fs/resctrl/mbm_L3_assignments
  mbm_total_bytes:0=e;1=e
  mbm_local_bytes:0=e;1=e

e.  To unassign the counter associated with the mbm_total_bytes event on domain 0.
::

  # echo "mbm_total_bytes:0=_" > /sys/fs/resctrl/mbm_L3_assignments
  # cat /sys/fs/resctrl/mbm_L3_assignments
  mbm_total_bytes:0=_;1=e
  mbm_local_bytes:0=e;1=e

f. To unassign the counter associated with the mbm_total_bytes event on all domains.
::

  # echo "mbm_total_bytes:*=_" > /sys/fs/resctrl/mbm_L3_assignments
  # cat /sys/fs/resctrl/mbm_L3_assignment
  mbm_total_bytes:0=_;1=_
  mbm_local_bytes:0=e;1=e

g. To assign a counter associated with the mbm_total_bytes event on all domains in
exclusive mode.
::

  # echo "mbm_total_bytes:*=e" > /sys/fs/resctrl/mbm_L3_assignments
  # cat /sys/fs/resctrl/mbm_L3_assignments
  mbm_total_bytes:0=e;1=e
  mbm_local_bytes:0=e;1=e

h. Read the events mbm_total_bytes and mbm_local_bytes of the default group. There is
no change in reading the events with the assignment.
::

  # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_total_bytes
  779247936
  # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_total_bytes
  562324232
  # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
  212122123
  # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
  121212144

i. Check the event configurations.
::

  # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_total_bytes/event_filter
  local_reads,remote_reads,local_non_temporal_writes,remote_non_temporal_writes,
  local_reads_slow_memory,remote_reads_slow_memory,dirty_victim_writes_all

  # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
  local_reads,local_non_temporal_writes,local_reads_slow_memory

j. Change the event configuration for mbm_local_bytes.
::

  # echo "local_reads, local_non_temporal_writes, local_reads_slow_memory, remote_reads" >
  /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter

  # cat /sys/fs/resctrl/info/L3_MON/event_configs/mbm_local_bytes/event_filter
  local_reads,local_non_temporal_writes,local_reads_slow_memory,remote_reads

k. Now read the local events again. The first read may come back with "Unavailable"
status. The subsequent read of mbm_local_bytes will display the current value.
::

  # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
  Unavailable
  # cat /sys/fs/resctrl/mon_data/mon_L3_00/mbm_local_bytes
  2252323
  # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
  Unavailable
  # cat /sys/fs/resctrl/mon_data/mon_L3_01/mbm_local_bytes
  1566565

l. Users have the option to go back to 'default' mbm_assign_mode if required. This can be
done using the following command. Note that switching the mbm_assign_mode may reset all
the MBM counters (and thus all MBM events) of all the resctrl groups.
::

  # echo "default" > /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
  # cat /sys/fs/resctrl/info/L3_MON/mbm_assign_mode
  mbm_event
  [default]

m. Unmount the resctrl filesystem.
::

  # umount /sys/fs/resctrl/

Intel MBM counter errata 보정

1757-1848

Skylake server의 SKX99와 Broadwell server의 BDF102 errata 때문에 Intel MBM counter가 특정 RMID에서 system memory bandwidth를 잘못 보고할 수 있다. Logical core에 할당된 RMID 기준 metric을 보고하는 `IA32_QM_CTR` register는 MSR `0xC8E`다.

이 문제로 실제 system bandwidth와 보고값이 일치하지 않을 수 있다. RMID가 표의 threshold보다 크면 MBM total과 local 값에 해당 correction factor를 곱해 보정한다.

Intel MBM correction factor
CoreRMIDThresholdFactor
1801.000000
21601.000000
324150.969650
43201.000000
648310.969650
756471.142857
86401.000000
972631.185115
1080631.066553
1188791.454545
129601.000000
13104951.230769
14112951.142857
15120951.066667
1612801.000000
171361271.254863
181441271.185255
1915201.000000
201601271.066667
2116801.000000
221761591.454334
2318401.000000
241921270.969744
252001911.280246
262081911.230921
2721601.000000
282241911.143118

Core 수에 따른 RMID 개수·threshold·보정 계수 전체 표다.

원문은 Intel Xeon Scalable Family specification update의 SKX99, Xeon E5-2600 v4 update의 BDF102, 2세대 Xeon Scalable용 Intel RDT reference manual 링크를 추가 근거로 제공한다.

MBM errata 적용
대상 Skylake/Broadwell server와 errata 확인Core count로 표 행 선택현재 RMID와 threshold 비교`RMID > threshold`면 total·local 값에 factor 곱하기보정된 system bandwidth 보고

Platform과 RMID 조건을 확인한 뒤 raw counter에 보정 계수를 적용한다.

Intel RDT Errata
================

Intel MBM Counters May Report System Memory Bandwidth Incorrectly
-----------------------------------------------------------------

Errata SKX99 for Skylake server and BDF102 for Broadwell server.

Problem: Intel Memory Bandwidth Monitoring (MBM) counters track metrics
according to the assigned Resource Monitor ID (RMID) for that logical
core. The IA32_QM_CTR register (MSR 0xC8E), used to report these
metrics, may report incorrect system bandwidth for certain RMID values.

Implication: Due to the errata, system memory bandwidth may not match
what is reported.

Workaround: MBM total and local readings are corrected according to the
following correction factor table:

+---------------+---------------+---------------+-----------------+
|core count        |rmid count        |rmid threshold        |correction factor|
+---------------+---------------+---------------+-----------------+
|1                |8                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|2                |16                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|3                |24                |15                |0.969650          |
+---------------+---------------+---------------+-----------------+
|4                |32                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|6                |48                |31                |0.969650          |
+---------------+---------------+---------------+-----------------+
|7                |56                |47                |1.142857          |
+---------------+---------------+---------------+-----------------+
|8                |64                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|9                |72                |63                |1.185115          |
+---------------+---------------+---------------+-----------------+
|10                |80                |63                |1.066553          |
+---------------+---------------+---------------+-----------------+
|11                |88                |79                |1.454545          |
+---------------+---------------+---------------+-----------------+
|12                |96                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|13                |104                |95                |1.230769          |
+---------------+---------------+---------------+-----------------+
|14                |112                |95                |1.142857          |
+---------------+---------------+---------------+-----------------+
|15                |120                |95                |1.066667          |
+---------------+---------------+---------------+-----------------+
|16                |128                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|17                |136                |127                |1.254863          |
+---------------+---------------+---------------+-----------------+
|18                |144                |127                |1.185255          |
+---------------+---------------+---------------+-----------------+
|19                |152                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|20                |160                |127                |1.066667          |
+---------------+---------------+---------------+-----------------+
|21                |168                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|22                |176                |159                |1.454334          |
+---------------+---------------+---------------+-----------------+
|23                |184                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|24                |192                |127                |0.969744          |
+---------------+---------------+---------------+-----------------+
|25                |200                |191                |1.280246          |
+---------------+---------------+---------------+-----------------+
|26                |208                |191                |1.230921          |
+---------------+---------------+---------------+-----------------+
|27                |216                |0                |1.000000          |
+---------------+---------------+---------------+-----------------+
|28                |224                |191                |1.143118          |
+---------------+---------------+---------------+-----------------+

If rmid > rmid threshold, MBM total and local values should be multiplied
by the correction factor.

See:

1. Erratum SKX99 in Intel Xeon Processor Scalable Family Specification Update:
http://web.archive.org/web/20200716124958/https://www.intel.com/content/www/us/en/processors/xeon/scalable/xeon-scalable-spec-update.html

2. Erratum BDF102 in Intel Xeon E5-2600 v4 Processor Product Family Specification Update:
http://web.archive.org/web/20191125200531/https://www.intel.com/content/dam/www/public/us/en/documents/specification-updates/xeon-e5-v4-spec-update.pdf

3. The errata in Intel Resource Director Technology (Intel RDT) on 2nd Generation Intel Xeon Scalable Processors Reference Manual:
https://software.intel.com/content/www/us/en/develop/articles/intel-resource-director-technology-rdt-reference-manual.html

for further information.