← Documents Documentation/filesystems/erofs.rst GitHub 원문 ↗

Linux 6.18.37 · Filesystems

EROFS - Enhanced Read-Only File System

EROFS의 기능, mount option, inode·xattr·directory on-disk layout, chunk와 fixed-output compression을 다룬 전문 번역입니다.

Source pathDocumentation/filesystems/erofs.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

erofs.rst:1-369

EROFS는 compact·extended inode, tail packing, xattr 공유, chunk deduplication, fixed-output compression, direct I/O·FSDAX와 Fscache distribution을 결합한 고성능 read-only filesystem입니다. metadata는 compact하지만 inode·xattr·directory·compression index의 주소식과 alignment를 명시적으로 유지합니다.

EROFS image data path
`mkfs.erofs`가 inode·xattr·directory metadata 구성file별 compression·chunk·tail-packing layout 선택block image 또는 Fscache·external blob으로 배포directory prefix search와 inode offset으로 lookupHEAD/NONHEAD index로 pcluster 탐색·decompressionuncompressed data는 direct I/O 또는 FSDAX 가능

image 생성부터 runtime lookup·decompression·direct access까지의 큰 흐름입니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 ======================================
4 EROFS - Enhanced Read-Only File System
5 ======================================
6
7 Overview
8 ========
9
10 EROFS filesystem stands for Enhanced Read-Only File System. It aims to form a
11 generic read-only filesystem solution for various read-only use cases instead
12 of just focusing on storage space saving without considering any side effects
13 of runtime performance.
14
15 It is designed to meet the needs of flexibility, feature extendability and user
16 payload friendly, etc. Apart from those, it is still kept as a simple
17 random-access friendly high-performance filesystem to get rid of unneeded I/O
18 amplification and memory-resident overhead compared to similar approaches.
19
20 It is implemented to be a better choice for the following scenarios:
21
22 - read-only storage media or
23
24 - part of a fully trusted read-only solution, which means it needs to be
25 immutable and bit-for-bit identical to the official golden image for
26 their releases due to security or other considerations and
27
28 - hope to minimize extra storage space with guaranteed end-to-end performance
29 by using compact layout, transparent file compression and direct access,
30 especially for those embedded devices with limited memory and high-density
31 hosts with numerous containers.
32
33 Here are the main features of EROFS:
34
35 - Little endian on-disk design;
36
37 - Block-based distribution and file-based distribution over fscache are
38 supported;
39
40 - Support multiple devices to refer to external blobs, which can be used
41 for container images;
42
43 - 32-bit block addresses for each device, therefore 16TiB address space at
44 most with 4KiB block size for now;
45
46 - Two inode layouts for different requirements:
47
48 ===================== ============ ======================================
49 compact (v1) extended (v2)
50 ===================== ============ ======================================
51 Inode metadata size 32 bytes 64 bytes
52 Max file size 4 GiB 16 EiB (also limited by max. vol size)
53 Max uids/gids 65536 4294967296
54 Per-inode timestamp no yes (64 + 32-bit timestamp)
55 Max hardlinks 65536 4294967296
56 Metadata reserved 8 bytes 18 bytes
57 ===================== ============ ======================================
58
59 - Support extended attributes as an option;
60
61 - Support a bloom filter that speeds up negative extended attribute lookups;
62
63 - Support POSIX.1e ACLs by using extended attributes;
64
65 - Support transparent data compression as an option:
66 LZ4, MicroLZMA and DEFLATE algorithms can be used on a per-file basis; In
67 addition, inplace decompression is also supported to avoid bounce compressed
68 buffers and unnecessary page cache thrashing.
69
70 - Support chunk-based data deduplication and rolling-hash compressed data
71 deduplication;
72
73 - Support tailpacking inline compared to byte-addressed unaligned metadata
74 or smaller block size alternatives;
75
76 - Support merging tail-end data into a special inode as fragments.
77
78 - Support large folios to make use of THPs (Transparent Hugepages);
79
80 - Support direct I/O on uncompressed files to avoid double caching for loop
81 devices;
82
83 - Support FSDAX on uncompressed images for secure containers and ramdisks in
84 order to get rid of unnecessary page cache.
85
86 - Support file-based on-demand loading with the Fscache infrastructure.
87
88 The following git tree provides the file system user-space tools under
89 development, such as a formatting tool (mkfs.erofs), an on-disk consistency &
90 compatibility checking tool (fsck.erofs), and a debugging tool (dump.erofs):
91
92 - git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs-utils.git
93
94 For more information, please also refer to the documentation site:
95
96 - https://erofs.docs.kernel.org
97
98 Bugs and patches are welcome, please kindly help us and send to the following
99 linux-erofs mailing list:
100
101 - linux-erofs mailing list <linux-erofs@lists.ozlabs.org>
102
103 Mount options
104 =============
105
106 =================== =========================================================
107 (no)user_xattr Setup Extended User Attributes. Note: xattr is enabled
108 by default if CONFIG_EROFS_FS_XATTR is selected.
109 (no)acl Setup POSIX Access Control List. Note: acl is enabled
110 by default if CONFIG_EROFS_FS_POSIX_ACL is selected.
111 cache_strategy=%s Select a strategy for cached decompression from now on:
112
113 ========== =============================================
114 disabled In-place I/O decompression only;
115 readahead Cache the last incomplete compressed physical
116 cluster for further reading. It still does
117 in-place I/O decompression for the rest
118 compressed physical clusters;
119 readaround Cache both ends of incomplete compressed
120 physical clusters for further reading.
121 It still does in-place I/O decompression
122 for the rest compressed physical clusters.
123 ========== =============================================
124 dax={always,never} Use direct access (no page cache). See
125 Documentation/filesystems/dax.rst.
126 dax A legacy option which is an alias for ``dax=always``.
127 device=%s Specify a path to an extra device to be used together.
128 fsid=%s Specify a filesystem image ID for Fscache back-end.
129 domain_id=%s Specify a domain ID in fscache mode so that different images
130 with the same blobs under a given domain ID can share storage.
131 fsoffset=%llu Specify block-aligned filesystem offset for the primary device.
132 =================== =========================================================
133
134 Sysfs Entries
135 =============
136
137 Information about mounted erofs file systems can be found in /sys/fs/erofs.
138 Each mounted filesystem will have a directory in /sys/fs/erofs based on its
139 device name (i.e., /sys/fs/erofs/sda).
140 (see also Documentation/ABI/testing/sysfs-fs-erofs)
141
142 On-disk details
143 ===============
144
145 Summary
146 -------
147 Different from other read-only file systems, an EROFS volume is designed
148 to be as simple as possible::
149
150 |-> aligned with the block size
151 ____________________________________________________________
152 | |SB| | ... | Metadata | ... | Data | Metadata | ... | Data |
153 |_|__|_|_____|__________|_____|______|__________|_____|______|
154 0 +1K
155
156 All data areas should be aligned with the block size, but metadata areas
157 may not. All metadatas can be now observed in two different spaces (views):
158
159 1. Inode metadata space
160
161 Each valid inode should be aligned with an inode slot, which is a fixed
162 value (32 bytes) and designed to be kept in line with compact inode size.
163
164 Each inode can be directly found with the following formula:
165 inode offset = meta_blkaddr * block_size + 32 * nid
166
167 ::
168
169 |-> aligned with 8B
170 |-> followed closely
171 + meta_blkaddr blocks |-> another slot
172 _____________________________________________________________________
173 | ... | inode | xattrs | extents | data inline | ... | inode ...
174 |________|_______|(optional)|(optional)|__(optional)_|_____|__________
175 |-> aligned with the inode slot size
176 . .
177 . .
178 . .
179 . .
180 . .
181 . .
182 .____________________________________________________|-> aligned with 4B
183 | xattr_ibody_header | shared xattrs | inline xattrs |
184 |____________________|_______________|_______________|
185 |-> 12 bytes <-|->x * 4 bytes<-| .
186 . . .
187 . . .
188 . . .
189 ._______________________________.______________________.
190 | id | id | id | id | ... | id | ent | ... | ent| ... |
191 |____|____|____|____|______|____|_____|_____|____|_____|
192 |-> aligned with 4B
193 |-> aligned with 4B
194
195 Inode could be 32 or 64 bytes, which can be distinguished from a common
196 field which all inode versions have -- i_format::
197
198 __________________ __________________
199 | i_format | | i_format |
200 |__________________| |__________________|
201 | ... | | ... |
202 | | | |
203 |__________________| 32 bytes | |
204 | |
205 |__________________| 64 bytes
206
207 Xattrs, extents, data inline are placed after the corresponding inode with
208 proper alignment, and they could be optional for different data mappings.
209 _currently_ total 5 data layouts are supported:
210
211 == ====================================================================
212 0 flat file data without data inline (no extent);
213 1 fixed-sized output data compression (with non-compacted indexes);
214 2 flat file data with tail packing data inline (no extent);
215 3 fixed-sized output data compression (with compacted indexes, v5.3+);
216 4 chunk-based file (v5.15+).
217 == ====================================================================
218
219 The size of the optional xattrs is indicated by i_xattr_count in inode
220 header. Large xattrs or xattrs shared by many different files can be
221 stored in shared xattrs metadata rather than inlined right after inode.
222
223 2. Shared xattrs metadata space
224
225 Shared xattrs space is similar to the above inode space, started with
226 a specific block indicated by xattr_blkaddr, organized one by one with
227 proper align.
228
229 Each share xattr can also be directly found by the following formula:
230 xattr offset = xattr_blkaddr * block_size + 4 * xattr_id
231
232 ::
233
234 |-> aligned by 4 bytes
235 + xattr_blkaddr blocks |-> aligned with 4 bytes
236 _________________________________________________________________________
237 | ... | xattr_entry | xattr data | ... | xattr_entry | xattr data ...
238 |________|_____________|_____________|_____|______________|_______________
239
240 Directories
241 -----------
242 All directories are now organized in a compact on-disk format. Note that
243 each directory block is divided into index and name areas in order to support
244 random file lookup, and all directory entries are _strictly_ recorded in
245 alphabetical order in order to support improved prefix binary search
246 algorithm (could refer to the related source code).
247
248 ::
249
250 ___________________________
251 / |
252 / ______________|________________
253 / / | nameoff1 | nameoffN-1
254 ____________.______________._______________v________________v__________
255 | dirent | dirent | ... | dirent | filename | filename | ... | filename |
256 |___.0___|____1___|_____|___N-1__|____0_____|____1_____|_____|___N-1____|
257 \ ^
258 \ | * could have
259 \ | trailing '\0'
260 \________________________| nameoff0
261 Directory block
262
263 Note that apart from the offset of the first filename, nameoff0 also indicates
264 the total number of directory entries in this block since it is no need to
265 introduce another on-disk field at all.
266
267 Chunk-based files
268 -----------------
269 In order to support chunk-based data deduplication, a new inode data layout has
270 been supported since Linux v5.15: Files are split in equal-sized data chunks
271 with ``extents`` area of the inode metadata indicating how to get the chunk
272 data: these can be simply as a 4-byte block address array or in the 8-byte
273 chunk index form (see struct erofs_inode_chunk_index in erofs_fs.h for more
274 details.)
275
276 By the way, chunk-based files are all uncompressed for now.
277
278 Long extended attribute name prefixes
279 -------------------------------------
280 There are use cases where extended attributes with different values can have
281 only a few common prefixes (such as overlayfs xattrs). The predefined prefixes
282 work inefficiently in both image size and runtime performance in such cases.
283
284 The long xattr name prefixes feature is introduced to address this issue. The
285 overall idea is that, apart from the existing predefined prefixes, the xattr
286 entry could also refer to user-specified long xattr name prefixes, e.g.
287 "trusted.overlay.".
288
289 When referring to a long xattr name prefix, the highest bit (bit 7) of
290 erofs_xattr_entry.e_name_index is set, while the lower bits (bit 0-6) as a whole
291 represent the index of the referred long name prefix among all long name
292 prefixes. Therefore, only the trailing part of the name apart from the long
293 xattr name prefix is stored in erofs_xattr_entry.e_name, which could be empty if
294 the full xattr name matches exactly as its long xattr name prefix.
295
296 All long xattr prefixes are stored one by one in the packed inode as long as
297 the packed inode is valid, or in the meta inode otherwise. The
298 xattr_prefix_count (of the on-disk superblock) indicates the total number of
299 long xattr name prefixes, while (xattr_prefix_start * 4) indicates the start
300 offset of long name prefixes in the packed/meta inode. Note that, long extended
301 attribute name prefixes are disabled if xattr_prefix_count is 0.
302
303 Each long name prefix is stored in the format: ALIGN({__le16 len, data}, 4),
304 where len represents the total size of the data part. The data part is actually
305 represented by 'struct erofs_xattr_long_prefix', where base_index represents the
306 index of the predefined xattr name prefix, e.g. EROFS_XATTR_INDEX_TRUSTED for
307 "trusted.overlay." long name prefix, while the infix string keeps the string
308 after stripping the short prefix, e.g. "overlay." for the example above.
309
310 Data compression
311 ----------------
312 EROFS implements fixed-sized output compression which generates fixed-sized
313 compressed data blocks from variable-sized input in contrast to other existing
314 fixed-sized input solutions. Relatively higher compression ratios can be gotten
315 by using fixed-sized output compression since nowadays popular data compression
316 algorithms are mostly LZ77-based and such fixed-sized output approach can be
317 benefited from the historical dictionary (aka. sliding window).
318
319 In details, original (uncompressed) data is turned into several variable-sized
320 extents and in the meanwhile, compressed into physical clusters (pclusters).
321 In order to record each variable-sized extent, logical clusters (lclusters) are
322 introduced as the basic unit of compress indexes to indicate whether a new
323 extent is generated within the range (HEAD) or not (NONHEAD). Lclusters are now
324 fixed in block size, as illustrated below::
325
326 |<- variable-sized extent ->|<- VLE ->|
327 clusterofs clusterofs clusterofs
328 | | |
329 _________v_________________________________v_______________________v________
330 ... | . | | . | | . ...
331 ____|____._________|______________|________.___ _|______________|__.________
332 |-> lcluster <-|-> lcluster <-|-> lcluster <-|-> lcluster <-|
333 (HEAD) (NONHEAD) (HEAD) (NONHEAD) .
334 . CBLKCNT . .
335 . . .
336 . . .
337 _______._____________________________.______________._________________
338 ... | | | | ...
339 _______|______________|______________|______________|_________________
340 |-> big pcluster <-|-> pcluster <-|
341
342 A physical cluster can be seen as a container of physical compressed blocks
343 which contains compressed data. Previously, only lcluster-sized (4KB) pclusters
344 were supported. After big pcluster feature is introduced (available since
345 Linux v5.13), pcluster can be a multiple of lcluster size.
346
347 For each HEAD lcluster, clusterofs is recorded to indicate where a new extent
348 starts and blkaddr is used to seek the compressed data. For each NONHEAD
349 lcluster, delta0 and delta1 are available instead of blkaddr to indicate the
350 distance to its HEAD lcluster and the next HEAD lcluster. A PLAIN lcluster is
351 also a HEAD lcluster except that its data is uncompressed. See the comments
352 around "struct z_erofs_vle_decompressed_index" in erofs_fs.h for more details.
353
354 If big pcluster is enabled, pcluster size in lclusters needs to be recorded as
355 well. Let the delta0 of the first NONHEAD lcluster store the compressed block
356 count with a special flag as a new called CBLKCNT NONHEAD lcluster. It's easy
357 to understand its delta0 is constantly 1, as illustrated below::
358
359 __________________________________________________________
360 | HEAD | NONHEAD | NONHEAD | ... | NONHEAD | HEAD | HEAD |
361 |__:___|_(CBLKCNT)_|_________|_____|_________|__:___|____:_|
362 |<----- a big pcluster (with CBLKCNT) ------>|<-- -->|
363 a lcluster-sized pcluster (without CBLKCNT) ^
364
365 If another HEAD follows a HEAD lcluster, there is no room to record CBLKCNT,
366 but it's easy to know the size of such pcluster is 1 lcluster as well.
367
368 Since Linux v6.1, each pcluster can be used for multiple variable-sized extents,
369 therefore it can be used for compressed data deduplication.
370

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

목표와 주요 사용 시나리오

1-32

EROFS는 Enhanced Read-Only File System의 약자입니다. 저장 공간 절약만 추구해 runtime 성능에 생기는 부작용을 무시하는 대신, 여러 read-only 용도에 적용할 수 있는 범용 filesystem을 목표로 합니다.

유연성, feature 확장성, user payload 친화성을 충족하면서도 구조를 단순하게 유지합니다. random access에 유리한 고성능 설계로 불필요한 I/O amplification과 memory-resident overhead를 비슷한 접근보다 줄입니다.

첫 번째 대상은 read-only storage media입니다. 두 번째는 security 등의 이유로 immutable하며 release의 공식 golden image와 bit-for-bit 동일해야 하는 완전히 신뢰된 read-only solution의 일부입니다.

세 번째는 compact layout, transparent file compression과 direct access를 이용해 end-to-end 성능을 보장하면서 추가 저장 공간을 최소화하려는 환경입니다. memory가 제한된 embedded device와 container가 많은 고밀도 host가 특히 해당합니다.

EROFS 설계 우선순위
read-only 또는 immutable image를 입력compact metadata layout 적용선택적 transparent compression·deduplication 적용random access와 direct access 유지I/O amplification·resident memory 절감embedded·container workload에 배포

저장 효율과 runtime 성능을 함께 달성하려는 목표를 정리합니다.

.. SPDX-License-Identifier: GPL-2.0

======================================
EROFS - Enhanced Read-Only File System
======================================

Overview
========

EROFS filesystem stands for Enhanced Read-Only File System.  It aims to form a
generic read-only filesystem solution for various read-only use cases instead
of just focusing on storage space saving without considering any side effects
of runtime performance.

It is designed to meet the needs of flexibility, feature extendability and user
payload friendly, etc.  Apart from those, it is still kept as a simple
random-access friendly high-performance filesystem to get rid of unneeded I/O
amplification and memory-resident overhead compared to similar approaches.

It is implemented to be a better choice for the following scenarios:

 - read-only storage media or

 - part of a fully trusted read-only solution, which means it needs to be
   immutable and bit-for-bit identical to the official golden image for
   their releases due to security or other considerations and

 - hope to minimize extra storage space with guaranteed end-to-end performance
   by using compact layout, transparent file compression and direct access,
   especially for those embedded devices with limited memory and high-density
   hosts with numerous containers.

On-disk 기능, inode 형식과 userspace 도구

33-102

EROFS on-disk format은 little-endian입니다. block-based distribution뿐 아니라 Fscache를 통한 file-based distribution도 지원하고, container image 등에 쓰는 external blob을 여러 device에서 참조할 수 있습니다.

device마다 32-bit block address를 사용하므로 현재 4-KiB block size에서는 최대 address space가 16 TiB입니다.

compact(v1) inode는 32 bytes이고 최대 file size 4 GiB, uid/gid와 hardlink 최대 65536, inode별 timestamp 없음, metadata reserve 8 bytes입니다. extended(v2) inode는 64 bytes이고 volume 한도 내 최대 16 EiB file, 32-bit uid/gid와 hardlink count, 64+32-bit timestamp, reserve 18 bytes를 제공합니다.

extended attribute는 선택 기능이며 negative xattr lookup을 빠르게 하는 bloom filter와 xattr 기반 POSIX.1e ACL을 지원합니다.

file별로 LZ4, MicroLZMA, DEFLATE transparent compression을 선택할 수 있습니다. inplace decompression은 bounce compressed buffer와 불필요한 page-cache thrashing을 피합니다.

chunk-based data deduplication과 rolling-hash compressed data deduplication, byte-addressed unaligned metadata나 더 작은 block size 대신 쓰는 tailpacking inline, tail-end data를 특별 inode의 fragment로 합치는 기능을 지원합니다.

large folio로 THP를 활용하며 uncompressed file direct I/O로 loop device의 double caching을 피합니다. secure container와 ramdisk용 uncompressed image에 FSDAX를 적용해 page cache를 제거하고, Fscache infrastructure를 통한 file-based on-demand loading도 지원합니다.

userspace 도구는 erofs-utils git tree에 있으며 `mkfs.erofs`, `fsck.erofs`, `dump.erofs`를 포함합니다. 문서 사이트는 `https://erofs.docs.kernel.org`, bug와 patch는 `linux-erofs@lists.ozlabs.org` mailing list로 보냅니다.

EROFS inode format 비교
항목compact (v1)extended (v2)
Metadata size32 bytes64 bytes
Max file size4 GiB16 EiB, volume 한도 적용
UID/GID655364294967296
Per-inode timestamp없음64+32-bit
Max hardlinks655364294967296
Reserved metadata8 bytes18 bytes

compact와 extended inode가 제공하는 범위입니다.

Here are the main features of EROFS:

 - Little endian on-disk design;

 - Block-based distribution and file-based distribution over fscache are
   supported;

 - Support multiple devices to refer to external blobs, which can be used
   for container images;

 - 32-bit block addresses for each device, therefore 16TiB address space at
   most with 4KiB block size for now;

 - Two inode layouts for different requirements:

   =====================  ============  ======================================
                          compact (v1)  extended (v2)
   =====================  ============  ======================================
   Inode metadata size    32 bytes      64 bytes
   Max file size          4 GiB         16 EiB (also limited by max. vol size)
   Max uids/gids          65536         4294967296
   Per-inode timestamp    no            yes (64 + 32-bit timestamp)
   Max hardlinks          65536         4294967296
   Metadata reserved      8 bytes       18 bytes
   =====================  ============  ======================================

 - Support extended attributes as an option;

 - Support a bloom filter that speeds up negative extended attribute lookups;

 - Support POSIX.1e ACLs by using extended attributes;

 - Support transparent data compression as an option:
   LZ4, MicroLZMA and DEFLATE algorithms can be used on a per-file basis; In
   addition, inplace decompression is also supported to avoid bounce compressed
   buffers and unnecessary page cache thrashing.

 - Support chunk-based data deduplication and rolling-hash compressed data
   deduplication;

 - Support tailpacking inline compared to byte-addressed unaligned metadata
   or smaller block size alternatives;

 - Support merging tail-end data into a special inode as fragments.

 - Support large folios to make use of THPs (Transparent Hugepages);

 - Support direct I/O on uncompressed files to avoid double caching for loop
   devices;

 - Support FSDAX on uncompressed images for secure containers and ramdisks in
   order to get rid of unnecessary page cache.

 - Support file-based on-demand loading with the Fscache infrastructure.

The following git tree provides the file system user-space tools under
development, such as a formatting tool (mkfs.erofs), an on-disk consistency &
compatibility checking tool (fsck.erofs), and a debugging tool (dump.erofs):

- git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs-utils.git

For more information, please also refer to the documentation site:

- https://erofs.docs.kernel.org

Bugs and patches are welcome, please kindly help us and send to the following
linux-erofs mailing list:

- linux-erofs mailing list   <linux-erofs@lists.ozlabs.org>

Mount options와 sysfs entry

103-140

`user_xattr`과 `nouser_xattr`은 extended user attribute를 제어합니다. `CONFIG_EROFS_FS_XATTR`을 선택하면 기본 활성화됩니다. `acl`과 `noacl`은 POSIX ACL을 제어하며 `CONFIG_EROFS_FS_POSIX_ACL`을 선택하면 기본 활성화됩니다.

`cache_strategy=disabled`는 in-place I/O decompression만 사용합니다. `readahead`는 마지막 incomplete compressed physical cluster를 이후 read용으로 cache하고 나머지는 in-place로 처리합니다. `readaround`는 incomplete cluster의 양 끝을 cache하며 나머지는 역시 in-place로 처리합니다.

`dax=always`와 `dax=never`는 page cache 없는 direct access 사용 여부를 정합니다. legacy `dax`는 `dax=always` alias입니다. 자세한 내용은 `Documentation/filesystems/dax.rst`를 참고합니다.

`device=%s`는 함께 사용할 extra device path, `fsid=%s`는 Fscache backend의 filesystem image ID, `domain_id=%s`는 같은 domain에서 blob이 같은 image들이 storage를 공유하도록 하는 domain ID입니다. `fsoffset=%llu`는 primary device에서 block-aligned filesystem offset을 지정합니다.

mount된 EROFS 정보는 `/sys/fs/erofs`에서 찾습니다. 각 filesystem은 device 이름을 사용한 directory, 예를 들어 `/sys/fs/erofs/sda`를 가집니다. ABI는 `Documentation/ABI/testing/sysfs-fs-erofs`에 정의되어 있습니다.

EROFS mount option
Option역할
`(no)user_xattr`extended user attributes
`(no)acl`POSIX ACL
`cache_strategy=disabled`in-place decompression만
`cache_strategy=readahead`마지막 incomplete pcluster cache
`cache_strategy=readaround`incomplete pcluster 양 끝 cache
`dax={always,never}`direct access 정책
`device=%s`extra device
`fsid`, `domain_id`Fscache image·공유 domain 식별
`fsoffset=%llu`primary device의 block-aligned offset

runtime data path와 external storage 식별에 쓰는 option입니다.

Mount options
=============

===================    =========================================================
(no)user_xattr         Setup Extended User Attributes. Note: xattr is enabled
                       by default if CONFIG_EROFS_FS_XATTR is selected.
(no)acl                Setup POSIX Access Control List. Note: acl is enabled
                       by default if CONFIG_EROFS_FS_POSIX_ACL is selected.
cache_strategy=%s      Select a strategy for cached decompression from now on:

                       ==========  =============================================
                         disabled  In-place I/O decompression only;
                        readahead  Cache the last incomplete compressed physical
                                   cluster for further reading. It still does
                                   in-place I/O decompression for the rest
                                   compressed physical clusters;
                       readaround  Cache both ends of incomplete compressed
                                   physical clusters for further reading.
                                   It still does in-place I/O decompression
                                   for the rest compressed physical clusters.
                       ==========  =============================================
dax={always,never}     Use direct access (no page cache).  See
                       Documentation/filesystems/dax.rst.
dax                    A legacy option which is an alias for ``dax=always``.
device=%s              Specify a path to an extra device to be used together.
fsid=%s                Specify a filesystem image ID for Fscache back-end.
domain_id=%s           Specify a domain ID in fscache mode so that different images
                       with the same blobs under a given domain ID can share storage.
fsoffset=%llu          Specify block-aligned filesystem offset for the primary device.
===================    =========================================================

Sysfs Entries
=============

Information about mounted erofs file systems can be found in /sys/fs/erofs.
Each mounted filesystem will have a directory in /sys/fs/erofs based on its
device name (i.e., /sys/fs/erofs/sda).
(see also Documentation/ABI/testing/sysfs-fs-erofs)

Volume layout와 inode 주소식

141-166

EROFS volume은 가능한 한 단순하게 설계됩니다. superblock은 offset 1 KiB 부근에 있고, 그 뒤 metadata와 data area가 교차해 배치될 수 있습니다.

모든 data area는 block size에 맞춰 align해야 하지만 metadata area는 반드시 block-aligned일 필요가 없습니다. metadata는 inode metadata space와 shared xattrs metadata space라는 두 view로 관찰합니다.

유효 inode는 compact inode 크기와 맞춘 고정 32-byte inode slot에 정렬됩니다.

inode의 직접 주소식은 `inode offset = meta_blkaddr * block_size + 32 * nid`입니다.

EROFS volume 기본 배치
영역배치·정렬
Superblockvolume 시작 +1 KiB
Metadatablock alignment가 필수는 아님
Datablock size에 정렬
Inode slot32-byte 고정 단위
Inode offset`meta_blkaddr * block_size + 32 * nid`

원문의 volume ASCII layout을 field와 alignment 제약으로 정리했습니다.


On-disk details
===============

Summary
-------
Different from other read-only file systems, an EROFS volume is designed
to be as simple as possible::

                                |-> aligned with the block size
   ____________________________________________________________
  | |SB| | ... | Metadata | ... | Data | Metadata | ... | Data |
  |_|__|_|_____|__________|_____|______|__________|_____|______|
  0 +1K

All data areas should be aligned with the block size, but metadata areas
may not. All metadatas can be now observed in two different spaces (views):

 1. Inode metadata space

    Each valid inode should be aligned with an inode slot, which is a fixed
    value (32 bytes) and designed to be kept in line with compact inode size.

    Each inode can be directly found with the following formula:
         inode offset = meta_blkaddr * block_size + 32 * nid

Inode metadata, data layout와 shared xattr

167-238

inode metadata space에서는 inode 뒤에 optional xattrs, extents, inline data가 각 형식의 alignment에 맞춰 이어집니다. inode slot은 32 bytes, 일부 구조는 8-byte 또는 4-byte alignment를 사용합니다.

xattr 영역은 12-byte `xattr_ibody_header`, shared xattr ID 배열, inline xattr entry와 data로 구성됩니다. shared ID와 entry는 4-byte 단위로 정렬됩니다.

모든 inode version에 공통인 `i_format` field로 32-byte compact inode와 64-byte extended inode를 구분합니다.

지원하는 data layout은 다섯 가지입니다. 0은 inline 없는 flat file, 1은 non-compacted index의 fixed-sized output compression, 2는 tail-packing inline을 가진 flat file, 3은 v5.3부터의 compacted index compression, 4는 v5.15부터의 chunk-based file입니다.

optional xattr 크기는 inode header의 `i_xattr_count`가 나타냅니다. 큰 xattr 또는 여러 file이 공유하는 xattr은 inode 바로 뒤에 inline하지 않고 shared xattrs metadata에 저장할 수 있습니다.

shared xattr space는 `xattr_blkaddr`가 가리키는 block에서 시작하고 각 xattr을 적절히 align해 연속 배치합니다. 개별 shared xattr의 주소식은 `xattr offset = xattr_blkaddr * block_size + 4 * xattr_id`입니다.

EROFS inode data layout 번호
IDData layout
0inline 없는 flat file
1non-compacted index fixed-output compression
2tail-packing inline flat file
3compacted index compression, v5.3+
4chunk-based file, v5.15+

inode 뒤 optional metadata와 file data mapping 형식을 구분합니다.

    ::

                                 |-> aligned with 8B
                                            |-> followed closely
     + meta_blkaddr blocks                                      |-> another slot
       _____________________________________________________________________
     |  ...   | inode |  xattrs  | extents  | data inline | ... | inode ...
     |________|_______|(optional)|(optional)|__(optional)_|_____|__________
              |-> aligned with the inode slot size
                   .                   .
                 .                         .
               .                              .
             .                                    .
           .                                         .
         .                                              .
       .____________________________________________________|-> aligned with 4B
       | xattr_ibody_header | shared xattrs | inline xattrs |
       |____________________|_______________|_______________|
       |->    12 bytes    <-|->x * 4 bytes<-|               .
                           .                .                 .
                     .                      .                   .
                .                           .                     .
            ._______________________________.______________________.
            | id | id | id | id |  ... | id | ent | ... | ent| ... |
            |____|____|____|____|______|____|_____|_____|____|_____|
                                            |-> aligned with 4B
                                                        |-> aligned with 4B

    Inode could be 32 or 64 bytes, which can be distinguished from a common
    field which all inode versions have -- i_format::

        __________________               __________________
       |     i_format     |             |     i_format     |
       |__________________|             |__________________|
       |        ...       |             |        ...       |
       |                  |             |                  |
       |__________________| 32 bytes    |                  |
                                        |                  |
                                        |__________________| 64 bytes

    Xattrs, extents, data inline are placed after the corresponding inode with
    proper alignment, and they could be optional for different data mappings.
    _currently_ total 5 data layouts are supported:

    ==  ====================================================================
     0  flat file data without data inline (no extent);
     1  fixed-sized output data compression (with non-compacted indexes);
     2  flat file data with tail packing data inline (no extent);
     3  fixed-sized output data compression (with compacted indexes, v5.3+);
     4  chunk-based file (v5.15+).
    ==  ====================================================================

    The size of the optional xattrs is indicated by i_xattr_count in inode
    header. Large xattrs or xattrs shared by many different files can be
    stored in shared xattrs metadata rather than inlined right after inode.

 2. Shared xattrs metadata space

    Shared xattrs space is similar to the above inode space, started with
    a specific block indicated by xattr_blkaddr, organized one by one with
    proper align.

    Each share xattr can also be directly found by the following formula:
         xattr offset = xattr_blkaddr * block_size + 4 * xattr_id

::

                           |-> aligned by  4 bytes
    + xattr_blkaddr blocks                     |-> aligned with 4 bytes
     _________________________________________________________________________
    |  ...   | xattr_entry |  xattr data | ... |  xattr_entry | xattr data  ...
    |________|_____________|_____________|_____|______________|_______________

Compact directory block와 nameoff0

239-266

모든 directory는 compact on-disk format으로 구성됩니다. random file lookup을 지원하기 위해 각 directory block을 index area와 name area로 나눕니다.

향상된 prefix binary search를 위해 모든 directory entry는 엄격한 alphabetical order로 기록합니다.

index area에는 `dirent 0...N-1`이 있고 name area에는 대응하는 `filename 0...N-1`이 이어집니다. 각 dirent의 name offset이 name area의 위치를 가리킵니다.

`nameoff0`은 첫 filename의 offset인 동시에 이 block의 전체 directory entry 수를 나타냅니다. 별도 on-disk field를 추가할 필요가 없기 때문입니다.

EROFS directory block
AreaContentSearch 역할
Index area`dirent[0..N-1]`각 name offset 보관
Name area`filename[0..N-1]`alphabetical order
`nameoff0`첫 filename offsetentry count도 함께 표현
Trailing byte선택적 `\0`마지막 name 뒤 가능

원문의 ASCII 구조를 index와 name area로 분리해 표현했습니다.


Directories
-----------
All directories are now organized in a compact on-disk format. Note that
each directory block is divided into index and name areas in order to support
random file lookup, and all directory entries are _strictly_ recorded in
alphabetical order in order to support improved prefix binary search
algorithm (could refer to the related source code).

::

                  ___________________________
                 /                           |
                /              ______________|________________
               /              /              | nameoff1       | nameoffN-1
  ____________.______________._______________v________________v__________
 | dirent | dirent | ... | dirent | filename | filename | ... | filename |
 |___.0___|____1___|_____|___N-1__|____0_____|____1_____|_____|___N-1____|
      \                           ^
       \                          |                           * could have
        \                         |                             trailing '\0'
         \________________________| nameoff0
                             Directory block

Note that apart from the offset of the first filename, nameoff0 also indicates
the total number of directory entries in this block since it is no need to
introduce another on-disk field at all.

Chunk file과 long xattr name prefix

267-309

chunk-based data deduplication을 위해 Linux v5.15부터 새 inode data layout을 지원합니다. file을 같은 크기의 data chunk로 나누고 inode metadata의 `extents` area가 chunk data 위치를 나타냅니다.

chunk 위치는 단순한 4-byte block address array 또는 8-byte chunk index 형식이며 자세한 구조는 `erofs_fs.h`의 `struct erofs_inode_chunk_index`를 참고합니다. 현재 chunk-based file은 모두 uncompressed입니다.

overlayfs xattr처럼 값은 달라도 공통 prefix가 몇 개뿐인 경우 predefined prefix만 쓰면 image size와 runtime 성능이 비효율적입니다. long xattr name prefix 기능은 `trusted.overlay.` 같은 사용자 지정 prefix를 xattr entry에서 참조하게 합니다.

long prefix를 참조하면 `erofs_xattr_entry.e_name_index`의 최상위 bit 7을 설정하고 bit 0-6 전체를 long prefix 배열 index로 사용합니다. `e_name`에는 prefix 뒤 trailing name만 저장하며 전체 이름이 prefix와 같으면 비어 있을 수 있습니다.

packed inode가 유효하면 long prefix를 그 안에, 아니면 meta inode에 차례로 저장합니다. superblock의 `xattr_prefix_count`는 prefix 수, `xattr_prefix_start * 4`는 packed/meta inode 안의 시작 offset입니다. count가 0이면 기능이 비활성화됩니다.

각 prefix 형식은 `ALIGN({__le16 len, data}, 4)`입니다. data는 `struct erofs_xattr_long_prefix`로 표현하며 `base_index`는 `EROFS_XATTR_INDEX_TRUSTED` 같은 predefined prefix index, infix는 short prefix를 제거한 뒤의 `overlay.` 같은 문자열입니다.

Long xattr prefix encoding
Field의미
`e_name_index[7]`long prefix 참조 표시
`e_name_index[0..6]`long prefix 배열 index
`e_name`prefix 뒤 trailing name
`xattr_prefix_count`전체 long prefix 수
`xattr_prefix_start * 4`packed/meta inode 내 시작 offset
Prefix record`ALIGN({__le16 len, data}, 4)`

full name을 predefined base와 사용자 infix로 분리하는 on-disk 표현입니다.

Chunk-based files
-----------------
In order to support chunk-based data deduplication, a new inode data layout has
been supported since Linux v5.15: Files are split in equal-sized data chunks
with ``extents`` area of the inode metadata indicating how to get the chunk
data: these can be simply as a 4-byte block address array or in the 8-byte
chunk index form (see struct erofs_inode_chunk_index in erofs_fs.h for more
details.)

By the way, chunk-based files are all uncompressed for now.

Long extended attribute name prefixes
-------------------------------------
There are use cases where extended attributes with different values can have
only a few common prefixes (such as overlayfs xattrs).  The predefined prefixes
work inefficiently in both image size and runtime performance in such cases.

The long xattr name prefixes feature is introduced to address this issue.  The
overall idea is that, apart from the existing predefined prefixes, the xattr
entry could also refer to user-specified long xattr name prefixes, e.g.
"trusted.overlay.".

When referring to a long xattr name prefix, the highest bit (bit 7) of
erofs_xattr_entry.e_name_index is set, while the lower bits (bit 0-6) as a whole
represent the index of the referred long name prefix among all long name
prefixes.  Therefore, only the trailing part of the name apart from the long
xattr name prefix is stored in erofs_xattr_entry.e_name, which could be empty if
the full xattr name matches exactly as its long xattr name prefix.

All long xattr prefixes are stored one by one in the packed inode as long as
the packed inode is valid, or in the meta inode otherwise.  The
xattr_prefix_count (of the on-disk superblock) indicates the total number of
long xattr name prefixes, while (xattr_prefix_start * 4) indicates the start
offset of long name prefixes in the packed/meta inode.  Note that, long extended
attribute name prefixes are disabled if xattr_prefix_count is 0.

Each long name prefix is stored in the format: ALIGN({__le16 len, data}, 4),
where len represents the total size of the data part.  The data part is actually
represented by 'struct erofs_xattr_long_prefix', where base_index represents the
index of the predefined xattr name prefix, e.g. EROFS_XATTR_INDEX_TRUSTED for
"trusted.overlay." long name prefix, while the infix string keeps the string
after stripping the short prefix, e.g. "overlay." for the example above.

Fixed-output compression과 big pcluster

310-369

EROFS는 variable-sized input에서 fixed-sized compressed data block을 만드는 fixed-sized output compression을 구현합니다. 기존 fixed-sized input 방식과 반대이며, 주로 LZ77 계열인 현대 compression algorithm이 historical dictionary 또는 sliding window를 활용하므로 더 높은 compression ratio를 얻을 수 있습니다.

원본 uncompressed data는 여러 variable-sized extent로 나뉘면서 physical cluster(pcluster)로 압축됩니다. 각 extent를 기록하기 위해 logical cluster(lcluster)를 compression index 기본 단위로 사용합니다.

lcluster 범위에서 새 extent가 시작하면 HEAD, 시작하지 않으면 NONHEAD입니다. lcluster 크기는 block size로 고정됩니다. `clusterofs`는 HEAD 안에서 새 extent가 시작하는 offset을 나타내고 `blkaddr`는 compressed data를 찾습니다.

NONHEAD에는 `blkaddr` 대신 `delta0`과 `delta1`이 있어 각각 자신의 HEAD와 다음 HEAD까지 거리를 나타냅니다. PLAIN lcluster도 data가 uncompressed라는 점만 제외하면 HEAD입니다. 자세한 내용은 `erofs_fs.h`의 `struct z_erofs_vle_decompressed_index` 주변 comment를 참고합니다.

pcluster는 compressed physical block을 담는 container입니다. 과거에는 lcluster 크기인 4 KiB pcluster만 지원했지만 Linux v5.13의 big pcluster 이후에는 lcluster 크기의 배수가 될 수 있습니다.

big pcluster의 lcluster 수를 기록하기 위해 첫 NONHEAD lcluster의 `delta0`에 특별 flag와 compressed block count를 저장하며 이를 CBLKCNT NONHEAD라고 부릅니다. 이 entry 자신의 `delta0`은 항상 1입니다.

HEAD 바로 뒤에 다른 HEAD가 오면 CBLKCNT를 기록할 공간이 없지만 이 pcluster가 1 lcluster 크기임을 알 수 있습니다. Linux v6.1부터 하나의 pcluster를 여러 variable-sized extent가 공유할 수 있어 compressed data deduplication에 사용합니다.

EROFS compressed extent mapping
원본 data를 variable-sized extent로 분할extent를 fixed-output pcluster로 압축시작 lcluster를 HEAD로 기록후속 lcluster를 NONHEAD와 `delta0`·`delta1`로 연결big pcluster는 첫 NONHEAD에 CBLKCNT 기록Linux v6.1+에서 pcluster를 여러 extent가 공유해 deduplication

logical extent index에서 physical compressed cluster를 찾는 흐름입니다.

Data compression
----------------
EROFS implements fixed-sized output compression which generates fixed-sized
compressed data blocks from variable-sized input in contrast to other existing
fixed-sized input solutions. Relatively higher compression ratios can be gotten
by using fixed-sized output compression since nowadays popular data compression
algorithms are mostly LZ77-based and such fixed-sized output approach can be
benefited from the historical dictionary (aka. sliding window).

In details, original (uncompressed) data is turned into several variable-sized
extents and in the meanwhile, compressed into physical clusters (pclusters).
In order to record each variable-sized extent, logical clusters (lclusters) are
introduced as the basic unit of compress indexes to indicate whether a new
extent is generated within the range (HEAD) or not (NONHEAD). Lclusters are now
fixed in block size, as illustrated below::

          |<-    variable-sized extent    ->|<-       VLE         ->|
        clusterofs                        clusterofs              clusterofs
          |                                 |                       |
 _________v_________________________________v_______________________v________
 ... |    .         |              |        .     |              |  .   ...
 ____|____._________|______________|________.___ _|______________|__.________
     |-> lcluster <-|-> lcluster <-|-> lcluster <-|-> lcluster <-|
          (HEAD)        (NONHEAD)       (HEAD)        (NONHEAD)    .
           .             CBLKCNT            .                    .
            .                               .                  .
             .                              .                .
       _______._____________________________.______________._________________
          ... |              |              |              | ...
       _______|______________|______________|______________|_________________
              |->      big pcluster       <-|-> pcluster <-|

A physical cluster can be seen as a container of physical compressed blocks
which contains compressed data. Previously, only lcluster-sized (4KB) pclusters
were supported. After big pcluster feature is introduced (available since
Linux v5.13), pcluster can be a multiple of lcluster size.

For each HEAD lcluster, clusterofs is recorded to indicate where a new extent
starts and blkaddr is used to seek the compressed data. For each NONHEAD
lcluster, delta0 and delta1 are available instead of blkaddr to indicate the
distance to its HEAD lcluster and the next HEAD lcluster. A PLAIN lcluster is
also a HEAD lcluster except that its data is uncompressed. See the comments
around "struct z_erofs_vle_decompressed_index" in erofs_fs.h for more details.

If big pcluster is enabled, pcluster size in lclusters needs to be recorded as
well. Let the delta0 of the first NONHEAD lcluster store the compressed block
count with a special flag as a new called CBLKCNT NONHEAD lcluster. It's easy
to understand its delta0 is constantly 1, as illustrated below::

   __________________________________________________________
  | HEAD |  NONHEAD  | NONHEAD | ... | NONHEAD | HEAD | HEAD |
  |__:___|_(CBLKCNT)_|_________|_____|_________|__:___|____:_|
     |<----- a big pcluster (with CBLKCNT) ------>|<--  -->|
           a lcluster-sized pcluster (without CBLKCNT) ^

If another HEAD follows a HEAD lcluster, there is no room to record CBLKCNT,
but it's easy to know the size of such pcluster is 1 lcluster as well.

Since Linux v6.1, each pcluster can be used for multiple variable-sized extents,
therefore it can be used for compressed data deduplication.