요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
.. SPDX-License-Identifier: GPL-2.0
======================================
EROFS - Enhanced Read-Only File System
======================================
Overview
========
EROFS filesystem stands for Enhanced Read-Only File System. It aims to form a
generic read-only filesystem solution for various read-only use cases instead
of just focusing on storage space saving without considering any side effects
of runtime performance.
It is designed to meet the needs of flexibility, feature extendability and user
payload friendly, etc. Apart from those, it is still kept as a simple
random-access friendly high-performance filesystem to get rid of unneeded I/O
amplification and memory-resident overhead compared to similar approaches.
It is implemented to be a better choice for the following scenarios:
- read-only storage media or
- part of a fully trusted read-only solution, which means it needs to be
immutable and bit-for-bit identical to the official golden image for
their releases due to security or other considerations and
- hope to minimize extra storage space with guaranteed end-to-end performance
by using compact layout, transparent file compression and direct access,
especially for those embedded devices with limited memory and high-density
hosts with numerous containers.
Here are the main features of EROFS:
- Little endian on-disk design;
- Block-based distribution and file-based distribution over fscache are
supported;
- Support multiple devices to refer to external blobs, which can be used
for container images;
- 32-bit block addresses for each device, therefore 16TiB address space at
most with 4KiB block size for now;
- Two inode layouts for different requirements:
===================== ============ ======================================
compact (v1) extended (v2)
===================== ============ ======================================
Inode metadata size 32 bytes 64 bytes
Max file size 4 GiB 16 EiB (also limited by max. vol size)
Max uids/gids 65536 4294967296
Per-inode timestamp no yes (64 + 32-bit timestamp)
Max hardlinks 65536 4294967296
Metadata reserved 8 bytes 18 bytes
===================== ============ ======================================
- Support extended attributes as an option;
- Support a bloom filter that speeds up negative extended attribute lookups;
- Support POSIX.1e ACLs by using extended attributes;
- Support transparent data compression as an option:
LZ4, MicroLZMA and DEFLATE algorithms can be used on a per-file basis; In
addition, inplace decompression is also supported to avoid bounce compressed
buffers and unnecessary page cache thrashing.
- Support chunk-based data deduplication and rolling-hash compressed data
deduplication;
- Support tailpacking inline compared to byte-addressed unaligned metadata
or smaller block size alternatives;
- Support merging tail-end data into a special inode as fragments.
- Support large folios to make use of THPs (Transparent Hugepages);
- Support direct I/O on uncompressed files to avoid double caching for loop
devices;
- Support FSDAX on uncompressed images for secure containers and ramdisks in
order to get rid of unnecessary page cache.
- Support file-based on-demand loading with the Fscache infrastructure.
The following git tree provides the file system user-space tools under
development, such as a formatting tool (mkfs.erofs), an on-disk consistency &
compatibility checking tool (fsck.erofs), and a debugging tool (dump.erofs):
- git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs-utils.git
For more information, please also refer to the documentation site:
- https://erofs.docs.kernel.org
Bugs and patches are welcome, please kindly help us and send to the following
linux-erofs mailing list:
- linux-erofs mailing list <linux-erofs@lists.ozlabs.org>
Mount options
=============
=================== =========================================================
(no)user_xattr Setup Extended User Attributes. Note: xattr is enabled
by default if CONFIG_EROFS_FS_XATTR is selected.
(no)acl Setup POSIX Access Control List. Note: acl is enabled
by default if CONFIG_EROFS_FS_POSIX_ACL is selected.
cache_strategy=%s Select a strategy for cached decompression from now on:
========== =============================================
disabled In-place I/O decompression only;
readahead Cache the last incomplete compressed physical
cluster for further reading. It still does
in-place I/O decompression for the rest
compressed physical clusters;
readaround Cache both ends of incomplete compressed
physical clusters for further reading.
It still does in-place I/O decompression
for the rest compressed physical clusters.
========== =============================================
dax={always,never} Use direct access (no page cache). See
Documentation/filesystems/dax.rst.
dax A legacy option which is an alias for ``dax=always``.
device=%s Specify a path to an extra device to be used together.
fsid=%s Specify a filesystem image ID for Fscache back-end.
domain_id=%s Specify a domain ID in fscache mode so that different images
with the same blobs under a given domain ID can share storage.
fsoffset=%llu Specify block-aligned filesystem offset for the primary device.
=================== =========================================================
Sysfs Entries
=============
Information about mounted erofs file systems can be found in /sys/fs/erofs.
Each mounted filesystem will have a directory in /sys/fs/erofs based on its
device name (i.e., /sys/fs/erofs/sda).
(see also Documentation/ABI/testing/sysfs-fs-erofs)
On-disk details
===============
Summary
-------
Different from other read-only file systems, an EROFS volume is designed
to be as simple as possible::
|-> aligned with the block size
____________________________________________________________
| |SB| | ... | Metadata | ... | Data | Metadata | ... | Data |
|_|__|_|_____|__________|_____|______|__________|_____|______|
0 +1K
All data areas should be aligned with the block size, but metadata areas
may not. All metadatas can be now observed in two different spaces (views):
1. Inode metadata space
Each valid inode should be aligned with an inode slot, which is a fixed
value (32 bytes) and designed to be kept in line with compact inode size.
Each inode can be directly found with the following formula:
inode offset = meta_blkaddr * block_size + 32 * nid
::
|-> aligned with 8B
|-> followed closely
+ meta_blkaddr blocks |-> another slot
_____________________________________________________________________
| ... | inode | xattrs | extents | data inline | ... | inode ...
|________|_______|(optional)|(optional)|__(optional)_|_____|__________
|-> aligned with the inode slot size
. .
. .
. .
. .
. .
. .
.____________________________________________________|-> aligned with 4B
| xattr_ibody_header | shared xattrs | inline xattrs |
|____________________|_______________|_______________|
|-> 12 bytes <-|->x * 4 bytes<-| .
. . .
. . .
. . .
._______________________________.______________________.
| id | id | id | id | ... | id | ent | ... | ent| ... |
|____|____|____|____|______|____|_____|_____|____|_____|
|-> aligned with 4B
|-> aligned with 4B
Inode could be 32 or 64 bytes, which can be distinguished from a common
field which all inode versions have -- i_format::
__________________ __________________
| i_format | | i_format |
|__________________| |__________________|
| ... | | ... |
| | | |
|__________________| 32 bytes | |
| |
|__________________| 64 bytes
Xattrs, extents, data inline are placed after the corresponding inode with
proper alignment, and they could be optional for different data mappings.
_currently_ total 5 data layouts are supported:
== ====================================================================
0 flat file data without data inline (no extent);
1 fixed-sized output data compression (with non-compacted indexes);
2 flat file data with tail packing data inline (no extent);
3 fixed-sized output data compression (with compacted indexes, v5.3+);
4 chunk-based file (v5.15+).
== ====================================================================
The size of the optional xattrs is indicated by i_xattr_count in inode
header. Large xattrs or xattrs shared by many different files can be
stored in shared xattrs metadata rather than inlined right after inode.
2. Shared xattrs metadata space
Shared xattrs space is similar to the above inode space, started with
a specific block indicated by xattr_blkaddr, organized one by one with
proper align.
Each share xattr can also be directly found by the following formula:
xattr offset = xattr_blkaddr * block_size + 4 * xattr_id
::
|-> aligned by 4 bytes
+ xattr_blkaddr blocks |-> aligned with 4 bytes
_________________________________________________________________________
| ... | xattr_entry | xattr data | ... | xattr_entry | xattr data ...
|________|_____________|_____________|_____|______________|_______________
Directories
-----------
All directories are now organized in a compact on-disk format. Note that
each directory block is divided into index and name areas in order to support
random file lookup, and all directory entries are _strictly_ recorded in
alphabetical order in order to support improved prefix binary search
algorithm (could refer to the related source code).
::
___________________________
/ |
/ ______________|________________
/ / | nameoff1 | nameoffN-1
____________.______________._______________v________________v__________
| dirent | dirent | ... | dirent | filename | filename | ... | filename |
|___.0___|____1___|_____|___N-1__|____0_____|____1_____|_____|___N-1____|
\ ^
\ | * could have
\ | trailing '\0'
\________________________| nameoff0
Directory block
Note that apart from the offset of the first filename, nameoff0 also indicates
the total number of directory entries in this block since it is no need to
introduce another on-disk field at all.
Chunk-based files
-----------------
In order to support chunk-based data deduplication, a new inode data layout has
been supported since Linux v5.15: Files are split in equal-sized data chunks
with ``extents`` area of the inode metadata indicating how to get the chunk
data: these can be simply as a 4-byte block address array or in the 8-byte
chunk index form (see struct erofs_inode_chunk_index in erofs_fs.h for more
details.)
By the way, chunk-based files are all uncompressed for now.
Long extended attribute name prefixes
-------------------------------------
There are use cases where extended attributes with different values can have
only a few common prefixes (such as overlayfs xattrs). The predefined prefixes
work inefficiently in both image size and runtime performance in such cases.
The long xattr name prefixes feature is introduced to address this issue. The
overall idea is that, apart from the existing predefined prefixes, the xattr
entry could also refer to user-specified long xattr name prefixes, e.g.
"trusted.overlay.".
When referring to a long xattr name prefix, the highest bit (bit 7) of
erofs_xattr_entry.e_name_index is set, while the lower bits (bit 0-6) as a whole
represent the index of the referred long name prefix among all long name
prefixes. Therefore, only the trailing part of the name apart from the long
xattr name prefix is stored in erofs_xattr_entry.e_name, which could be empty if
the full xattr name matches exactly as its long xattr name prefix.
All long xattr prefixes are stored one by one in the packed inode as long as
the packed inode is valid, or in the meta inode otherwise. The
xattr_prefix_count (of the on-disk superblock) indicates the total number of
long xattr name prefixes, while (xattr_prefix_start * 4) indicates the start
offset of long name prefixes in the packed/meta inode. Note that, long extended
attribute name prefixes are disabled if xattr_prefix_count is 0.
Each long name prefix is stored in the format: ALIGN({__le16 len, data}, 4),
where len represents the total size of the data part. The data part is actually
represented by 'struct erofs_xattr_long_prefix', where base_index represents the
index of the predefined xattr name prefix, e.g. EROFS_XATTR_INDEX_TRUSTED for
"trusted.overlay." long name prefix, while the infix string keeps the string
after stripping the short prefix, e.g. "overlay." for the example above.
Data compression
----------------
EROFS implements fixed-sized output compression which generates fixed-sized
compressed data blocks from variable-sized input in contrast to other existing
fixed-sized input solutions. Relatively higher compression ratios can be gotten
by using fixed-sized output compression since nowadays popular data compression
algorithms are mostly LZ77-based and such fixed-sized output approach can be
benefited from the historical dictionary (aka. sliding window).
In details, original (uncompressed) data is turned into several variable-sized
extents and in the meanwhile, compressed into physical clusters (pclusters).
In order to record each variable-sized extent, logical clusters (lclusters) are
introduced as the basic unit of compress indexes to indicate whether a new
extent is generated within the range (HEAD) or not (NONHEAD). Lclusters are now
fixed in block size, as illustrated below::
|<- variable-sized extent ->|<- VLE ->|
clusterofs clusterofs clusterofs
| | |
_________v_________________________________v_______________________v________
... | . | | . | | . ...
____|____._________|______________|________.___ _|______________|__.________
|-> lcluster <-|-> lcluster <-|-> lcluster <-|-> lcluster <-|
(HEAD) (NONHEAD) (HEAD) (NONHEAD) .
. CBLKCNT . .
. . .
. . .
_______._____________________________.______________._________________
... | | | | ...
_______|______________|______________|______________|_________________
|-> big pcluster <-|-> pcluster <-|
A physical cluster can be seen as a container of physical compressed blocks
which contains compressed data. Previously, only lcluster-sized (4KB) pclusters
were supported. After big pcluster feature is introduced (available since
Linux v5.13), pcluster can be a multiple of lcluster size.
For each HEAD lcluster, clusterofs is recorded to indicate where a new extent
starts and blkaddr is used to seek the compressed data. For each NONHEAD
lcluster, delta0 and delta1 are available instead of blkaddr to indicate the
distance to its HEAD lcluster and the next HEAD lcluster. A PLAIN lcluster is
also a HEAD lcluster except that its data is uncompressed. See the comments
around "struct z_erofs_vle_decompressed_index" in erofs_fs.h for more details.
If big pcluster is enabled, pcluster size in lclusters needs to be recorded as
well. Let the delta0 of the first NONHEAD lcluster store the compressed block
count with a special flag as a new called CBLKCNT NONHEAD lcluster. It's easy
to understand its delta0 is constantly 1, as illustrated below::
__________________________________________________________
| HEAD | NONHEAD | NONHEAD | ... | NONHEAD | HEAD | HEAD |
|__:___|_(CBLKCNT)_|_________|_____|_________|__:___|____:_|
|<----- a big pcluster (with CBLKCNT) ------>|<-- -->|
a lcluster-sized pcluster (without CBLKCNT) ^
If another HEAD follows a HEAD lcluster, there is no room to record CBLKCNT,
but it's easy to know the size of such pcluster is 1 lcluster as well.
Since Linux v6.1, each pcluster can be used for multiple variable-sized extents,
therefore it can be used for compressed data deduplication.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
목표와 주요 사용 시나리오
1-32EROFS는 Enhanced Read-Only File System의 약자입니다. 저장 공간 절약만 추구해 runtime 성능에 생기는 부작용을 무시하는 대신, 여러 read-only 용도에 적용할 수 있는 범용 filesystem을 목표로 합니다.
유연성, feature 확장성, user payload 친화성을 충족하면서도 구조를 단순하게 유지합니다. random access에 유리한 고성능 설계로 불필요한 I/O amplification과 memory-resident overhead를 비슷한 접근보다 줄입니다.
첫 번째 대상은 read-only storage media입니다. 두 번째는 security 등의 이유로 immutable하며 release의 공식 golden image와 bit-for-bit 동일해야 하는 완전히 신뢰된 read-only solution의 일부입니다.
세 번째는 compact layout, transparent file compression과 direct access를 이용해 end-to-end 성능을 보장하면서 추가 저장 공간을 최소화하려는 환경입니다. memory가 제한된 embedded device와 container가 많은 고밀도 host가 특히 해당합니다.
저장 효율과 runtime 성능을 함께 달성하려는 목표를 정리합니다.
.. SPDX-License-Identifier: GPL-2.0
======================================
EROFS - Enhanced Read-Only File System
======================================
Overview
========
EROFS filesystem stands for Enhanced Read-Only File System. It aims to form a
generic read-only filesystem solution for various read-only use cases instead
of just focusing on storage space saving without considering any side effects
of runtime performance.
It is designed to meet the needs of flexibility, feature extendability and user
payload friendly, etc. Apart from those, it is still kept as a simple
random-access friendly high-performance filesystem to get rid of unneeded I/O
amplification and memory-resident overhead compared to similar approaches.
It is implemented to be a better choice for the following scenarios:
- read-only storage media or
- part of a fully trusted read-only solution, which means it needs to be
immutable and bit-for-bit identical to the official golden image for
their releases due to security or other considerations and
- hope to minimize extra storage space with guaranteed end-to-end performance
by using compact layout, transparent file compression and direct access,
especially for those embedded devices with limited memory and high-density
hosts with numerous containers.
On-disk 기능, inode 형식과 userspace 도구
33-102EROFS on-disk format은 little-endian입니다. block-based distribution뿐 아니라 Fscache를 통한 file-based distribution도 지원하고, container image 등에 쓰는 external blob을 여러 device에서 참조할 수 있습니다.
device마다 32-bit block address를 사용하므로 현재 4-KiB block size에서는 최대 address space가 16 TiB입니다.
compact(v1) inode는 32 bytes이고 최대 file size 4 GiB, uid/gid와 hardlink 최대 65536, inode별 timestamp 없음, metadata reserve 8 bytes입니다. extended(v2) inode는 64 bytes이고 volume 한도 내 최대 16 EiB file, 32-bit uid/gid와 hardlink count, 64+32-bit timestamp, reserve 18 bytes를 제공합니다.
extended attribute는 선택 기능이며 negative xattr lookup을 빠르게 하는 bloom filter와 xattr 기반 POSIX.1e ACL을 지원합니다.
file별로 LZ4, MicroLZMA, DEFLATE transparent compression을 선택할 수 있습니다. inplace decompression은 bounce compressed buffer와 불필요한 page-cache thrashing을 피합니다.
chunk-based data deduplication과 rolling-hash compressed data deduplication, byte-addressed unaligned metadata나 더 작은 block size 대신 쓰는 tailpacking inline, tail-end data를 특별 inode의 fragment로 합치는 기능을 지원합니다.
large folio로 THP를 활용하며 uncompressed file direct I/O로 loop device의 double caching을 피합니다. secure container와 ramdisk용 uncompressed image에 FSDAX를 적용해 page cache를 제거하고, Fscache infrastructure를 통한 file-based on-demand loading도 지원합니다.
userspace 도구는 erofs-utils git tree에 있으며 `mkfs.erofs`, `fsck.erofs`, `dump.erofs`를 포함합니다. 문서 사이트는 `https://erofs.docs.kernel.org`, bug와 patch는 `linux-erofs@lists.ozlabs.org` mailing list로 보냅니다.
compact와 extended inode가 제공하는 범위입니다.
Here are the main features of EROFS:
- Little endian on-disk design;
- Block-based distribution and file-based distribution over fscache are
supported;
- Support multiple devices to refer to external blobs, which can be used
for container images;
- 32-bit block addresses for each device, therefore 16TiB address space at
most with 4KiB block size for now;
- Two inode layouts for different requirements:
===================== ============ ======================================
compact (v1) extended (v2)
===================== ============ ======================================
Inode metadata size 32 bytes 64 bytes
Max file size 4 GiB 16 EiB (also limited by max. vol size)
Max uids/gids 65536 4294967296
Per-inode timestamp no yes (64 + 32-bit timestamp)
Max hardlinks 65536 4294967296
Metadata reserved 8 bytes 18 bytes
===================== ============ ======================================
- Support extended attributes as an option;
- Support a bloom filter that speeds up negative extended attribute lookups;
- Support POSIX.1e ACLs by using extended attributes;
- Support transparent data compression as an option:
LZ4, MicroLZMA and DEFLATE algorithms can be used on a per-file basis; In
addition, inplace decompression is also supported to avoid bounce compressed
buffers and unnecessary page cache thrashing.
- Support chunk-based data deduplication and rolling-hash compressed data
deduplication;
- Support tailpacking inline compared to byte-addressed unaligned metadata
or smaller block size alternatives;
- Support merging tail-end data into a special inode as fragments.
- Support large folios to make use of THPs (Transparent Hugepages);
- Support direct I/O on uncompressed files to avoid double caching for loop
devices;
- Support FSDAX on uncompressed images for secure containers and ramdisks in
order to get rid of unnecessary page cache.
- Support file-based on-demand loading with the Fscache infrastructure.
The following git tree provides the file system user-space tools under
development, such as a formatting tool (mkfs.erofs), an on-disk consistency &
compatibility checking tool (fsck.erofs), and a debugging tool (dump.erofs):
- git://git.kernel.org/pub/scm/linux/kernel/git/xiang/erofs-utils.git
For more information, please also refer to the documentation site:
- https://erofs.docs.kernel.org
Bugs and patches are welcome, please kindly help us and send to the following
linux-erofs mailing list:
- linux-erofs mailing list <linux-erofs@lists.ozlabs.org>
Mount options와 sysfs entry
103-140`user_xattr`과 `nouser_xattr`은 extended user attribute를 제어합니다. `CONFIG_EROFS_FS_XATTR`을 선택하면 기본 활성화됩니다. `acl`과 `noacl`은 POSIX ACL을 제어하며 `CONFIG_EROFS_FS_POSIX_ACL`을 선택하면 기본 활성화됩니다.
`cache_strategy=disabled`는 in-place I/O decompression만 사용합니다. `readahead`는 마지막 incomplete compressed physical cluster를 이후 read용으로 cache하고 나머지는 in-place로 처리합니다. `readaround`는 incomplete cluster의 양 끝을 cache하며 나머지는 역시 in-place로 처리합니다.
`dax=always`와 `dax=never`는 page cache 없는 direct access 사용 여부를 정합니다. legacy `dax`는 `dax=always` alias입니다. 자세한 내용은 `Documentation/filesystems/dax.rst`를 참고합니다.
`device=%s`는 함께 사용할 extra device path, `fsid=%s`는 Fscache backend의 filesystem image ID, `domain_id=%s`는 같은 domain에서 blob이 같은 image들이 storage를 공유하도록 하는 domain ID입니다. `fsoffset=%llu`는 primary device에서 block-aligned filesystem offset을 지정합니다.
mount된 EROFS 정보는 `/sys/fs/erofs`에서 찾습니다. 각 filesystem은 device 이름을 사용한 directory, 예를 들어 `/sys/fs/erofs/sda`를 가집니다. ABI는 `Documentation/ABI/testing/sysfs-fs-erofs`에 정의되어 있습니다.
runtime data path와 external storage 식별에 쓰는 option입니다.
Mount options
=============
=================== =========================================================
(no)user_xattr Setup Extended User Attributes. Note: xattr is enabled
by default if CONFIG_EROFS_FS_XATTR is selected.
(no)acl Setup POSIX Access Control List. Note: acl is enabled
by default if CONFIG_EROFS_FS_POSIX_ACL is selected.
cache_strategy=%s Select a strategy for cached decompression from now on:
========== =============================================
disabled In-place I/O decompression only;
readahead Cache the last incomplete compressed physical
cluster for further reading. It still does
in-place I/O decompression for the rest
compressed physical clusters;
readaround Cache both ends of incomplete compressed
physical clusters for further reading.
It still does in-place I/O decompression
for the rest compressed physical clusters.
========== =============================================
dax={always,never} Use direct access (no page cache). See
Documentation/filesystems/dax.rst.
dax A legacy option which is an alias for ``dax=always``.
device=%s Specify a path to an extra device to be used together.
fsid=%s Specify a filesystem image ID for Fscache back-end.
domain_id=%s Specify a domain ID in fscache mode so that different images
with the same blobs under a given domain ID can share storage.
fsoffset=%llu Specify block-aligned filesystem offset for the primary device.
=================== =========================================================
Sysfs Entries
=============
Information about mounted erofs file systems can be found in /sys/fs/erofs.
Each mounted filesystem will have a directory in /sys/fs/erofs based on its
device name (i.e., /sys/fs/erofs/sda).
(see also Documentation/ABI/testing/sysfs-fs-erofs)
Volume layout와 inode 주소식
141-166EROFS volume은 가능한 한 단순하게 설계됩니다. superblock은 offset 1 KiB 부근에 있고, 그 뒤 metadata와 data area가 교차해 배치될 수 있습니다.
모든 data area는 block size에 맞춰 align해야 하지만 metadata area는 반드시 block-aligned일 필요가 없습니다. metadata는 inode metadata space와 shared xattrs metadata space라는 두 view로 관찰합니다.
유효 inode는 compact inode 크기와 맞춘 고정 32-byte inode slot에 정렬됩니다.
inode의 직접 주소식은 `inode offset = meta_blkaddr * block_size + 32 * nid`입니다.
원문의 volume ASCII layout을 field와 alignment 제약으로 정리했습니다.
On-disk details
===============
Summary
-------
Different from other read-only file systems, an EROFS volume is designed
to be as simple as possible::
|-> aligned with the block size
____________________________________________________________
| |SB| | ... | Metadata | ... | Data | Metadata | ... | Data |
|_|__|_|_____|__________|_____|______|__________|_____|______|
0 +1K
All data areas should be aligned with the block size, but metadata areas
may not. All metadatas can be now observed in two different spaces (views):
1. Inode metadata space
Each valid inode should be aligned with an inode slot, which is a fixed
value (32 bytes) and designed to be kept in line with compact inode size.
Each inode can be directly found with the following formula:
inode offset = meta_blkaddr * block_size + 32 * nid
Inode metadata, data layout와 shared xattr
167-238inode metadata space에서는 inode 뒤에 optional xattrs, extents, inline data가 각 형식의 alignment에 맞춰 이어집니다. inode slot은 32 bytes, 일부 구조는 8-byte 또는 4-byte alignment를 사용합니다.
xattr 영역은 12-byte `xattr_ibody_header`, shared xattr ID 배열, inline xattr entry와 data로 구성됩니다. shared ID와 entry는 4-byte 단위로 정렬됩니다.
모든 inode version에 공통인 `i_format` field로 32-byte compact inode와 64-byte extended inode를 구분합니다.
지원하는 data layout은 다섯 가지입니다. 0은 inline 없는 flat file, 1은 non-compacted index의 fixed-sized output compression, 2는 tail-packing inline을 가진 flat file, 3은 v5.3부터의 compacted index compression, 4는 v5.15부터의 chunk-based file입니다.
optional xattr 크기는 inode header의 `i_xattr_count`가 나타냅니다. 큰 xattr 또는 여러 file이 공유하는 xattr은 inode 바로 뒤에 inline하지 않고 shared xattrs metadata에 저장할 수 있습니다.
shared xattr space는 `xattr_blkaddr`가 가리키는 block에서 시작하고 각 xattr을 적절히 align해 연속 배치합니다. 개별 shared xattr의 주소식은 `xattr offset = xattr_blkaddr * block_size + 4 * xattr_id`입니다.
inode 뒤 optional metadata와 file data mapping 형식을 구분합니다.
::
|-> aligned with 8B
|-> followed closely
+ meta_blkaddr blocks |-> another slot
_____________________________________________________________________
| ... | inode | xattrs | extents | data inline | ... | inode ...
|________|_______|(optional)|(optional)|__(optional)_|_____|__________
|-> aligned with the inode slot size
. .
. .
. .
. .
. .
. .
.____________________________________________________|-> aligned with 4B
| xattr_ibody_header | shared xattrs | inline xattrs |
|____________________|_______________|_______________|
|-> 12 bytes <-|->x * 4 bytes<-| .
. . .
. . .
. . .
._______________________________.______________________.
| id | id | id | id | ... | id | ent | ... | ent| ... |
|____|____|____|____|______|____|_____|_____|____|_____|
|-> aligned with 4B
|-> aligned with 4B
Inode could be 32 or 64 bytes, which can be distinguished from a common
field which all inode versions have -- i_format::
__________________ __________________
| i_format | | i_format |
|__________________| |__________________|
| ... | | ... |
| | | |
|__________________| 32 bytes | |
| |
|__________________| 64 bytes
Xattrs, extents, data inline are placed after the corresponding inode with
proper alignment, and they could be optional for different data mappings.
_currently_ total 5 data layouts are supported:
== ====================================================================
0 flat file data without data inline (no extent);
1 fixed-sized output data compression (with non-compacted indexes);
2 flat file data with tail packing data inline (no extent);
3 fixed-sized output data compression (with compacted indexes, v5.3+);
4 chunk-based file (v5.15+).
== ====================================================================
The size of the optional xattrs is indicated by i_xattr_count in inode
header. Large xattrs or xattrs shared by many different files can be
stored in shared xattrs metadata rather than inlined right after inode.
2. Shared xattrs metadata space
Shared xattrs space is similar to the above inode space, started with
a specific block indicated by xattr_blkaddr, organized one by one with
proper align.
Each share xattr can also be directly found by the following formula:
xattr offset = xattr_blkaddr * block_size + 4 * xattr_id
::
|-> aligned by 4 bytes
+ xattr_blkaddr blocks |-> aligned with 4 bytes
_________________________________________________________________________
| ... | xattr_entry | xattr data | ... | xattr_entry | xattr data ...
|________|_____________|_____________|_____|______________|_______________
Compact directory block와 nameoff0
239-266모든 directory는 compact on-disk format으로 구성됩니다. random file lookup을 지원하기 위해 각 directory block을 index area와 name area로 나눕니다.
향상된 prefix binary search를 위해 모든 directory entry는 엄격한 alphabetical order로 기록합니다.
index area에는 `dirent 0...N-1`이 있고 name area에는 대응하는 `filename 0...N-1`이 이어집니다. 각 dirent의 name offset이 name area의 위치를 가리킵니다.
`nameoff0`은 첫 filename의 offset인 동시에 이 block의 전체 directory entry 수를 나타냅니다. 별도 on-disk field를 추가할 필요가 없기 때문입니다.
원문의 ASCII 구조를 index와 name area로 분리해 표현했습니다.
Directories
-----------
All directories are now organized in a compact on-disk format. Note that
each directory block is divided into index and name areas in order to support
random file lookup, and all directory entries are _strictly_ recorded in
alphabetical order in order to support improved prefix binary search
algorithm (could refer to the related source code).
::
___________________________
/ |
/ ______________|________________
/ / | nameoff1 | nameoffN-1
____________.______________._______________v________________v__________
| dirent | dirent | ... | dirent | filename | filename | ... | filename |
|___.0___|____1___|_____|___N-1__|____0_____|____1_____|_____|___N-1____|
\ ^
\ | * could have
\ | trailing '\0'
\________________________| nameoff0
Directory block
Note that apart from the offset of the first filename, nameoff0 also indicates
the total number of directory entries in this block since it is no need to
introduce another on-disk field at all.
Chunk file과 long xattr name prefix
267-309chunk-based data deduplication을 위해 Linux v5.15부터 새 inode data layout을 지원합니다. file을 같은 크기의 data chunk로 나누고 inode metadata의 `extents` area가 chunk data 위치를 나타냅니다.
chunk 위치는 단순한 4-byte block address array 또는 8-byte chunk index 형식이며 자세한 구조는 `erofs_fs.h`의 `struct erofs_inode_chunk_index`를 참고합니다. 현재 chunk-based file은 모두 uncompressed입니다.
overlayfs xattr처럼 값은 달라도 공통 prefix가 몇 개뿐인 경우 predefined prefix만 쓰면 image size와 runtime 성능이 비효율적입니다. long xattr name prefix 기능은 `trusted.overlay.` 같은 사용자 지정 prefix를 xattr entry에서 참조하게 합니다.
long prefix를 참조하면 `erofs_xattr_entry.e_name_index`의 최상위 bit 7을 설정하고 bit 0-6 전체를 long prefix 배열 index로 사용합니다. `e_name`에는 prefix 뒤 trailing name만 저장하며 전체 이름이 prefix와 같으면 비어 있을 수 있습니다.
packed inode가 유효하면 long prefix를 그 안에, 아니면 meta inode에 차례로 저장합니다. superblock의 `xattr_prefix_count`는 prefix 수, `xattr_prefix_start * 4`는 packed/meta inode 안의 시작 offset입니다. count가 0이면 기능이 비활성화됩니다.
각 prefix 형식은 `ALIGN({__le16 len, data}, 4)`입니다. data는 `struct erofs_xattr_long_prefix`로 표현하며 `base_index`는 `EROFS_XATTR_INDEX_TRUSTED` 같은 predefined prefix index, infix는 short prefix를 제거한 뒤의 `overlay.` 같은 문자열입니다.
full name을 predefined base와 사용자 infix로 분리하는 on-disk 표현입니다.
Chunk-based files
-----------------
In order to support chunk-based data deduplication, a new inode data layout has
been supported since Linux v5.15: Files are split in equal-sized data chunks
with ``extents`` area of the inode metadata indicating how to get the chunk
data: these can be simply as a 4-byte block address array or in the 8-byte
chunk index form (see struct erofs_inode_chunk_index in erofs_fs.h for more
details.)
By the way, chunk-based files are all uncompressed for now.
Long extended attribute name prefixes
-------------------------------------
There are use cases where extended attributes with different values can have
only a few common prefixes (such as overlayfs xattrs). The predefined prefixes
work inefficiently in both image size and runtime performance in such cases.
The long xattr name prefixes feature is introduced to address this issue. The
overall idea is that, apart from the existing predefined prefixes, the xattr
entry could also refer to user-specified long xattr name prefixes, e.g.
"trusted.overlay.".
When referring to a long xattr name prefix, the highest bit (bit 7) of
erofs_xattr_entry.e_name_index is set, while the lower bits (bit 0-6) as a whole
represent the index of the referred long name prefix among all long name
prefixes. Therefore, only the trailing part of the name apart from the long
xattr name prefix is stored in erofs_xattr_entry.e_name, which could be empty if
the full xattr name matches exactly as its long xattr name prefix.
All long xattr prefixes are stored one by one in the packed inode as long as
the packed inode is valid, or in the meta inode otherwise. The
xattr_prefix_count (of the on-disk superblock) indicates the total number of
long xattr name prefixes, while (xattr_prefix_start * 4) indicates the start
offset of long name prefixes in the packed/meta inode. Note that, long extended
attribute name prefixes are disabled if xattr_prefix_count is 0.
Each long name prefix is stored in the format: ALIGN({__le16 len, data}, 4),
where len represents the total size of the data part. The data part is actually
represented by 'struct erofs_xattr_long_prefix', where base_index represents the
index of the predefined xattr name prefix, e.g. EROFS_XATTR_INDEX_TRUSTED for
"trusted.overlay." long name prefix, while the infix string keeps the string
after stripping the short prefix, e.g. "overlay." for the example above.
Fixed-output compression과 big pcluster
310-369EROFS는 variable-sized input에서 fixed-sized compressed data block을 만드는 fixed-sized output compression을 구현합니다. 기존 fixed-sized input 방식과 반대이며, 주로 LZ77 계열인 현대 compression algorithm이 historical dictionary 또는 sliding window를 활용하므로 더 높은 compression ratio를 얻을 수 있습니다.
원본 uncompressed data는 여러 variable-sized extent로 나뉘면서 physical cluster(pcluster)로 압축됩니다. 각 extent를 기록하기 위해 logical cluster(lcluster)를 compression index 기본 단위로 사용합니다.
lcluster 범위에서 새 extent가 시작하면 HEAD, 시작하지 않으면 NONHEAD입니다. lcluster 크기는 block size로 고정됩니다. `clusterofs`는 HEAD 안에서 새 extent가 시작하는 offset을 나타내고 `blkaddr`는 compressed data를 찾습니다.
NONHEAD에는 `blkaddr` 대신 `delta0`과 `delta1`이 있어 각각 자신의 HEAD와 다음 HEAD까지 거리를 나타냅니다. PLAIN lcluster도 data가 uncompressed라는 점만 제외하면 HEAD입니다. 자세한 내용은 `erofs_fs.h`의 `struct z_erofs_vle_decompressed_index` 주변 comment를 참고합니다.
pcluster는 compressed physical block을 담는 container입니다. 과거에는 lcluster 크기인 4 KiB pcluster만 지원했지만 Linux v5.13의 big pcluster 이후에는 lcluster 크기의 배수가 될 수 있습니다.
big pcluster의 lcluster 수를 기록하기 위해 첫 NONHEAD lcluster의 `delta0`에 특별 flag와 compressed block count를 저장하며 이를 CBLKCNT NONHEAD라고 부릅니다. 이 entry 자신의 `delta0`은 항상 1입니다.
HEAD 바로 뒤에 다른 HEAD가 오면 CBLKCNT를 기록할 공간이 없지만 이 pcluster가 1 lcluster 크기임을 알 수 있습니다. Linux v6.1부터 하나의 pcluster를 여러 variable-sized extent가 공유할 수 있어 compressed data deduplication에 사용합니다.
logical extent index에서 physical compressed cluster를 찾는 흐름입니다.
Data compression
----------------
EROFS implements fixed-sized output compression which generates fixed-sized
compressed data blocks from variable-sized input in contrast to other existing
fixed-sized input solutions. Relatively higher compression ratios can be gotten
by using fixed-sized output compression since nowadays popular data compression
algorithms are mostly LZ77-based and such fixed-sized output approach can be
benefited from the historical dictionary (aka. sliding window).
In details, original (uncompressed) data is turned into several variable-sized
extents and in the meanwhile, compressed into physical clusters (pclusters).
In order to record each variable-sized extent, logical clusters (lclusters) are
introduced as the basic unit of compress indexes to indicate whether a new
extent is generated within the range (HEAD) or not (NONHEAD). Lclusters are now
fixed in block size, as illustrated below::
|<- variable-sized extent ->|<- VLE ->|
clusterofs clusterofs clusterofs
| | |
_________v_________________________________v_______________________v________
... | . | | . | | . ...
____|____._________|______________|________.___ _|______________|__.________
|-> lcluster <-|-> lcluster <-|-> lcluster <-|-> lcluster <-|
(HEAD) (NONHEAD) (HEAD) (NONHEAD) .
. CBLKCNT . .
. . .
. . .
_______._____________________________.______________._________________
... | | | | ...
_______|______________|______________|______________|_________________
|-> big pcluster <-|-> pcluster <-|
A physical cluster can be seen as a container of physical compressed blocks
which contains compressed data. Previously, only lcluster-sized (4KB) pclusters
were supported. After big pcluster feature is introduced (available since
Linux v5.13), pcluster can be a multiple of lcluster size.
For each HEAD lcluster, clusterofs is recorded to indicate where a new extent
starts and blkaddr is used to seek the compressed data. For each NONHEAD
lcluster, delta0 and delta1 are available instead of blkaddr to indicate the
distance to its HEAD lcluster and the next HEAD lcluster. A PLAIN lcluster is
also a HEAD lcluster except that its data is uncompressed. See the comments
around "struct z_erofs_vle_decompressed_index" in erofs_fs.h for more details.
If big pcluster is enabled, pcluster size in lclusters needs to be recorded as
well. Let the delta0 of the first NONHEAD lcluster store the compressed block
count with a special flag as a new called CBLKCNT NONHEAD lcluster. It's easy
to understand its delta0 is constantly 1, as illustrated below::
__________________________________________________________
| HEAD | NONHEAD | NONHEAD | ... | NONHEAD | HEAD | HEAD |
|__:___|_(CBLKCNT)_|_________|_____|_________|__:___|____:_|
|<----- a big pcluster (with CBLKCNT) ------>|<-- -->|
a lcluster-sized pcluster (without CBLKCNT) ^
If another HEAD follows a HEAD lcluster, there is no room to record CBLKCNT,
but it's easy to know the size of such pcluster is 1 lcluster as well.
Since Linux v6.1, each pcluster can be used for multiple variable-sized extents,
therefore it can be used for compressed data deduplication.
요약·해설
erofs.rst:1-369EROFS는 compact·extended inode, tail packing, xattr 공유, chunk deduplication, fixed-output compression, direct I/O·FSDAX와 Fscache distribution을 결합한 고성능 read-only filesystem입니다. metadata는 compact하지만 inode·xattr·directory·compression index의 주소식과 alignment를 명시적으로 유지합니다.
image 생성부터 runtime lookup·decompression·direct access까지의 큰 흐름입니다.