요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
Buffer Sharing and Synchronization (dma-buf)
============================================
The dma-buf subsystem provides the framework for sharing buffers for
hardware (DMA) access across multiple device drivers and subsystems, and
for synchronizing asynchronous hardware access.
As an example, it is used extensively by the DRM subsystem to exchange
buffers between processes, contexts, library APIs within the same
process, and also to exchange buffers with other subsystems such as
V4L2.
This document describes the way in which kernel subsystems can use and
interact with the three main primitives offered by dma-buf:
- dma-buf, representing a sg_table and exposed to userspace as a file
descriptor to allow passing between processes, subsystems, devices,
etc;
- dma-fence, providing a mechanism to signal when an asynchronous
hardware operation has completed; and
- dma-resv, which manages a set of dma-fences for a particular dma-buf
allowing implicit (kernel-ordered) synchronization of work to
preserve the illusion of coherent access
Userspace API principles and use
--------------------------------
For more details on how to design your subsystem's API for dma-buf use, please
see Documentation/userspace-api/dma-buf-alloc-exchange.rst.
Shared DMA Buffers
------------------
This document serves as a guide to device-driver writers on what is the dma-buf
buffer sharing API, how to use it for exporting and using shared buffers.
Any device driver which wishes to be a part of DMA buffer sharing, can do so as
either the 'exporter' of buffers, or the 'user' or 'importer' of buffers.
Say a driver A wants to use buffers created by driver B, then we call B as the
exporter, and A as buffer-user/importer.
The exporter
- implements and manages operations in :c:type:`struct dma_buf_ops
<dma_buf_ops>` for the buffer,
- allows other users to share the buffer by using dma_buf sharing APIs,
- manages the details of buffer allocation, wrapped in a :c:type:`struct
dma_buf <dma_buf>`,
- decides about the actual backing storage where this allocation happens,
- and takes care of any migration of scatterlist - for all (shared) users of
this buffer.
The buffer-user
- is one of (many) sharing users of the buffer.
- doesn't need to worry about how the buffer is allocated, or where.
- and needs a mechanism to get access to the scatterlist that makes up this
buffer in memory, mapped into its own address space, so it can access the
same area of memory. This interface is provided by :c:type:`struct
dma_buf_attachment <dma_buf_attachment>`.
Any exporters or users of the dma-buf buffer sharing framework must have a
'select DMA_SHARED_BUFFER' in their respective Kconfigs.
Userspace Interface Notes
~~~~~~~~~~~~~~~~~~~~~~~~~
Mostly a DMA buffer file descriptor is simply an opaque object for userspace,
and hence the generic interface exposed is very minimal. There's a few things to
consider though:
- Since kernel 3.12 the dma-buf FD supports the llseek system call, but only
with offset=0 and whence=SEEK_END|SEEK_SET. SEEK_SET is supported to allow
the usual size discover pattern size = SEEK_END(0); SEEK_SET(0). Every other
llseek operation will report -EINVAL.
If llseek on dma-buf FDs isn't supported the kernel will report -ESPIPE for all
cases. Userspace can use this to detect support for discovering the dma-buf
size using llseek.
- In order to avoid fd leaks on exec, the FD_CLOEXEC flag must be set
on the file descriptor. This is not just a resource leak, but a
potential security hole. It could give the newly exec'd application
access to buffers, via the leaked fd, to which it should otherwise
not be permitted access.
The problem with doing this via a separate fcntl() call, versus doing it
atomically when the fd is created, is that this is inherently racy in a
multi-threaded app[3]. The issue is made worse when it is library code
opening/creating the file descriptor, as the application may not even be
aware of the fd's.
To avoid this problem, userspace must have a way to request O_CLOEXEC
flag be set when the dma-buf fd is created. So any API provided by
the exporting driver to create a dmabuf fd must provide a way to let
userspace control setting of O_CLOEXEC flag passed in to dma_buf_fd().
- Memory mapping the contents of the DMA buffer is also supported. See the
discussion below on `CPU Access to DMA Buffer Objects`_ for the full details.
- The DMA buffer FD is also pollable, see `Implicit Fence Poll Support`_ below for
details.
- The DMA buffer FD also supports a few dma-buf-specific ioctls, see
`DMA Buffer ioctls`_ below for details.
Basic Operation and Device DMA Access
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: dma buf device access
CPU Access to DMA Buffer Objects
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: cpu access
Implicit Fence Poll Support
~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: implicit fence polling
DMA-BUF statistics
~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf-sysfs-stats.c
:doc: overview
DMA Buffer ioctls
~~~~~~~~~~~~~~~~~
.. kernel-doc:: include/uapi/linux/dma-buf.h
DMA-BUF locking convention
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: locking convention
Kernel Functions and Structures Reference
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:export:
.. kernel-doc:: include/linux/dma-buf.h
:internal:
Reservation Objects
-------------------
.. kernel-doc:: drivers/dma-buf/dma-resv.c
:doc: Reservation Object Overview
.. kernel-doc:: drivers/dma-buf/dma-resv.c
:export:
.. kernel-doc:: include/linux/dma-resv.h
:internal:
DMA Fences
----------
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: DMA fences overview
DMA Fence Cross-Driver Contract
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: fence cross-driver contract
DMA Fence Signalling Annotations
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: fence signalling annotation
DMA Fence Deadline Hints
~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: deadline hints
DMA Fences Functions Reference
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:export:
.. kernel-doc:: include/linux/dma-fence.h
:internal:
DMA Fence Array
~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence-array.c
:export:
.. kernel-doc:: include/linux/dma-fence-array.h
:internal:
DMA Fence Chain
~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence-chain.c
:export:
.. kernel-doc:: include/linux/dma-fence-chain.h
:internal:
DMA Fence unwrap
~~~~~~~~~~~~~~~~
.. kernel-doc:: include/linux/dma-fence-unwrap.h
:internal:
DMA Fence Sync File
~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/sync_file.c
:export:
.. kernel-doc:: include/linux/sync_file.h
:internal:
DMA Fence Sync File uABI
~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: include/uapi/linux/sync_file.h
:internal:
Indefinite DMA Fences
~~~~~~~~~~~~~~~~~~~~~
At various times struct dma_fence with an indefinite time until dma_fence_wait()
finishes have been proposed. Examples include:
* Future fences, used in HWC1 to signal when a buffer isn't used by the display
any longer, and created with the screen update that makes the buffer visible.
The time this fence completes is entirely under userspace's control.
* Proxy fences, proposed to handle &drm_syncobj for which the fence has not yet
been set. Used to asynchronously delay command submission.
* Userspace fences or gpu futexes, fine-grained locking within a command buffer
that userspace uses for synchronization across engines or with the CPU, which
are then imported as a DMA fence for integration into existing winsys
protocols.
* Long-running compute command buffers, while still using traditional end of
batch DMA fences for memory management instead of context preemption DMA
fences which get reattached when the compute job is rescheduled.
Common to all these schemes is that userspace controls the dependencies of these
fences and controls when they fire. Mixing indefinite fences with normal
in-kernel DMA fences does not work, even when a fallback timeout is included to
protect against malicious userspace:
* Only the kernel knows about all DMA fence dependencies, userspace is not aware
of dependencies injected due to memory management or scheduler decisions.
* Only userspace knows about all dependencies in indefinite fences and when
exactly they will complete, the kernel has no visibility.
Furthermore the kernel has to be able to hold up userspace command submission
for memory management needs, which means we must support indefinite fences being
dependent upon DMA fences. If the kernel also support indefinite fences in the
kernel like a DMA fence, like any of the above proposal would, there is the
potential for deadlocks.
.. kernel-render:: DOT
:alt: Indefinite Fencing Dependency Cycle
:caption: Indefinite Fencing Dependency Cycle
digraph "Fencing Cycle" {
node [shape=box bgcolor=grey style=filled]
kernel [label="Kernel DMA Fences"]
userspace [label="userspace controlled fences"]
kernel -> userspace [label="memory management"]
userspace -> kernel [label="Future fence, fence proxy, ..."]
{ rank=same; kernel userspace }
}
This means that the kernel might accidentally create deadlocks
through memory management dependencies which userspace is unaware of, which
randomly hangs workloads until the timeout kicks in. Workloads, which from
userspace's perspective, do not contain a deadlock. In such a mixed fencing
architecture there is no single entity with knowledge of all dependencies.
Therefore preventing such deadlocks from within the kernel is not possible.
The only solution to avoid dependencies loops is by not allowing indefinite
fences in the kernel. This means:
* No future fences, proxy fences or userspace fences imported as DMA fences,
with or without a timeout.
* No DMA fences that signal end of batchbuffer for command submission where
userspace is allowed to use userspace fencing or long running compute
workloads. This also means no implicit fencing for shared buffers in these
cases.
Recoverable Hardware Page Faults Implications
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Modern hardware supports recoverable page faults, which has a lot of
implications for DMA fences.
First, a pending page fault obviously holds up the work that's running on the
accelerator and a memory allocation is usually required to resolve the fault.
But memory allocations are not allowed to gate completion of DMA fences, which
means any workload using recoverable page faults cannot use DMA fences for
synchronization. Synchronization fences controlled by userspace must be used
instead.
On GPUs this poses a problem, because current desktop compositor protocols on
Linux rely on DMA fences, which means without an entirely new userspace stack
built on top of userspace fences, they cannot benefit from recoverable page
faults. Specifically this means implicit synchronization will not be possible.
The exception is when page faults are only used as migration hints and never to
on-demand fill a memory request. For now this means recoverable page
faults on GPUs are limited to pure compute workloads.
Furthermore GPUs usually have shared resources between the 3D rendering and
compute side, like compute units or command submission engines. If both a 3D
job with a DMA fence and a compute workload using recoverable page faults are
pending they could deadlock:
- The 3D workload might need to wait for the compute job to finish and release
hardware resources first.
- The compute workload might be stuck in a page fault, because the memory
allocation is waiting for the DMA fence of the 3D workload to complete.
There are a few options to prevent this problem, one of which drivers need to
ensure:
- Compute workloads can always be preempted, even when a page fault is pending
and not yet repaired. Not all hardware supports this.
- DMA fence workloads and workloads which need page fault handling have
independent hardware resources to guarantee forward progress. This could be
achieved through e.g. through dedicated engines and minimal compute unit
reservations for DMA fence workloads.
- The reservation approach could be further refined by only reserving the
hardware resources for DMA fence workloads when they are in-flight. This must
cover the time from when the DMA fence is visible to other threads up to
moment when fence is completed through dma_fence_signal().
- As a last resort, if the hardware provides no useful reservation mechanics,
all workloads must be flushed from the GPU when switching between jobs
requiring DMA fences or jobs requiring page fault handling: This means all DMA
fences must complete before a compute job with page fault handling can be
inserted into the scheduler queue. And vice versa, before a DMA fence can be
made visible anywhere in the system, all compute workloads must be preempted
to guarantee all pending GPU page faults are flushed.
- Only a fairly theoretical option would be to untangle these dependencies when
allocating memory to repair hardware page faults, either through separate
memory blocks or runtime tracking of the full dependency graph of all DMA
fences. This results very wide impact on the kernel, since resolving the page
on the CPU side can itself involve a page fault. It is much more feasible and
robust to limit the impact of handling hardware page faults to the specific
driver.
Note that workloads that run on independent hardware like copy engines or other
GPUs do not have any impact. This allows us to keep using DMA fences internally
in the kernel even for resolving hardware page faults, e.g. by using copy
engines to clear or copy memory needed to resolve the page fault.
In some ways this page fault problem is a special case of the `Infinite DMA
Fences` discussions: Infinite fences from compute workloads are allowed to
depend on DMA fences, but not the other way around. And not even the page fault
problem is new, because some other CPU thread in userspace might
hit a page fault which holds up a userspace fence - supporting page faults on
GPUs doesn't anything fundamentally new.
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
Buffer sharing과 synchronization
1-24dma-buf subsystem은 여러 device driver와 subsystem 사이에서 hardware(DMA) access용 buffer를 공유하고 asynchronous hardware access를 synchronize하는 framework를 제공합니다.
대표적으로 DRM subsystem은 process, context, 같은 process 안의 library API 사이에서 buffer를 교환하고 V4L2 같은 다른 subsystem과도 buffer를 교환하는 데 dma-buf를 널리 사용합니다.
kernel subsystem이 상호작용하는 세 가지 핵심 primitive는 다음과 같습니다.
- `dma-buf`: `sg_table`을 나타내며 userspace에는 file descriptor로 노출되어 process, subsystem, device 사이에서 전달됩니다.
- `dma-fence`: asynchronous hardware operation이 끝났음을 signal하는 mechanism입니다.
- `dma-resv`: 특정 dma-buf의 dma-fence set을 관리해 coherent access처럼 보이도록 work를 implicit, 즉 kernel-ordered 방식으로 synchronize합니다.
buffer identity, asynchronous completion, implicit synchronization의 역할을 분리했습니다.
Userspace API 원칙
25-32subsystem의 dma-buf userspace API를 설계하는 자세한 방법은 `Documentation/userspace-api/dma-buf-alloc-exchange.rst`를 참조합니다.
Userspace interface 주의사항
68-109DMA buffer file descriptor는 userspace에서 대체로 opaque object이므로 generic interface는 매우 작지만 다음 사항을 고려해야 합니다.
- kernel 3.12부터 dma-buf FD는 offset=0과 `SEEK_END|SEEK_SET`에 한해 `llseek`를 지원합니다. `size = SEEK_END(0); SEEK_SET(0)` pattern으로 size를 찾을 수 있고 다른 operation은 `-EINVAL`입니다. 지원하지 않는 kernel은 모든 경우 `-ESPIPE`을 반환하므로 userspace가 size discovery support를 감지할 수 있습니다.
- exec 때 fd leak을 막으려면 `FD_CLOEXEC`를 설정해야 합니다. leaked fd가 새 application에 허용되지 않은 buffer access를 줄 수 있어 security hole이 됩니다. multi-threaded application에서 별도 `fcntl()` 호출은 race가 있으므로 exporter API는 fd 생성 시 `dma_buf_fd()`에 전달할 `O_CLOEXEC`를 userspace가 제어하게 해야 합니다.
- DMA buffer content의 memory mapping도 지원합니다. `CPU Access to DMA Buffer Objects` section을 참조합니다.
- DMA buffer FD는 poll할 수 있습니다. `Implicit Fence Poll Support`를 참조합니다.
- DMA buffer FD는 dma-buf-specific ioctl도 지원합니다. `DMA Buffer ioctls`를 참조합니다.
size discovery, exec safety, mapping, polling, ioctl 지원을 정리했습니다.
Shared DMA buffer core API
110-152kernel-doc는 basic device DMA access, CPU access, implicit fence polling, sysfs statistics, ioctl, locking convention, exported function과 internal structure를 다음 source에서 가져옵니다.
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: dma buf device access
CPU Access to DMA Buffer Objects
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: cpu access
Implicit Fence Poll Support
~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: implicit fence polling
DMA-BUF statistics
~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf-sysfs-stats.c
:doc: overview
DMA Buffer ioctls
~~~~~~~~~~~~~~~~~
.. kernel-doc:: include/uapi/linux/dma-buf.h
DMA-BUF locking convention
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:doc: locking convention
Kernel Functions and Structures Reference
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-buf.c
:export:
.. kernel-doc:: include/linux/dma-buf.h
:internal:
core operation과 UAPI/statistics source를 역할별로 정리했습니다.
Reservation object
153-164`drivers/dma-buf/dma-resv.c`의 kernel-doc는 Reservation Object overview와 exported API를 제공하고 `include/linux/dma-resv.h`는 internal structure를 제공합니다.
.. kernel-doc:: drivers/dma-buf/dma-resv.c
:doc: Reservation Object Overview
.. kernel-doc:: drivers/dma-buf/dma-resv.c
:export:
.. kernel-doc:: include/linux/dma-resv.h
:internal:
DMA fence family와 sync file
165-236DMA fence documentation은 base fence overview, cross-driver contract, signalling annotation, deadline hint, exported function과 internal definition을 포함합니다. array와 chain은 여러 fence composition을 제공하고 unwrap helper는 wrapper를 해석합니다. sync file은 fence를 file descriptor UABI로 노출합니다.
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: DMA fences overview
DMA Fence Cross-Driver Contract
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: fence cross-driver contract
DMA Fence Signalling Annotations
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: fence signalling annotation
DMA Fence Deadline Hints
~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:doc: deadline hints
DMA Fences Functions Reference
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence.c
:export:
.. kernel-doc:: include/linux/dma-fence.h
:internal:
DMA Fence Array
~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence-array.c
:export:
.. kernel-doc:: include/linux/dma-fence-array.h
:internal:
DMA Fence Chain
~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/dma-fence-chain.c
:export:
.. kernel-doc:: include/linux/dma-fence-chain.h
:internal:
DMA Fence unwrap
~~~~~~~~~~~~~~~~
.. kernel-doc:: include/linux/dma-fence-unwrap.h
:internal:
DMA Fence Sync File
~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: drivers/dma-buf/sync_file.c
:export:
.. kernel-doc:: include/linux/sync_file.h
:internal:
DMA Fence Sync File uABI
~~~~~~~~~~~~~~~~~~~~~~~~
.. kernel-doc:: include/uapi/linux/sync_file.h
:internal:
- `drivers/dma-buf/dma-fence.c` / `include/linux/dma-fence.h`: base fence contract와 API
- `dma-fence-array.c` / `dma-fence-array.h`: fence array
- `dma-fence-chain.c` / `dma-fence-chain.h`: fence chain
- `dma-fence-unwrap.h`: unwrap internal helper
- `sync_file.c` / `sync_file.h`: sync file kernel interface
- `include/uapi/linux/sync_file.h`: sync file uABI
Indefinite DMA fence 제안과 근본 제약
237-275`dma_fence_wait()`가 언제 끝날지 정해지지 않은 `struct dma_fence`가 여러 번 제안되었습니다.
- Future fence: HWC1에서 display가 buffer 사용을 끝내는 시점을 signal하며 buffer를 visible하게 만드는 screen update와 함께 생성됩니다. completion time은 userspace가 완전히 제어합니다.
- Proxy fence: 아직 fence가 설정되지 않은 `drm_syncobj`를 처리해 command submission을 비동기로 지연하려는 제안입니다.
- Userspace fence 또는 GPU futex: userspace가 engine 사이 또는 CPU와 synchronize하는 command buffer 내부 fine-grained lock을 기존 winsys protocol과 통합하려고 DMA fence로 import합니다.
- Long-running compute command buffer: context preemption fence 대신 memory management에 traditional end-of-batch DMA fence를 계속 사용합니다.
공통점은 userspace가 dependency와 signal 시점을 제어한다는 것입니다. malicious userspace를 막는 timeout이 있어도 indefinite fence와 normal in-kernel DMA fence를 섞을 수 없습니다.
모든 DMA fence dependency는 kernel만 알고 memory management나 scheduler decision이 주입한 dependency를 userspace는 모릅니다. 반대로 indefinite fence의 모든 dependency와 정확한 completion 시점은 userspace만 알고 kernel은 볼 수 없습니다.
memory management를 위해 kernel이 userspace command submission을 멈출 수 있어야 하므로 indefinite fence가 DMA fence에 의존하는 것은 허용해야 합니다. kernel이 indefinite fence를 DMA fence처럼 지원하면 반대 dependency도 생겨 deadlock 가능성이 생깁니다.
Indefinite fencing dependency cycle
276-307.. kernel-render:: DOT
:alt: Indefinite Fencing Dependency Cycle
:caption: Indefinite Fencing Dependency Cycle
digraph "Fencing Cycle" {
node [shape=box bgcolor=grey style=filled]
kernel [label="Kernel DMA Fences"]
userspace [label="userspace controlled fences"]
kernel -> userspace [label="memory management"]
userspace -> kernel [label="Future fence, fence proxy, ..."]
{ rank=same; kernel userspace }
}
원문 DOT graph의 양방향 dependency를 동일한 두 edge로 재구성했습니다.
kernel은 userspace가 모르는 memory-management dependency를 우연히 만들어 timeout까지 workload를 hang시킬 수 있습니다. userspace 관점에는 deadlock이 없어도 mixed fencing architecture에는 전체 dependency를 아는 단일 entity가 없으므로 kernel 내부에서 예방할 수 없습니다.
dependency loop를 피하는 유일한 해법은 kernel에 indefinite fence를 허용하지 않는 것입니다.
- future fence, proxy fence, userspace fence를 timeout 유무와 관계없이 DMA fence로 import하지 않습니다.
- userspace fencing 또는 long-running compute workload를 허용하는 command submission에서는 end-of-batchbuffer DMA fence를 사용하지 않습니다. 이 경우 shared buffer의 implicit fencing도 사용할 수 없습니다.
Recoverable hardware page fault의 fence 영향
308-328modern hardware의 recoverable page fault는 DMA fence에 큰 영향을 줍니다. pending page fault는 accelerator work를 멈추고 fault 해결에는 보통 memory allocation이 필요합니다.
memory allocation이 DMA fence completion을 gate해서는 안 되므로 recoverable page fault를 사용하는 workload는 synchronization에 DMA fence를 사용할 수 없습니다. 대신 userspace-controlled synchronization fence를 사용해야 합니다.
GPU에서는 Linux desktop compositor protocol이 DMA fence에 의존하므로 userspace fence 기반의 완전히 새로운 stack 없이는 recoverable page fault의 이점을 얻지 못하고 implicit synchronization도 불가능합니다. page fault를 on-demand allocation이 아닌 migration hint로만 쓸 때는 예외입니다. 현재 GPU recoverable fault는 pure compute workload로 제한됩니다.
GPU shared resource deadlock과 forward progress
329-375GPU는 3D rendering과 compute가 compute unit 또는 command submission engine 같은 resource를 공유하는 경우가 많습니다. DMA fence를 가진 3D job과 recoverable page fault를 쓰는 compute workload가 함께 pending이면 deadlock이 생길 수 있습니다.
- 3D workload는 compute job이 끝나 hardware resource를 release할 때까지 기다릴 수 있습니다.
- compute workload는 page fault에 걸려 있고 memory allocation은 3D workload의 DMA fence completion을 기다릴 수 있습니다.
driver는 다음 option 가운데 하나로 forward progress를 보장해야 합니다.
- page fault가 pending이고 아직 repair되지 않았어도 compute workload를 항상 preempt할 수 있어야 합니다. 모든 hardware가 지원하지는 않습니다.
- DMA fence workload와 page-fault handling workload에 독립 hardware resource를 배정합니다. dedicated engine과 DMA fence workload용 최소 compute unit reservation으로 구현할 수 있습니다.
- DMA fence workload가 in-flight인 동안에만 resource를 reserve하도록 개선할 수 있습니다. fence가 다른 thread에 visible해진 때부터 `dma_fence_signal()`로 완료될 때까지를 포함해야 합니다.
- reservation mechanism이 없으면 두 job 유형을 전환할 때 GPU workload를 모두 flush합니다. page-fault compute job을 scheduler queue에 넣기 전 모든 DMA fence를 완료하고, DMA fence를 system에 visible하게 하기 전 모든 compute workload를 preempt해 pending GPU fault를 flush합니다.
- 별도 memory block 또는 모든 DMA fence dependency graph의 runtime tracking으로 fault repair allocation dependency를 풀 수 있지만 CPU-side fault까지 포함해 kernel 전체에 영향이 매우 큽니다. 특정 driver 안에서 hardware page-fault handling 영향을 제한하는 편이 훨씬 실현 가능하고 robust합니다.
preemption, resource partition, dynamic reservation, full flush, dependency untangling을 비교했습니다.
copy engine이나 다른 GPU처럼 독립 hardware에서 실행되는 workload는 영향을 주지 않습니다. 따라서 hardware page fault를 해결할 때도 copy engine으로 memory를 clear/copy하는 등 kernel 내부 DMA fence를 계속 사용할 수 있습니다.
Infinite fence discussion과의 관계
376-382이 page fault 문제는 `Infinite DMA Fences` discussion의 특수 사례입니다. compute workload의 infinite fence가 DMA fence에 의존하는 것은 허용되지만 DMA fence가 그 반대로 의존해서는 안 됩니다.
userspace의 다른 CPU thread도 userspace fence를 막는 page fault를 만날 수 있으므로 문제 자체가 새롭지는 않습니다. GPU page fault support가 근본적으로 새로운 dependency를 추가하는 것은 아닙니다.
요약과 해설
dma-buf.rst:1-382dma-buf는 exporter가 소유한 scatterlist-backed buffer를 file descriptor와 attachment로 여러 driver·process에 공유하고, dma-fence와 dma-resv로 asynchronous access를 synchronize합니다. FD creation의 `O_CLOEXEC`, CPU mapping, poll/ioctl contract와 exporter/importer 책임을 지켜야 합니다. userspace가 completion을 제어하는 indefinite fence를 in-kernel DMA fence와 섞으면 전체 dependency를 아는 주체가 없어 cycle이 생길 수 있으며, recoverable GPU page fault workload는 DMA fence와 shared hardware resource 사이 forward progress를 별도로 보장해야 합니다.