요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.
1. 요약·해설
원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.
2. 영어 원문 전체
번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.
원문 전체 펼치기
=========================
Dynamic DMA mapping Guide
=========================
:Author: David S. Miller <davem@redhat.com>
:Author: Richard Henderson <rth@cygnus.com>
:Author: Jakub Jelinek <jakub@redhat.com>
This is a guide to device driver writers on how to use the DMA API
with example pseudo-code. For a concise description of the API, see
Documentation/core-api/dma-api.rst.
CPU and DMA addresses
=====================
There are several kinds of addresses involved in the DMA API, and it's
important to understand the differences.
The kernel normally uses virtual addresses. Any address returned by
kmalloc(), vmalloc(), and similar interfaces is a virtual address and can
be stored in a ``void *``.
The virtual memory system (TLB, page tables, etc.) translates virtual
addresses to CPU physical addresses, which are stored as "phys_addr_t" or
"resource_size_t". The kernel manages device resources like registers as
physical addresses. These are the addresses in /proc/iomem. The physical
address is not directly useful to a driver; it must use ioremap() to map
the space and produce a virtual address.
I/O devices use a third kind of address: a "bus address". If a device has
registers at an MMIO address, or if it performs DMA to read or write system
memory, the addresses used by the device are bus addresses. In some
systems, bus addresses are identical to CPU physical addresses, but in
general they are not. IOMMUs and host bridges can produce arbitrary
mappings between physical and bus addresses.
From a device's point of view, DMA uses the bus address space, but it may
be restricted to a subset of that space. For example, even if a system
supports 64-bit addresses for main memory and PCI BARs, it may use an IOMMU
so devices only need to use 32-bit DMA addresses.
Here's a picture and some examples::
CPU CPU Bus
Virtual Physical Address
Address Address Space
Space Space
+-------+ +------+ +------+
| | |MMIO | Offset | |
| | Virtual |Space | applied | |
C +-------+ --------> B +------+ ----------> +------+ A
| | mapping | | by host | |
+-----+ | | | | bridge | | +--------+
| | | | +------+ | | | |
| CPU | | | | RAM | | | | Device |
| | | | | | | | | |
+-----+ +-------+ +------+ +------+ +--------+
| | Virtual |Buffer| Mapping | |
X +-------+ --------> Y +------+ <---------- +------+ Z
| | mapping | RAM | by IOMMU
| | | |
| | | |
+-------+ +------+
During the enumeration process, the kernel learns about I/O devices and
their MMIO space and the host bridges that connect them to the system. For
example, if a PCI device has a BAR, the kernel reads the bus address (A)
from the BAR and converts it to a CPU physical address (B). The address B
is stored in a struct resource and usually exposed via /proc/iomem. When a
driver claims a device, it typically uses ioremap() to map physical address
B at a virtual address (C). It can then use, e.g., ioread32(C), to access
the device registers at bus address A.
If the device supports DMA, the driver sets up a buffer using kmalloc() or
a similar interface, which returns a virtual address (X). The virtual
memory system maps X to a physical address (Y) in system RAM. The driver
can use virtual address X to access the buffer, but the device itself
cannot because DMA doesn't go through the CPU virtual memory system.
In some simple systems, the device can do DMA directly to physical address
Y. But in many others, there is IOMMU hardware that translates DMA
addresses to physical addresses, e.g., it translates Z to Y. This is part
of the reason for the DMA API: the driver can give a virtual address X to
an interface like dma_map_single(), which sets up any required IOMMU
mapping and returns the DMA address Z. The driver then tells the device to
do DMA to Z, and the IOMMU maps it to the buffer at address Y in system
RAM.
So that Linux can use the dynamic DMA mapping, it needs some help from the
drivers, namely it has to take into account that DMA addresses should be
mapped only for the time they are actually used and unmapped after the DMA
transfer.
The following API will work of course even on platforms where no such
hardware exists.
Note that the DMA API works with any bus independent of the underlying
microprocessor architecture. You should use the DMA API rather than the
bus-specific DMA API, i.e., use the dma_map_*() interfaces rather than the
pci_map_*() interfaces.
First of all, you should make sure::
#include <linux/dma-mapping.h>
is in your driver, which provides the definition of dma_addr_t. This type
can hold any valid DMA address for the platform and should be used
everywhere you hold a DMA address returned from the DMA mapping functions.
What memory is DMA'able?
========================
The first piece of information you must know is what kernel memory can
be used with the DMA mapping facilities. There has been an unwritten
set of rules regarding this, and this text is an attempt to finally
write them down.
If you acquired your memory via the page allocator
(i.e. __get_free_page*()) or the generic memory allocators
(i.e. kmalloc() or kmem_cache_alloc()) then you may DMA to/from
that memory using the addresses returned from those routines.
This means specifically that you may _not_ use the memory/addresses
returned from vmalloc() for DMA. It is possible to DMA to the
_underlying_ memory mapped into a vmalloc() area, but this requires
walking page tables to get the physical addresses, and then
translating each of those pages back to a kernel address using
something like __va(). [ EDIT: Update this when we integrate
Gerd Knorr's generic code which does this. ]
This rule also means that you may use neither kernel image addresses
(items in data/text/bss segments), nor module image addresses, nor
stack addresses for DMA. These could all be mapped somewhere entirely
different than the rest of physical memory. Even if those classes of
memory could physically work with DMA, you'd need to ensure the I/O
buffers were cacheline-aligned. Without that, you'd see cacheline
sharing problems (data corruption) on CPUs with DMA-incoherent caches.
(The CPU could write to one word, DMA would write to a different one
in the same cache line, and one of them could be overwritten.)
Also, this means that you cannot take the return of a kmap()
call and DMA to/from that. This is similar to vmalloc().
What about block I/O and networking buffers? The block I/O and
networking subsystems make sure that the buffers they use are valid
for you to DMA from/to.
DMA addressing capabilities
===========================
By default, the kernel assumes that your device can address 32-bits of DMA
addressing. For a 64-bit capable device, this needs to be increased, and for
a device with limitations, it needs to be decreased.
Special note about PCI: PCI-X specification requires PCI-X devices to support
64-bit addressing (DAC) for all transactions. And at least one platform (SGI
SN2) requires 64-bit coherent allocations to operate correctly when the IO
bus is in PCI-X mode.
For correct operation, you must set the DMA mask to inform the kernel about
your devices DMA addressing capabilities.
This is performed via a call to dma_set_mask_and_coherent()::
int dma_set_mask_and_coherent(struct device *dev, u64 mask);
which will set the mask for both streaming and coherent APIs together. If you
have some special requirements, then the following two separate calls can be
used instead:
The setup for streaming mappings is performed via a call to
dma_set_mask()::
int dma_set_mask(struct device *dev, u64 mask);
The setup for coherent allocations is performed via a call
to dma_set_coherent_mask()::
int dma_set_coherent_mask(struct device *dev, u64 mask);
Here, dev is a pointer to the device struct of your device, and mask is a bit
mask describing which bits of an address your device supports. Often the
device struct of your device is embedded in the bus-specific device struct of
your device. For example, &pdev->dev is a pointer to the device struct of a
PCI device (pdev is a pointer to the PCI device struct of your device).
These calls usually return zero to indicate your device can perform DMA
properly on the machine given the address mask you provided, but they might
return an error if the mask is too small to be supportable on the given
system. If it returns non-zero, your device cannot perform DMA properly on
this platform, and attempting to do so will result in undefined behavior.
You must not use DMA on this device unless the dma_set_mask family of
functions has returned success.
This means that in the failure case, you have two options:
1) Use some non-DMA mode for data transfer, if possible.
2) Ignore this device and do not initialize it.
It is recommended that your driver print a kernel KERN_WARNING message when
setting the DMA mask fails. In this manner, if a user of your driver reports
that performance is bad or that the device is not even detected, you can ask
them for the kernel messages to find out exactly why.
The 24-bit addressing device would do something like this::
if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(24))) {
dev_warn(dev, "mydev: No suitable DMA available\n");
goto ignore_this_device;
}
The standard 64-bit addressing device would do something like this::
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64))
dma_set_mask_and_coherent() never return fail when DMA_BIT_MASK(64). Typical
error code like::
/* Wrong code */
if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64)))
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32))
dma_set_mask_and_coherent() will never return failure when bigger than 32.
So typical code like::
/* Recommended code */
if (support_64bit)
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
else
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32));
If the device only supports 32-bit addressing for descriptors in the
coherent allocations, but supports full 64-bits for streaming mappings
it would look like this::
if (dma_set_mask(dev, DMA_BIT_MASK(64))) {
dev_warn(dev, "mydev: No suitable DMA available\n");
goto ignore_this_device;
}
The coherent mask will always be able to set the same or a smaller mask as
the streaming mask. However for the rare case that a device driver only
uses coherent allocations, one would have to check the return value from
dma_set_coherent_mask().
Finally, if your device can only drive the low 24-bits of
address you might do something like::
if (dma_set_mask(dev, DMA_BIT_MASK(24))) {
dev_warn(dev, "mydev: 24-bit DMA addressing not available\n");
goto ignore_this_device;
}
When dma_set_mask() or dma_set_mask_and_coherent() is successful, and
returns zero, the kernel saves away this mask you have provided. The
kernel will use this information later when you make DMA mappings.
There is a case which we are aware of at this time, which is worth
mentioning in this documentation. If your device supports multiple
functions (for example a sound card provides playback and record
functions) and the various different functions have _different_
DMA addressing limitations, you may wish to probe each mask and
only provide the functionality which the machine can handle. It
is important that the last call to dma_set_mask() be for the
most specific mask.
Here is pseudo-code showing how this might be done::
#define PLAYBACK_ADDRESS_BITS DMA_BIT_MASK(32)
#define RECORD_ADDRESS_BITS DMA_BIT_MASK(24)
struct my_sound_card *card;
struct device *dev;
...
if (!dma_set_mask(dev, PLAYBACK_ADDRESS_BITS)) {
card->playback_enabled = 1;
} else {
card->playback_enabled = 0;
dev_warn(dev, "%s: Playback disabled due to DMA limitations\n",
card->name);
}
if (!dma_set_mask(dev, RECORD_ADDRESS_BITS)) {
card->record_enabled = 1;
} else {
card->record_enabled = 0;
dev_warn(dev, "%s: Record disabled due to DMA limitations\n",
card->name);
}
A sound card was used as an example here because this genre of PCI
devices seems to be littered with ISA chips given a PCI front end,
and thus retaining the 16MB DMA addressing limitations of ISA.
Types of DMA mappings
=====================
There are two types of DMA mappings:
- Coherent DMA mappings which are usually mapped at driver
initialization, unmapped at the end and for which the hardware should
guarantee that the device and the CPU can access the data
in parallel and will see updates made by each other without any
explicit software flushing.
Think of "coherent" as "synchronous".
The current default is to return coherent memory in the low 32
bits of the DMA space. However, for future compatibility you should
set the coherent mask even if this default is fine for your
driver.
Good examples of what to use coherent mappings for are:
- Network card DMA ring descriptors.
- SCSI adapter mailbox command data structures.
- Device firmware microcode executed out of
main memory.
The invariant these examples all require is that any CPU store
to memory is immediately visible to the device, and vice
versa. Coherent mappings guarantee this.
.. important::
Coherent DMA memory does not preclude the usage of
proper memory barriers. The CPU may reorder stores to
coherent memory just as it may normal memory. Example:
if it is important for the device to see the first word
of a descriptor updated before the second, you must do
something like::
desc->word0 = address;
wmb();
desc->word1 = DESC_VALID;
in order to get correct behavior on all platforms.
Also, on some platforms your driver may need to flush CPU write
buffers in much the same way as it needs to flush write buffers
found in PCI bridges (such as by reading a register's value
after writing it).
- Streaming DMA mappings which are usually mapped for one DMA
transfer, unmapped right after it (unless you use dma_sync_* below)
and for which hardware can optimize for sequential accesses.
Think of "streaming" as "asynchronous" or "outside the coherency
domain".
Good examples of what to use streaming mappings for are:
- Networking buffers transmitted/received by a device.
- Filesystem buffers written/read by a SCSI device.
The interfaces for using this type of mapping were designed in
such a way that an implementation can make whatever performance
optimizations the hardware allows. To this end, when using
such mappings you must be explicit about what you want to happen.
Neither type of DMA mapping has alignment restrictions that come from
the underlying bus, although some devices may have such restrictions.
Also, systems with caches that aren't DMA-coherent will work better
when the underlying buffers don't share cache lines with other data.
Using Coherent DMA mappings
===========================
To allocate and map large (PAGE_SIZE or so) coherent DMA regions,
you should do::
dma_addr_t dma_handle;
cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, gfp);
where device is a ``struct device *``. This may be called in interrupt
context with the GFP_ATOMIC flag.
Size is the length of the region you want to allocate, in bytes.
This routine will allocate RAM for that region, so it acts similarly to
__get_free_pages() (but takes size instead of a page order). If your
driver needs regions sized smaller than a page, you may prefer using
the dma_pool interface, described below.
The coherent DMA mapping interfaces, will by default return a DMA address
which is 32-bit addressable. Even if the device indicates (via the DMA mask)
that it may address the upper 32-bits, coherent allocation will only
return > 32-bit addresses for DMA if the coherent DMA mask has been
explicitly changed via dma_set_coherent_mask(). This is true of the
dma_pool interface as well.
dma_alloc_coherent() returns two values: the virtual address which you
can use to access it from the CPU and dma_handle which you pass to the
card.
The CPU virtual address and the DMA address are both
guaranteed to be aligned to the smallest PAGE_SIZE order which
is greater than or equal to the requested size. This invariant
exists (for example) to guarantee that if you allocate a chunk
which is smaller than or equal to 64 kilobytes, the extent of the
buffer you receive will not cross a 64K boundary.
To unmap and free such a DMA region, you call::
dma_free_coherent(dev, size, cpu_addr, dma_handle);
where dev, size are the same as in the above call and cpu_addr and
dma_handle are the values dma_alloc_coherent() returned to you.
This function may not be called in interrupt context.
If your driver needs lots of smaller memory regions, you can write
custom code to subdivide pages returned by dma_alloc_coherent(),
or you can use the dma_pool API to do that. A dma_pool is like
a kmem_cache, but it uses dma_alloc_coherent(), not __get_free_pages().
Also, it understands common hardware constraints for alignment,
like queue heads needing to be aligned on N byte boundaries.
Create a dma_pool like this::
struct dma_pool *pool;
pool = dma_pool_create(name, dev, size, align, boundary);
The "name" is for diagnostics (like a kmem_cache name); dev and size
are as above. The device's hardware alignment requirement for this
type of data is "align" (which is expressed in bytes, and must be a
power of two). If your device has no boundary crossing restrictions,
pass 0 for boundary; passing 4096 says memory allocated from this pool
must not cross 4KByte boundaries (but at that time it may be better to
use dma_alloc_coherent() directly instead).
Allocate memory from a DMA pool like this::
cpu_addr = dma_pool_alloc(pool, flags, &dma_handle);
flags are GFP_KERNEL if blocking is permitted (not in_interrupt nor
holding SMP locks), GFP_ATOMIC otherwise. Like dma_alloc_coherent(),
this returns two values, cpu_addr and dma_handle.
Free memory that was allocated from a dma_pool like this::
dma_pool_free(pool, cpu_addr, dma_handle);
where pool is what you passed to dma_pool_alloc(), and cpu_addr and
dma_handle are the values dma_pool_alloc() returned. This function
may be called in interrupt context.
Destroy a dma_pool by calling::
dma_pool_destroy(pool);
Make sure you've called dma_pool_free() for all memory allocated
from a pool before you destroy the pool. This function may not
be called in interrupt context.
DMA Direction
=============
The interfaces described in subsequent portions of this document
take a DMA direction argument, which is an integer and takes on
one of the following values::
DMA_BIDIRECTIONAL
DMA_TO_DEVICE
DMA_FROM_DEVICE
DMA_NONE
You should provide the exact DMA direction if you know it.
DMA_TO_DEVICE means "from main memory to the device"
DMA_FROM_DEVICE means "from the device to main memory"
It is the direction in which the data moves during the DMA
transfer.
You are _strongly_ encouraged to specify this as precisely
as you possibly can.
If you absolutely cannot know the direction of the DMA transfer,
specify DMA_BIDIRECTIONAL. It means that the DMA can go in
either direction. The platform guarantees that you may legally
specify this, and that it will work, but this may be at the
cost of performance for example.
The value DMA_NONE is to be used for debugging. One can
hold this in a data structure before you come to know the
precise direction, and this will help catch cases where your
direction tracking logic has failed to set things up properly.
Another advantage of specifying this value precisely (outside of
potential platform-specific optimizations of such) is for debugging.
Some platforms actually have a write permission boolean which DMA
mappings can be marked with, much like page protections in the user
program address space. Such platforms can and do report errors in the
kernel logs when the DMA controller hardware detects violation of the
permission setting.
Only streaming mappings specify a direction, coherent mappings
implicitly have a direction attribute setting of
DMA_BIDIRECTIONAL.
The SCSI subsystem tells you the direction to use in the
'sc_data_direction' member of the SCSI command your driver is
working on.
For Networking drivers, it's a rather simple affair. For transmit
packets, map/unmap them with the DMA_TO_DEVICE direction
specifier. For receive packets, just the opposite, map/unmap them
with the DMA_FROM_DEVICE direction specifier.
Using Streaming DMA mappings
============================
The streaming DMA mapping routines can be called from interrupt
context. There are two versions of each map/unmap, one which will
map/unmap a single memory region, and one which will map/unmap a
scatterlist.
To map a single region, you do::
struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
void *addr = buffer->ptr;
size_t size = buffer->len;
dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
and to unmap it::
dma_unmap_single(dev, dma_handle, size, direction);
You should call dma_mapping_error() as dma_map_single() could fail and return
error. Doing so will ensure that the mapping code will work correctly on all
DMA implementations without any dependency on the specifics of the underlying
implementation. Using the returned address without checking for errors could
result in failures ranging from panics to silent data corruption. The same
applies to dma_map_page() as well.
You should call dma_unmap_single() when the DMA activity is finished, e.g.,
from the interrupt which told you that the DMA transfer is done.
Using CPU pointers like this for single mappings has a disadvantage:
you cannot reference HIGHMEM memory in this way. Thus, there is a
map/unmap interface pair akin to dma_{map,unmap}_single(). These
interfaces deal with page/offset pairs instead of CPU pointers.
Specifically::
struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
struct page *page = buffer->page;
unsigned long offset = buffer->offset;
size_t size = buffer->len;
dma_handle = dma_map_page(dev, page, offset, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
...
dma_unmap_page(dev, dma_handle, size, direction);
Here, "offset" means byte offset within the given page.
You should call dma_mapping_error() as dma_map_page() could fail and return
error as outlined under the dma_map_single() discussion.
You should call dma_unmap_page() when the DMA activity is finished, e.g.,
from the interrupt which told you that the DMA transfer is done.
With scatterlists, you map a region gathered from several regions by::
int i, count = dma_map_sg(dev, sglist, nents, direction);
struct scatterlist *sg;
for_each_sg(sglist, sg, count, i) {
hw_address[i] = sg_dma_address(sg);
hw_len[i] = sg_dma_len(sg);
}
where nents is the number of entries in the sglist.
The implementation is free to merge several consecutive sglist entries
into one (e.g. if DMA mapping is done with PAGE_SIZE granularity, any
consecutive sglist entries can be merged into one provided the first one
ends and the second one starts on a page boundary - in fact this is a huge
advantage for cards which either cannot do scatter-gather or have very
limited number of scatter-gather entries) and returns the actual number
of sg entries it mapped them to. On failure 0 is returned.
Then you should loop count times (note: this can be less than nents times)
and use sg_dma_address() and sg_dma_len() macros where you previously
accessed sg->address and sg->length as shown above.
To unmap a scatterlist, just call::
dma_unmap_sg(dev, sglist, nents, direction);
Again, make sure DMA activity has already finished.
.. note::
The 'nents' argument to the dma_unmap_sg call must be
the _same_ one you passed into the dma_map_sg call,
it should _NOT_ be the 'count' value _returned_ from the
dma_map_sg call.
Every dma_map_{single,sg}() call should have its dma_unmap_{single,sg}()
counterpart, because the DMA address space is a shared resource and
you could render the machine unusable by consuming all DMA addresses.
If you need to use the same streaming DMA region multiple times and touch
the data in between the DMA transfers, the buffer needs to be synced
properly in order for the CPU and device to see the most up-to-date and
correct copy of the DMA buffer.
So, firstly, just map it with dma_map_{single,sg}(), and after each DMA
transfer call either::
dma_sync_single_for_cpu(dev, dma_handle, size, direction);
or::
dma_sync_sg_for_cpu(dev, sglist, nents, direction);
as appropriate.
Then, if you wish to let the device get at the DMA area again,
finish accessing the data with the CPU, and then before actually
giving the buffer to the hardware call either::
dma_sync_single_for_device(dev, dma_handle, size, direction);
or::
dma_sync_sg_for_device(dev, sglist, nents, direction);
as appropriate.
.. note::
The 'nents' argument to dma_sync_sg_for_cpu() and
dma_sync_sg_for_device() must be the same passed to
dma_map_sg(). It is _NOT_ the count returned by
dma_map_sg().
After the last DMA transfer call one of the DMA unmap routines
dma_unmap_{single,sg}(). If you don't touch the data from the first
dma_map_*() call till dma_unmap_*(), then you don't have to call the
dma_sync_*() routines at all.
Here is pseudo code which shows a situation in which you would need
to use the dma_sync_*() interfaces::
my_card_setup_receive_buffer(struct my_card *cp, char *buffer, int len)
{
dma_addr_t mapping;
mapping = dma_map_single(cp->dev, buffer, len, DMA_FROM_DEVICE);
if (dma_mapping_error(cp->dev, mapping)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
cp->rx_buf = buffer;
cp->rx_len = len;
cp->rx_dma = mapping;
give_rx_buf_to_card(cp);
}
...
my_card_interrupt_handler(int irq, void *devid, struct pt_regs *regs)
{
struct my_card *cp = devid;
...
if (read_card_status(cp) == RX_BUF_TRANSFERRED) {
struct my_card_header *hp;
/* Examine the header to see if we wish
* to accept the data. But synchronize
* the DMA transfer with the CPU first
* so that we see updated contents.
*/
dma_sync_single_for_cpu(&cp->dev, cp->rx_dma,
cp->rx_len,
DMA_FROM_DEVICE);
/* Now it is safe to examine the buffer. */
hp = (struct my_card_header *) cp->rx_buf;
if (header_is_ok(hp)) {
dma_unmap_single(&cp->dev, cp->rx_dma, cp->rx_len,
DMA_FROM_DEVICE);
pass_to_upper_layers(cp->rx_buf);
make_and_setup_new_rx_buf(cp);
} else {
/* CPU should not write to
* DMA_FROM_DEVICE-mapped area,
* so dma_sync_single_for_device() is
* not needed here. It would be required
* for DMA_BIDIRECTIONAL mapping if
* the memory was modified.
*/
give_rx_buf_to_card(cp);
}
}
}
Handling Errors
===============
DMA address space is limited on some architectures and an allocation
failure can be determined by:
- checking if dma_alloc_coherent() returns NULL or dma_map_sg returns 0
- checking the dma_addr_t returned from dma_map_single() and dma_map_page()
by using dma_mapping_error()::
dma_addr_t dma_handle;
dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
- unmap pages that are already mapped, when mapping error occurs in the middle
of a multiple page mapping attempt. These example are applicable to
dma_map_page() as well.
Example 1::
dma_addr_t dma_handle1;
dma_addr_t dma_handle2;
dma_handle1 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle1)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling1;
}
dma_handle2 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle2)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling2;
}
...
map_error_handling2:
dma_unmap_single(dma_handle1);
map_error_handling1:
Example 2::
/*
* if buffers are allocated in a loop, unmap all mapped buffers when
* mapping error is detected in the middle
*/
dma_addr_t dma_addr;
dma_addr_t array[DMA_BUFFERS];
int save_index = 0;
for (i = 0; i < DMA_BUFFERS; i++) {
...
dma_addr = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_addr)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
array[i].dma_addr = dma_addr;
save_index++;
}
...
map_error_handling:
for (i = 0; i < save_index; i++) {
...
dma_unmap_single(array[i].dma_addr);
}
Networking drivers must call dev_kfree_skb() to free the socket buffer
and return NETDEV_TX_OK if the DMA mapping fails on the transmit hook
(ndo_start_xmit). This means that the socket buffer is just dropped in
the failure case.
SCSI drivers must return SCSI_MLQUEUE_HOST_BUSY if the DMA mapping
fails in the queuecommand hook. This means that the SCSI subsystem
passes the command to the driver again later.
Optimizing Unmap State Space Consumption
========================================
On many platforms, dma_unmap_{single,page}() is simply a nop.
Therefore, keeping track of the mapping address and length is a waste
of space. Instead of filling your drivers up with ifdefs and the like
to "work around" this (which would defeat the whole purpose of a
portable API) the following facilities are provided.
Actually, instead of describing the macros one by one, we'll
transform some example code.
1) Use DEFINE_DMA_UNMAP_{ADDR,LEN} in state saving structures.
Example, before::
struct ring_state {
struct sk_buff *skb;
dma_addr_t mapping;
__u32 len;
};
after::
struct ring_state {
struct sk_buff *skb;
DEFINE_DMA_UNMAP_ADDR(mapping);
DEFINE_DMA_UNMAP_LEN(len);
};
2) Use dma_unmap_{addr,len}_set() to set these values.
Example, before::
ringp->mapping = FOO;
ringp->len = BAR;
after::
dma_unmap_addr_set(ringp, mapping, FOO);
dma_unmap_len_set(ringp, len, BAR);
3) Use dma_unmap_{addr,len}() to access these values.
Example, before::
dma_unmap_single(dev, ringp->mapping, ringp->len,
DMA_FROM_DEVICE);
after::
dma_unmap_single(dev,
dma_unmap_addr(ringp, mapping),
dma_unmap_len(ringp, len),
DMA_FROM_DEVICE);
It really should be self-explanatory. We treat the ADDR and LEN
separately, because it is possible for an implementation to only
need the address in order to perform the unmap operation.
Platform Issues
===============
If you are just writing drivers for Linux and do not maintain
an architecture port for the kernel, you can safely skip down
to "Closing".
1) Struct scatterlist requirements.
You need to enable CONFIG_NEED_SG_DMA_LENGTH if the architecture
supports IOMMUs (including software IOMMU).
2) ARCH_DMA_MINALIGN
Architectures must ensure that kmalloc'ed buffer is
DMA-safe. Drivers and subsystems depend on it. If an architecture
isn't fully DMA-coherent (i.e. hardware doesn't ensure that data in
the CPU cache is identical to data in main memory),
ARCH_DMA_MINALIGN must be set so that the memory allocator
makes sure that kmalloc'ed buffer doesn't share a cache line with
the others. See arch/arm/include/asm/cache.h as an example.
Note that ARCH_DMA_MINALIGN is about DMA memory alignment
constraints. You don't need to worry about the architecture data
alignment constraints (e.g. the alignment constraints about 64-bit
objects).
Closing
=======
This document, and the API itself, would not be in its current
form without the feedback and suggestions from numerous individuals.
We would like to specifically mention, in no particular order, the
following people::
Russell King <rmk@arm.linux.org.uk>
Leo Dagum <dagum@barrel.engr.sgi.com>
Ralf Baechle <ralf@oss.sgi.com>
Grant Grundler <grundler@cup.hp.com>
Jay Estabrook <Jay.Estabrook@compaq.com>
Thomas Sailer <sailer@ife.ee.ethz.ch>
Andrea Arcangeli <andrea@suse.de>
Jens Axboe <jens.axboe@oracle.com>
David Mosberger-Tang <davidm@hpl.hp.com>
3. 한국어 전문 번역
영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.
Dynamic DMA mapping 안내서
1-12Dynamic DMA mapping Guide (동적 DMA mapping 안내서)
저자는 David S. Miller <davem@redhat.com>, Richard Henderson <rth@cygnus.com>, Jakub Jelinek <jakub@redhat.com>입니다.
이 문서는 device driver 작성자가 DMA API를 사용하는 방법을 예제 pseudo-code와 함께 설명합니다. API의 간결한 설명은 `Documentation/core-api/dma-api.rst`를 참고하십시오.
CPU 주소와 DMA 주소
13-110CPU 주소와 DMA 주소
DMA API에는 여러 종류의 주소가 관여하므로 그 차이를 이해하는 것이 중요합니다.
Kernel은 보통 virtual address를 사용합니다. `kmalloc()`, `vmalloc()` 및 비슷한 interface가 반환한 주소는 virtual address이며 `void *`에 저장할 수 있습니다.
Virtual memory system인 TLB와 page table 등은 virtual address를 CPU physical address로 변환합니다. CPU physical address는 `phys_addr_t` 또는 `resource_size_t`에 저장됩니다. Kernel은 register 같은 device resource를 physical address로 관리하고 `/proc/iomem`에 노출합니다. Physical address는 driver가 직접 사용할 수 없으므로 `ioremap()`으로 공간을 mapping해 virtual address를 만들어야 합니다.
I/O device는 세 번째 종류인 bus address를 사용합니다. Device의 register가 MMIO address에 있거나 device가 DMA로 system memory를 읽고 쓸 때 사용하는 주소가 bus address입니다. 일부 시스템에서는 bus address와 CPU physical address가 같지만 일반적으로는 다르며, IOMMU와 host bridge가 둘 사이에 임의의 mapping을 만들 수 있습니다.
Device 관점에서 DMA는 bus address space를 사용하지만 그 일부로 제한될 수 있습니다. 예를 들어 main memory와 PCI BAR가 64-bit address를 지원해도 IOMMU를 사용하면 device는 32-bit DMA address만 다루면 됩니다.
MMIO 접근은 bus address A에서 host bridge offset을 거쳐 CPU physical address B로, 다시 virtual mapping을 거쳐 CPU virtual address C로 이어집니다. DMA buffer는 CPU virtual address X가 physical RAM address Y로 mapping되고, IOMMU가 device의 DMA address Z를 Y로 변환합니다.
Enumeration 과정에서 kernel은 I/O device, MMIO space, 시스템에 연결하는 host bridge를 파악합니다. PCI device에 BAR가 있으면 kernel은 BAR의 bus address A를 읽어 CPU physical address B로 변환합니다. B는 `struct resource`에 저장되고 보통 `/proc/iomem`에 표시됩니다. Driver가 device를 점유하면 `ioremap()`으로 B를 virtual address C에 mapping하고, 예를 들어 `ioread32(C)`로 bus address A의 register에 접근합니다.
Device가 DMA를 지원하면 driver는 `kmalloc()` 같은 interface로 virtual address X를 반환받아 buffer를 준비합니다. Virtual memory system은 X를 system RAM의 physical address Y로 mapping합니다. Driver는 X로 buffer에 접근하지만 DMA는 CPU virtual memory system을 통과하지 않으므로 device는 X를 사용할 수 없습니다.
단순한 시스템에서는 device가 Y에 직접 DMA할 수 있습니다. 많은 시스템에서는 IOMMU가 DMA address를 physical address로 변환하며, 예를 들어 Z를 Y로 mapping합니다. 이것이 DMA API가 필요한 이유입니다. Driver가 virtual address X를 `dma_map_single()` 같은 interface에 주면 필요한 IOMMU mapping을 만들고 DMA address Z를 반환합니다. Driver는 device에 Z로 DMA하도록 지시하고 IOMMU는 이를 system RAM의 Y buffer로 연결합니다.
Linux의 dynamic DMA mapping을 사용하려면 driver가 DMA address를 실제 사용하는 동안에만 mapping하고 DMA transfer 뒤에는 unmap해야 합니다. 아래 API는 이런 hardware가 없는 platform에서도 동작합니다.
DMA API는 microprocessor architecture와 무관하게 모든 bus에서 동작합니다. Bus 전용 DMA API인 `pci_map_*()` 대신 `dma_map_*()` interface를 사용해야 합니다.
먼저 driver에 다음 header가 포함되어 있는지 확인합니다.
#include <linux/dma-mapping.h>
이 header는 `dma_addr_t`를 정의합니다. 이 type은 platform에서 유효한 모든 DMA address를 담을 수 있으므로 DMA mapping function이 반환한 DMA address를 보관하는 모든 곳에서 사용해야 합니다.
DMA에 사용할 수 있는 메모리
111-148어떤 메모리를 DMA에 사용할 수 있는가?
먼저 DMA mapping 기능에 사용할 수 있는 kernel memory를 알아야 합니다. 과거에는 기록되지 않은 규칙이었으며 이 절에서 이를 명문화합니다.
Page allocator인 `__get_free_page*()` 또는 generic memory allocator인 `kmalloc()`, `kmem_cache_alloc()`으로 얻은 메모리는 해당 routine이 반환한 주소로 DMA 송수신에 사용할 수 있습니다.
`vmalloc()`이 반환한 메모리나 주소는 DMA에 사용할 수 없습니다. `vmalloc()` 영역에 mapping된 실제 하부 메모리에 DMA하는 것은 가능하지만, page table을 순회해 physical address를 얻고 각 page를 `__va()` 같은 방법으로 다시 kernel address로 변환해야 합니다. 원문의 편집 메모리는 Gerd Knorr의 generic code가 통합되면 이 설명을 갱신하라고 적고 있습니다.
Kernel image의 data/text/bss segment 주소, module image 주소, stack 주소 역시 DMA에 사용할 수 없습니다. 이들은 나머지 physical memory와 완전히 다른 곳에 mapping될 수 있습니다. 물리적으로 DMA가 가능하더라도 I/O buffer를 cacheline에 맞춰 정렬해야 합니다. 그렇지 않으면 DMA-incoherent cache를 쓰는 CPU에서 cacheline sharing으로 data corruption이 생길 수 있습니다. CPU와 DMA가 같은 cache line의 서로 다른 word를 쓰면 둘 중 하나가 덮어써질 수 있습니다.
`kmap()` 반환값도 DMA 송수신에 사용할 수 없습니다. 이는 `vmalloc()`과 같은 종류의 제약입니다.
Block I/O와 networking subsystem은 자신들이 제공하는 buffer가 DMA 송수신에 유효하도록 보장합니다.
DMA 주소 지정 능력과 mask 설정
149-205DMA 주소 지정 능력
기본적으로 kernel은 device가 32-bit DMA address를 지정할 수 있다고 가정합니다. 64-bit device는 범위를 늘려야 하고 제약이 있는 device는 줄여야 합니다.
PCI에 관한 특별한 주의 사항으로, PCI-X 명세는 PCI-X device가 모든 transaction에서 64-bit addressing인 DAC를 지원하도록 요구합니다. 적어도 SGI SN2 platform은 I/O bus가 PCI-X mode일 때 올바르게 동작하려면 64-bit coherent allocation이 필요합니다.
올바른 동작을 위해 DMA mask를 설정하여 device의 DMA addressing capability를 kernel에 알려야 합니다. Streaming과 coherent API의 mask를 함께 설정하려면 다음을 호출합니다.
int dma_set_mask_and_coherent(struct device *dev, u64 mask);
특별한 요구가 있으면 두 API를 나누어 설정할 수 있습니다. Streaming mapping에는 다음을 사용합니다.
int dma_set_mask(struct device *dev, u64 mask);
Coherent allocation에는 다음을 사용합니다.
int dma_set_coherent_mask(struct device *dev, u64 mask);
`dev`는 device의 `struct device` pointer이고 `mask`는 device가 지원하는 address bit를 나타내는 bit mask입니다. `struct device`는 보통 bus 전용 device structure 안에 들어 있습니다. 예를 들어 PCI device structure pointer가 `pdev`라면 `&pdev->dev`가 해당 device structure의 pointer입니다.
이 함수들은 주어진 address mask로 해당 시스템에서 DMA를 올바르게 수행할 수 있으면 보통 0을 반환합니다. Mask가 시스템에서 지원하기에 너무 작으면 오류를 반환할 수 있습니다. 0이 아니면 이 platform에서 DMA를 수행할 수 없고 시도 결과는 undefined behavior입니다. `dma_set_mask` 계열이 성공하기 전에는 해당 device에서 DMA를 사용하면 안 됩니다.
실패하면 가능한 선택은 두 가지입니다.
- 가능하다면 data transfer에 DMA가 아닌 mode를 사용합니다.
- 해당 device를 무시하고 초기화하지 않습니다.
DMA mask 설정 실패 시 driver가 kernel `KERN_WARNING` message를 출력하는 것이 좋습니다. 사용자가 성능 저하나 device 미검출을 보고할 때 kernel message를 통해 정확한 원인을 확인할 수 있습니다.
DMA mask 설정 예제와 다기능 device
206-29524-bit addressing device는 다음처럼 설정합니다.
if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(24))) {
dev_warn(dev, "mydev: No suitable DMA available\n");
goto ignore_this_device;
}
표준 64-bit addressing device는 다음과 같이 설정합니다.
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64))
`dma_set_mask_and_coherent()`는 `DMA_BIT_MASK(64)`에서 실패를 반환하지 않습니다. 따라서 다음과 같은 fallback code는 잘못되었습니다.
/* Wrong code */
if (dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64)))
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32))
`dma_set_mask_and_coherent()`는 32-bit보다 큰 mask에서 실패하지 않으므로 device의 capability를 알고 다음과 같이 선택하는 것이 권장됩니다.
/* Recommended code */
if (support_64bit)
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(64));
else
dma_set_mask_and_coherent(dev, DMA_BIT_MASK(32));
Device가 coherent allocation의 descriptor에는 32-bit addressing만 지원하지만 streaming mapping에는 전체 64-bit를 지원한다면 streaming mask를 다음처럼 설정합니다.
if (dma_set_mask(dev, DMA_BIT_MASK(64))) {
dev_warn(dev, "mydev: No suitable DMA available\n");
goto ignore_this_device;
}
Coherent mask는 항상 streaming mask와 같거나 더 작은 mask로 설정할 수 있습니다. 드물게 coherent allocation만 사용하는 driver라면 `dma_set_coherent_mask()`의 반환값을 검사해야 합니다.
Device가 주소의 낮은 24-bit만 구동한다면 다음처럼 설정합니다.
if (dma_set_mask(dev, DMA_BIT_MASK(24))) {
dev_warn(dev, "mydev: 24-bit DMA addressing not available\n");
goto ignore_this_device;
}
`dma_set_mask()` 또는 `dma_set_mask_and_coherent()`가 성공해 0을 반환하면 kernel은 제공된 mask를 저장하고 이후 DMA mapping을 만들 때 사용합니다.
Sound card의 playback과 record처럼 여러 기능의 DMA addressing limit가 서로 다르면 각 mask를 probe하고 시스템이 처리할 수 있는 기능만 제공할 수 있습니다. 마지막 `dma_set_mask()` 호출에는 가장 구체적인 mask를 사용해야 합니다.
다음 pseudo-code가 그 방법을 보여 줍니다.
#define PLAYBACK_ADDRESS_BITS DMA_BIT_MASK(32)
#define RECORD_ADDRESS_BITS DMA_BIT_MASK(24)
struct my_sound_card *card;
struct device *dev;
...
if (!dma_set_mask(dev, PLAYBACK_ADDRESS_BITS)) {
card->playback_enabled = 1;
} else {
card->playback_enabled = 0;
dev_warn(dev, "%s: Playback disabled due to DMA limitations\n",
card->name);
}
if (!dma_set_mask(dev, RECORD_ADDRESS_BITS)) {
card->record_enabled = 1;
} else {
card->record_enabled = 0;
dev_warn(dev, "%s: Record disabled due to DMA limitations\n",
card->name);
}
Sound card를 예로 든 이유는 이런 PCI device에 PCI front end를 붙인 ISA chip이 흔하여 ISA의 16MB DMA addressing limit를 그대로 유지하는 경우가 많기 때문입니다.
Coherent와 streaming DMA mapping
296-367DMA mapping의 종류
DMA mapping은 `Coherent DMA mappings`와 `Streaming DMA mappings` 두 종류입니다.
Coherent DMA mapping은 보통 driver initialization 때 mapping하고 종료 때 unmap합니다. Hardware는 device와 CPU가 data에 병렬로 접근하면서 명시적인 software flush 없이도 서로의 update를 볼 수 있도록 보장해야 합니다. `coherent`를 `synchronous`로 이해하면 됩니다.
현재 기본값은 DMA space의 낮은 32-bit에서 coherent memory를 반환하는 것입니다. Driver가 이 기본값으로 충분하더라도 향후 호환성을 위해 coherent mask를 설정해야 합니다.
Coherent mapping의 좋은 사용 예는 다음과 같습니다.
- Network card DMA ring descriptor
- SCSI adapter mailbox command data structure
- Main memory에서 실행되는 device firmware microcode
이 예들의 공통 불변 조건은 CPU의 memory store가 device에 즉시 보이고 그 반대도 성립해야 한다는 것입니다. Coherent mapping이 이를 보장합니다.
Coherent DMA memory에서도 올바른 memory barrier가 필요합니다. CPU는 일반 memory와 마찬가지로 coherent memory에 대한 store 순서를 바꿀 수 있습니다. Device가 descriptor의 첫 word 갱신을 두 번째 word보다 먼저 봐야 한다면 다음 순서를 사용해야 합니다.
desc->word0 = address;
wmb();
desc->word1 = DESC_VALID;
이렇게 해야 모든 platform에서 올바르게 동작합니다. 일부 platform에서는 PCI bridge의 write buffer를 flush하듯 register를 쓴 뒤 값을 읽는 방식 등으로 CPU write buffer도 flush해야 할 수 있습니다.
Streaming DMA mapping은 보통 한 번의 DMA transfer를 위해 mapping하고 직후 unmap합니다. 아래의 `dma_sync_*`를 쓰는 경우는 예외입니다. Hardware는 sequential access에 맞게 최적화할 수 있습니다. `streaming`을 `asynchronous` 또는 `coherency domain 밖`으로 이해하면 됩니다.
Streaming mapping의 좋은 사용 예는 다음과 같습니다.
- Device가 송수신하는 networking buffer
- SCSI device가 읽고 쓰는 filesystem buffer
Streaming interface는 구현이 hardware가 허용하는 모든 성능 최적화를 할 수 있도록 설계되었습니다. 따라서 사용자는 원하는 동작을 명시해야 합니다.
두 종류 모두 하부 bus에서 비롯된 alignment 제한은 없지만 device 자체의 제한은 있을 수 있습니다. DMA-coherent가 아닌 cache를 쓰는 시스템에서는 buffer가 다른 data와 cache line을 공유하지 않을 때 더 잘 동작합니다.
Coherent DMA mapping 사용법
368-458Coherent DMA mapping 사용
`PAGE_SIZE` 정도의 큰 coherent DMA region을 allocate하고 mapping하려면 다음과 같이 합니다.
dma_addr_t dma_handle;
cpu_addr = dma_alloc_coherent(dev, size, &dma_handle, gfp);
`dev`는 `struct device *`입니다. `GFP_ATOMIC` flag를 사용하면 interrupt context에서 호출할 수 있습니다. `size`는 allocate할 region의 byte 길이입니다.
이 routine은 해당 region의 RAM을 allocate하므로 page order 대신 size를 받는 `__get_free_pages()`와 비슷합니다. 한 page보다 작은 region이 필요하면 아래의 `dma_pool` interface가 더 적합할 수 있습니다.
Coherent DMA mapping interface는 기본적으로 32-bit로 address 가능한 DMA address를 반환합니다. Device가 DMA mask로 상위 32-bit 접근을 표시해도 `dma_set_coherent_mask()`로 coherent DMA mask를 명시적으로 바꾼 경우에만 32-bit보다 큰 DMA address를 반환합니다. `dma_pool`에도 같은 규칙이 적용됩니다.
`dma_alloc_coherent()`는 CPU에서 접근할 virtual address와 card에 전달할 `dma_handle` 두 값을 반환합니다.
CPU virtual address와 DMA address는 모두 요청 size 이상인 가장 작은 `PAGE_SIZE` order에 맞춰 정렬됩니다. 예를 들어 64KB 이하 chunk를 allocate하면 반환된 buffer 범위가 64K boundary를 넘지 않습니다.
이 DMA region을 unmap하고 free하려면 `dma_free_coherent()`를 호출합니다.
dma_free_coherent(dev, size, cpu_addr, dma_handle);
`dev`와 `size`는 allocate 때와 같고 `cpu_addr`와 `dma_handle`은 `dma_alloc_coherent()`의 반환값입니다. 이 함수는 interrupt context에서 호출할 수 없습니다.
작은 memory region이 많이 필요하면 `dma_alloc_coherent()`가 반환한 page를 직접 나누거나 `dma_pool` API를 사용할 수 있습니다. `dma_pool`은 `kmem_cache`와 비슷하지만 `__get_free_pages()` 대신 `dma_alloc_coherent()`를 사용하며, queue head가 N-byte boundary에 정렬되어야 하는 것 같은 hardware alignment constraint도 이해합니다.
다음과 같이 `dma_pool`을 만듭니다.
struct dma_pool *pool;
pool = dma_pool_create(name, dev, size, align, boundary);
`name`은 `kmem_cache` 이름처럼 진단에 쓰입니다. `dev`와 `size`는 위와 같습니다. `align`은 이 data type에 대한 device hardware의 byte 단위 alignment requirement이며 2의 거듭제곱이어야 합니다. Boundary crossing 제한이 없으면 `boundary`에 0을 전달합니다. 4096은 pool allocation이 4KByte boundary를 넘지 못한다는 뜻이며, 이 경우에는 `dma_alloc_coherent()`를 직접 쓰는 편이 나을 수 있습니다.
DMA pool에서 memory를 allocate하려면 다음을 사용합니다.
cpu_addr = dma_pool_alloc(pool, flags, &dma_handle);
Blocking이 허용되면, 즉 interrupt 안이나 SMP lock 보유 중이 아니라면 `flags`는 `GFP_KERNEL`이고 그 밖에는 `GFP_ATOMIC`입니다. `dma_alloc_coherent()`처럼 `cpu_addr`와 `dma_handle`을 반환합니다.
Pool에서 allocate한 memory는 다음과 같이 free합니다.
dma_pool_free(pool, cpu_addr, dma_handle);
`pool`은 `dma_pool_alloc()`에 전달한 값이고 나머지는 그 함수가 반환한 값입니다. 이 함수는 interrupt context에서 호출할 수 있습니다.
DMA pool은 다음과 같이 destroy합니다.
dma_pool_destroy(pool);
Pool을 destroy하기 전에 그 pool에서 allocate한 모든 memory에 `dma_pool_free()`를 호출했는지 확인해야 합니다. `dma_pool_destroy()`는 interrupt context에서 호출할 수 없습니다.
DMA 전송 방향
459-512DMA 방향
이 문서 뒤에서 설명하는 interface는 정수형 DMA direction argument로 다음 값 중 하나를 받습니다.
DMA_BIDIRECTIONAL
DMA_TO_DEVICE
DMA_FROM_DEVICE
DMA_NONE
방향을 안다면 정확히 지정해야 합니다. `DMA_TO_DEVICE`는 main memory에서 device로, `DMA_FROM_DEVICE`는 device에서 main memory로 data가 이동한다는 뜻이며 DMA transfer의 실제 이동 방향을 가리킵니다.
가능한 한 정확한 방향을 지정할 것을 강하게 권장합니다.
DMA transfer 방향을 정말 알 수 없다면 양방향을 뜻하는 `DMA_BIDIRECTIONAL`을 지정합니다. Platform은 이 값이 합법적이며 동작함을 보장하지만 성능 비용이 들 수 있습니다.
`DMA_NONE`은 debugging에 사용합니다. 정확한 방향을 알기 전 data structure에 넣어 두면 direction tracking logic이 올바르게 설정하지 못한 경우를 잡는 데 도움이 됩니다.
정확한 방향은 platform 전용 최적화뿐 아니라 debugging에도 유리합니다. 일부 platform은 user address space의 page protection처럼 DMA mapping에 write permission boolean을 표시할 수 있고, DMA controller가 permission 위반을 감지하면 kernel log에 오류를 보고합니다.
Direction을 지정하는 것은 streaming mapping뿐입니다. Coherent mapping은 암묵적으로 `DMA_BIDIRECTIONAL` direction attribute를 가집니다.
SCSI subsystem은 driver가 처리 중인 SCSI command의 `sc_data_direction` member로 사용할 방향을 알려 줍니다.
Networking driver에서는 transmit packet을 `DMA_TO_DEVICE`로 map/unmap하고 receive packet은 반대로 `DMA_FROM_DEVICE`로 map/unmap합니다.
Streaming single 및 page mapping
513-585Streaming DMA mapping 사용
Streaming DMA mapping routine은 interrupt context에서 호출할 수 있습니다. 각 map/unmap에는 단일 memory region용과 scatterlist용 두 형태가 있습니다.
단일 region은 다음과 같이 mapping합니다.
struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
void *addr = buffer->ptr;
size_t size = buffer->len;
dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
다음과 같이 unmap합니다.
dma_unmap_single(dev, dma_handle, size, direction);
`dma_map_single()`은 실패해 오류를 반환할 수 있으므로 `dma_mapping_error()`를 호출해야 합니다. 그래야 하부 구현 세부 사항에 의존하지 않고 모든 DMA 구현에서 올바르게 동작합니다. 오류 검사 없이 반환 주소를 사용하면 panic부터 조용한 data corruption까지 발생할 수 있습니다. `dma_map_page()`에도 같은 규칙이 적용됩니다.
DMA activity가 끝나면, 예를 들어 DMA transfer 완료를 알린 interrupt에서 `dma_unmap_single()`을 호출해야 합니다.
CPU pointer를 쓰는 single mapping은 HIGHMEM memory를 참조할 수 없다는 단점이 있습니다. 이를 위해 `dma_{map,unmap}_single()`과 비슷하지만 CPU pointer 대신 page/offset pair를 다루는 interface가 있습니다.
struct device *dev = &my_dev->dev;
dma_addr_t dma_handle;
struct page *page = buffer->page;
unsigned long offset = buffer->offset;
size_t size = buffer->len;
dma_handle = dma_map_page(dev, page, offset, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
...
dma_unmap_page(dev, dma_handle, size, direction);
`offset`은 주어진 page 안의 byte offset입니다. `dma_map_page()`도 실패할 수 있으므로 `dma_mapping_error()`를 호출해야 하며, DMA activity가 끝나면 `dma_unmap_page()`를 호출합니다.
Streaming scatterlist mapping
586-626여러 region에서 모은 영역은 scatterlist로 다음처럼 mapping합니다.
int i, count = dma_map_sg(dev, sglist, nents, direction);
struct scatterlist *sg;
for_each_sg(sglist, sg, count, i) {
hw_address[i] = sg_dma_address(sg);
hw_len[i] = sg_dma_len(sg);
}
`nents`는 `sglist` entry 수입니다. 구현은 연속된 여러 entry를 하나로 합칠 수 있습니다. 예를 들어 `PAGE_SIZE` 단위로 mapping할 때 앞 entry가 page boundary에서 끝나고 다음 entry가 그 boundary에서 시작하면 합칠 수 있습니다. Scatter-gather를 지원하지 않거나 entry 수가 매우 제한된 card에 큰 장점입니다. 반환값은 실제 mapping된 sg entry 수이며 실패하면 0입니다.
그 뒤에는 `nents`가 아니라 반환된 `count`만큼 순회하며, 기존의 `sg->address`, `sg->length` 대신 `sg_dma_address()`와 `sg_dma_len()` macro를 사용합니다.
Scatterlist는 다음과 같이 unmap합니다.
dma_unmap_sg(dev, sglist, nents, direction);
이때도 DMA activity가 이미 끝났는지 확인해야 합니다.
`dma_unmap_sg()`의 `nents` argument는 `dma_map_sg()`에 전달했던 값과 같아야 합니다. `dma_map_sg()`가 반환한 `count`를 사용하면 안 됩니다.
DMA address space는 공유 자원이므로 모든 `dma_map_{single,sg}()` 호출에는 대응하는 `dma_unmap_{single,sg}()`가 있어야 합니다. DMA address를 모두 소비하면 시스템을 사용할 수 없게 만들 수 있습니다.
반복 사용하는 streaming buffer 동기화
627-729같은 streaming DMA region을 여러 번 사용하고 transfer 사이에 data를 만진다면 CPU와 device가 최신의 올바른 DMA buffer 사본을 보도록 buffer를 정확히 동기화해야 합니다.
먼저 `dma_map_{single,sg}()`로 mapping하고 각 DMA transfer 뒤 CPU가 접근하기 전에 상황에 맞는 다음 함수 중 하나를 호출합니다.
dma_sync_single_for_cpu(dev, dma_handle, size, direction);
dma_sync_sg_for_cpu(dev, sglist, nents, direction);
Device가 DMA area에 다시 접근하게 하려면 CPU 접근을 마친 뒤 buffer를 hardware에 넘기기 전에 다음 중 맞는 함수를 호출합니다.
dma_sync_single_for_device(dev, dma_handle, size, direction);
dma_sync_sg_for_device(dev, sglist, nents, direction);
`dma_sync_sg_for_cpu()`와 `dma_sync_sg_for_device()`의 `nents`는 `dma_map_sg()`에 전달한 값과 같아야 하며, `dma_map_sg()`가 반환한 `count`가 아닙니다.
마지막 DMA transfer 뒤에는 `dma_unmap_{single,sg}()` 중 하나를 호출합니다. 첫 `dma_map_*()`부터 `dma_unmap_*()`까지 CPU가 data를 건드리지 않는다면 `dma_sync_*()`를 호출할 필요가 없습니다.
다음 pseudo-code는 `dma_sync_*()` interface가 필요한 상황을 보여 줍니다.
my_card_setup_receive_buffer(struct my_card *cp, char *buffer, int len)
{
dma_addr_t mapping;
mapping = dma_map_single(cp->dev, buffer, len, DMA_FROM_DEVICE);
if (dma_mapping_error(cp->dev, mapping)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
cp->rx_buf = buffer;
cp->rx_len = len;
cp->rx_dma = mapping;
give_rx_buf_to_card(cp);
}
...
my_card_interrupt_handler(int irq, void *devid, struct pt_regs *regs)
{
struct my_card *cp = devid;
...
if (read_card_status(cp) == RX_BUF_TRANSFERRED) {
struct my_card_header *hp;
/* Examine the header to see if we wish
* to accept the data. But synchronize
* the DMA transfer with the CPU first
* so that we see updated contents.
*/
dma_sync_single_for_cpu(&cp->dev, cp->rx_dma,
cp->rx_len,
DMA_FROM_DEVICE);
/* Now it is safe to examine the buffer. */
hp = (struct my_card_header *) cp->rx_buf;
if (header_is_ok(hp)) {
dma_unmap_single(&cp->dev, cp->rx_dma, cp->rx_len,
DMA_FROM_DEVICE);
pass_to_upper_layers(cp->rx_buf);
make_and_setup_new_rx_buf(cp);
} else {
/* CPU should not write to
* DMA_FROM_DEVICE-mapped area,
* so dma_sync_single_for_device() is
* not needed here. It would be required
* for DMA_BIDIRECTIONAL mapping if
* the memory was modified.
*/
give_rx_buf_to_card(cp);
}
}
}
예제는 receive buffer를 `DMA_FROM_DEVICE`로 mapping하고 card에 넘깁니다. Interrupt handler는 header를 검사하기 전에 `dma_sync_single_for_cpu()`로 CPU 관찰 상태를 갱신합니다. Header가 유효하면 unmap하고 상위 layer로 넘깁니다. 유효하지 않으면 buffer를 card에 다시 주며, CPU가 `DMA_FROM_DEVICE` 영역을 쓰지 않았으므로 `dma_sync_single_for_device()`는 필요 없습니다. `DMA_BIDIRECTIONAL` mapping에서 memory를 수정했다면 필요합니다.
DMA mapping 오류 처리
730-834오류 처리
일부 architecture에서는 DMA address space가 제한되어 있습니다. `dma_alloc_coherent()`가 `NULL`을 반환하거나 `dma_map_sg()`가 0을 반환하는지 검사하여 allocation failure를 확인합니다.
`dma_map_single()`과 `dma_map_page()`가 반환한 `dma_addr_t`는 `dma_mapping_error()`로 검사합니다.
dma_addr_t dma_handle;
dma_handle = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
여러 page를 mapping하는 도중 오류가 나면 이미 mapping한 page를 unmap해야 합니다. 다음 예들은 `dma_map_page()`에도 적용됩니다.
예제 1은 두 번째 mapping 실패 시 첫 번째 mapping을 해제합니다.
dma_addr_t dma_handle1;
dma_addr_t dma_handle2;
dma_handle1 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle1)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling1;
}
dma_handle2 = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_handle2)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling2;
}
...
map_error_handling2:
dma_unmap_single(dma_handle1);
map_error_handling1:
예제 2는 loop 도중 오류가 나면 성공한 index 수를 이용해 앞서 mapping한 모든 buffer를 해제합니다.
/*
* if buffers are allocated in a loop, unmap all mapped buffers when
* mapping error is detected in the middle
*/
dma_addr_t dma_addr;
dma_addr_t array[DMA_BUFFERS];
int save_index = 0;
for (i = 0; i < DMA_BUFFERS; i++) {
...
dma_addr = dma_map_single(dev, addr, size, direction);
if (dma_mapping_error(dev, dma_addr)) {
/*
* reduce current DMA mapping usage,
* delay and try again later or
* reset driver.
*/
goto map_error_handling;
}
array[i].dma_addr = dma_addr;
save_index++;
}
...
map_error_handling:
for (i = 0; i < save_index; i++) {
...
dma_unmap_single(array[i].dma_addr);
}
Networking driver는 transmit hook인 `ndo_start_xmit`에서 DMA mapping이 실패하면 `dev_kfree_skb()`로 socket buffer를 free하고 `NETDEV_TX_OK`를 반환해야 합니다. 실패한 socket buffer는 drop됩니다.
SCSI driver는 `queuecommand` hook에서 DMA mapping이 실패하면 `SCSI_MLQUEUE_HOST_BUSY`를 반환해야 합니다. 그러면 SCSI subsystem이 나중에 command를 다시 driver에 전달합니다.
Unmap 상태 공간 사용량 최적화
835-891Unmap 상태 공간 사용량 최적화
많은 platform에서 `dma_unmap_{single,page}()`는 단순한 nop입니다. 이 경우 mapping address와 length를 계속 저장하는 것은 공간 낭비입니다. Portable API의 목적을 해치는 `#ifdef` 우회 대신 다음 기능을 사용합니다.
Macro를 하나씩 설명하는 대신 예제 code를 변환해 보겠습니다. 첫째, state 저장 structure에서 `DEFINE_DMA_UNMAP_{ADDR,LEN}`을 사용합니다. 변경 전:
struct ring_state {
struct sk_buff *skb;
dma_addr_t mapping;
__u32 len;
};
변경 후:
struct ring_state {
struct sk_buff *skb;
DEFINE_DMA_UNMAP_ADDR(mapping);
DEFINE_DMA_UNMAP_LEN(len);
};
둘째, 값을 설정할 때 `dma_unmap_{addr,len}_set()`을 사용합니다. 변경 전:
ringp->mapping = FOO;
ringp->len = BAR;
변경 후:
dma_unmap_addr_set(ringp, mapping, FOO);
dma_unmap_len_set(ringp, len, BAR);
셋째, 값에 접근할 때 `dma_unmap_{addr,len}()`을 사용합니다. 변경 전:
dma_unmap_single(dev, ringp->mapping, ringp->len,
DMA_FROM_DEVICE);
변경 후:
dma_unmap_single(dev,
dma_unmap_addr(ringp, mapping),
dma_unmap_len(ringp, len),
DMA_FROM_DEVICE);
동작은 이름 그대로입니다. 구현에 따라 unmap operation에 address만 필요할 수 있으므로 `ADDR`과 `LEN`은 별도로 취급합니다.
Architecture port의 platform 고려 사항
892-918Platform 문제
Linux driver만 작성하고 kernel architecture port를 유지하지 않는다면 이 절을 건너뛰고 `Closing`으로 가도 됩니다.
1) `struct scatterlist` 요구 사항: Architecture가 software IOMMU를 포함한 IOMMU를 지원한다면 `CONFIG_NEED_SG_DMA_LENGTH`를 활성화해야 합니다.
2) `ARCH_DMA_MINALIGN`: Architecture는 `kmalloc()` buffer가 DMA-safe하도록 보장해야 하며 driver와 subsystem은 이를 전제로 합니다. Hardware가 CPU cache의 data와 main memory의 data가 같음을 보장하지 않는 등 완전히 DMA-coherent하지 않다면 memory allocator가 `kmalloc()` buffer를 다른 data와 같은 cache line에 두지 않도록 `ARCH_DMA_MINALIGN`을 설정해야 합니다. 예제는 `arch/arm/include/asm/cache.h`를 참고하십시오.
`ARCH_DMA_MINALIGN`은 DMA memory alignment constraint를 위한 것입니다. 64-bit object 정렬 같은 architecture data alignment constraint는 여기서 고려할 필요가 없습니다.
맺음말과 기여자
919-935맺음말
수많은 사람의 feedback과 제안이 없었다면 이 문서와 API는 현재 형태가 되지 못했을 것입니다. 특별히 다음 기여자들을 순서 없이 언급합니다.
Russell King <rmk@arm.linux.org.uk>
Leo Dagum <dagum@barrel.engr.sgi.com>
Ralf Baechle <ralf@oss.sgi.com>
Grant Grundler <grundler@cup.hp.com>
Jay Estabrook <Jay.Estabrook@compaq.com>
Thomas Sailer <sailer@ife.ee.ethz.ch>
Andrea Arcangeli <andrea@suse.de>
Jens Axboe <jens.axboe@oracle.com>
David Mosberger-Tang <davidm@hpl.hp.com>
요약과 해설
dma-api-howto.rst:1-935DMA driver는 CPU virtual address, CPU physical address, device bus address를 구분해야 합니다. `dma_map_*()` API는 필요한 IOMMU mapping을 만들고 device가 실제로 사용할 `dma_addr_t`를 반환하므로 bus 전용 API 대신 공통 DMA API를 사용해야 합니다.
Coherent mapping은 CPU와 device가 서로의 갱신을 즉시 볼 수 있는 장기 공유 영역에, streaming mapping은 개별 transfer에 적합합니다. Coherent memory도 CPU store ordering을 위한 memory barrier가 필요하며 streaming buffer를 transfer 사이에 CPU가 만지면 `dma_sync_*_for_cpu()`와 `dma_sync_*_for_device()`로 ownership을 전환해야 합니다.
모든 mapping 결과는 오류를 검사하고 반드시 대응하는 unmap을 수행해야 합니다. DMA mask는 device 능력에 맞게 probe하고, scatterlist의 `nents`에는 mapping 입력 개수와 반환된 segment 개수의 차이를 엄격히 지켜야 합니다. 이 규칙을 어기면 address space 고갈, panic, 조용한 data corruption으로 이어질 수 있습니다.