← Documents Documentation/PCI/pci.rst GitHub 원문 ↗

Linux 6.18.37 · PCI

Linux PCI driver 작성법

PCI driver 등록부터 device·resource·DMA·IRQ 초기화, shutdown, config access와 MMIO write posting까지 전체 생명주기를 설명합니다.

Source pathDocumentation/PCI/pci.rst
Source versionLinux v6.18.37
TranslationDUJINLABS 전문 번역 + 해설

요약·해설과 원문, 전문 번역을 서로 분리했습니다. API 이름, symbol, source path는 원문 표기를 사용합니다.

1. 요약·해설

원문의 핵심 논리와 kernel programming 관점의 보충 설명입니다. 아래의 전문 번역과는 별도로 작성했습니다.

요약·해설

pci.rst:1-578

PCI driver는 `pci_register_driver()`로 match·hotplug를 PCI core에 맡기고, device enable→resource→DMA mask→control data→IRQ→engine 순으로 초기화합니다.

제거 시 IRQ source와 DMA를 먼저 완전히 멈춘 뒤 buffer·subsystem·mapping·region을 역순으로 해제해야 memory corruption과 screaming interrupt를 피할 수 있습니다.

MMIO write는 posted될 수 있으므로 timing-sensitive path와 reset에서는 side effect 없는 read 또는 config-space read로 flush해야 합니다.

2. 영어 원문 전체

번역 기준이 된 Linux v6.18.37 원문입니다. 줄 번호는 이 버전의 파일 좌표입니다.

원문 전체 펼치기
1 .. SPDX-License-Identifier: GPL-2.0
2
3 ==============================
4 How To Write Linux PCI Drivers
5 ==============================
6
7 :Authors: - Martin Mares <mj@ucw.cz>
8 - Grant Grundler <grundler@parisc-linux.org>
9
10 The world of PCI is vast and full of (mostly unpleasant) surprises.
11 Since each CPU architecture implements different chip-sets and PCI devices
12 have different requirements (erm, "features"), the result is the PCI support
13 in the Linux kernel is not as trivial as one would wish. This short paper
14 tries to introduce all potential driver authors to Linux APIs for
15 PCI device drivers.
16
17 A more complete resource is the third edition of "Linux Device Drivers"
18 by Jonathan Corbet, Alessandro Rubini, and Greg Kroah-Hartman.
19 LDD3 is available for free (under Creative Commons License) from:
20 https://lwn.net/Kernel/LDD3/.
21
22 However, keep in mind that all documents are subject to "bit rot".
23 Refer to the source code if things are not working as described here.
24
25 Please send questions/comments/patches about Linux PCI API to the
26 "Linux PCI" <linux-pci@atrey.karlin.mff.cuni.cz> mailing list.
27
28
29 Structure of PCI drivers
30 ========================
31 PCI drivers "discover" PCI devices in a system via pci_register_driver().
32 Actually, it's the other way around. When the PCI generic code discovers
33 a new device, the driver with a matching "description" will be notified.
34 Details on this below.
35
36 pci_register_driver() leaves most of the probing for devices to
37 the PCI layer and supports online insertion/removal of devices [thus
38 supporting hot-pluggable PCI, CardBus, and Express-Card in a single driver].
39 pci_register_driver() call requires passing in a table of function
40 pointers and thus dictates the high level structure of a driver.
41
42 Once the driver knows about a PCI device and takes ownership, the
43 driver generally needs to perform the following initialization:
44
45 - Enable the device
46 - Request MMIO/IOP resources
47 - Set the DMA mask size (for both coherent and streaming DMA)
48 - Allocate and initialize shared control data (pci_allocate_coherent())
49 - Access device configuration space (if needed)
50 - Register IRQ handler (request_irq())
51 - Initialize non-PCI (i.e. LAN/SCSI/etc parts of the chip)
52 - Enable DMA/processing engines
53
54 When done using the device, and perhaps the module needs to be unloaded,
55 the driver needs to take the following steps:
56
57 - Disable the device from generating IRQs
58 - Release the IRQ (free_irq())
59 - Stop all DMA activity
60 - Release DMA buffers (both streaming and coherent)
61 - Unregister from other subsystems (e.g. scsi or netdev)
62 - Release MMIO/IOP resources
63 - Disable the device
64
65 Most of these topics are covered in the following sections.
66 For the rest look at LDD3 or <linux/pci.h> .
67
68 If the PCI subsystem is not configured (CONFIG_PCI is not set), most of
69 the PCI functions described below are defined as inline functions either
70 completely empty or just returning an appropriate error codes to avoid
71 lots of ifdefs in the drivers.
72
73
74 pci_register_driver() call
75 ==========================
76
77 PCI device drivers call ``pci_register_driver()`` during their
78 initialization with a pointer to a structure describing the driver
79 (``struct pci_driver``):
80
81 .. kernel-doc:: include/linux/pci.h
82 :functions: pci_driver
83
84 The ID table is an array of ``struct pci_device_id`` entries ending with an
85 all-zero entry. Definitions with static const are generally preferred.
86
87 .. kernel-doc:: include/linux/mod_devicetable.h
88 :functions: pci_device_id
89
90 Most drivers only need ``PCI_DEVICE()`` or ``PCI_DEVICE_CLASS()`` to set up
91 a pci_device_id table.
92
93 New PCI IDs may be added to a device driver pci_ids table at runtime
94 as shown below::
95
96 echo "vendor device subvendor subdevice class class_mask driver_data" > \
97 /sys/bus/pci/drivers/{driver}/new_id
98
99 All fields are passed in as hexadecimal values (no leading 0x).
100 The vendor and device fields are mandatory, the others are optional. Users
101 need pass only as many optional fields as necessary:
102
103 - subvendor and subdevice fields default to PCI_ANY_ID (FFFFFFFF)
104 - class and classmask fields default to 0
105 - driver_data defaults to 0UL.
106 - override_only field defaults to 0.
107
108 Note that driver_data must match the value used by any of the pci_device_id
109 entries defined in the driver. This makes the driver_data field mandatory
110 if all the pci_device_id entries have a non-zero driver_data value.
111
112 Once added, the driver probe routine will be invoked for any unclaimed
113 PCI devices listed in its (newly updated) pci_ids list.
114
115 When the driver exits, it just calls pci_unregister_driver() and the PCI layer
116 automatically calls the remove hook for all devices handled by the driver.
117
118
119 "Attributes" for driver functions/data
120 --------------------------------------
121
122 Please mark the initialization and cleanup functions where appropriate
123 (the corresponding macros are defined in <linux/init.h>):
124
125 ====== =================================================
126 __init Initialization code. Thrown away after the driver
127 initializes.
128 __exit Exit code. Ignored for non-modular drivers.
129 ====== =================================================
130
131 Tips on when/where to use the above attributes:
132 - The module_init()/module_exit() functions (and all
133 initialization functions called _only_ from these)
134 should be marked __init/__exit.
135
136 - Do not mark the struct pci_driver.
137
138 - Do NOT mark a function if you are not sure which mark to use.
139 Better to not mark the function than mark the function wrong.
140
141
142 How to find PCI devices manually
143 ================================
144
145 PCI drivers should have a really good reason for not using the
146 pci_register_driver() interface to search for PCI devices.
147 The main reason PCI devices are controlled by multiple drivers
148 is because one PCI device implements several different HW services.
149 E.g. combined serial/parallel port/floppy controller.
150
151 A manual search may be performed using the following constructs:
152
153 Searching by vendor and device ID::
154
155 struct pci_dev *dev = NULL;
156 while (dev = pci_get_device(VENDOR_ID, DEVICE_ID, dev))
157 configure_device(dev);
158
159 Searching by class ID (iterate in a similar way)::
160
161 pci_get_class(CLASS_ID, dev)
162
163 Searching by both vendor/device and subsystem vendor/device ID::
164
165 pci_get_subsys(VENDOR_ID,DEVICE_ID, SUBSYS_VENDOR_ID, SUBSYS_DEVICE_ID, dev).
166
167 You can use the constant PCI_ANY_ID as a wildcard replacement for
168 VENDOR_ID or DEVICE_ID. This allows searching for any device from a
169 specific vendor, for example.
170
171 These functions are hotplug-safe. They increment the reference count on
172 the pci_dev that they return. You must eventually (possibly at module unload)
173 decrement the reference count on these devices by calling pci_dev_put().
174
175
176 Device Initialization Steps
177 ===========================
178
179 As noted in the introduction, most PCI drivers need the following steps
180 for device initialization:
181
182 - Enable the device
183 - Request MMIO/IOP resources
184 - Set the DMA mask size (for both coherent and streaming DMA)
185 - Allocate and initialize shared control data (pci_allocate_coherent())
186 - Access device configuration space (if needed)
187 - Register IRQ handler (request_irq())
188 - Initialize non-PCI (i.e. LAN/SCSI/etc parts of the chip)
189 - Enable DMA/processing engines.
190
191 The driver can access PCI config space registers at any time.
192 (Well, almost. When running BIST, config space can go away...but
193 that will just result in a PCI Bus Master Abort and config reads
194 will return garbage).
195
196
197 Enable the PCI device
198 ---------------------
199 Before touching any device registers, the driver needs to enable
200 the PCI device by calling pci_enable_device(). This will:
201
202 - wake up the device if it was in suspended state,
203 - allocate I/O and memory regions of the device (if BIOS did not),
204 - allocate an IRQ (if BIOS did not).
205
206 .. note::
207 pci_enable_device() can fail! Check the return value.
208
209 .. warning::
210 OS BUG: we don't check resource allocations before enabling those
211 resources. The sequence would make more sense if we called
212 pci_request_resources() before calling pci_enable_device().
213 Currently, the device drivers can't detect the bug when two
214 devices have been allocated the same range. This is not a common
215 problem and unlikely to get fixed soon.
216
217 This has been discussed before but not changed as of 2.6.19:
218 https://lore.kernel.org/r/20060302180025.GC28895@flint.arm.linux.org.uk/
219
220
221 pci_set_master() will enable DMA by setting the bus master bit
222 in the PCI_COMMAND register. It also fixes the latency timer value if
223 it's set to something bogus by the BIOS. pci_clear_master() will
224 disable DMA by clearing the bus master bit.
225
226 If the PCI device can use the PCI Memory-Write-Invalidate transaction,
227 call pci_set_mwi(). This enables the PCI_COMMAND bit for Mem-Wr-Inval
228 and also ensures that the cache line size register is set correctly.
229 Check the return value of pci_set_mwi() as not all architectures
230 or chip-sets may support Memory-Write-Invalidate. Alternatively,
231 if Mem-Wr-Inval would be nice to have but is not required, call
232 pci_try_set_mwi() to have the system do its best effort at enabling
233 Mem-Wr-Inval.
234
235
236 Request MMIO/IOP resources
237 --------------------------
238 Memory (MMIO), and I/O port addresses should NOT be read directly
239 from the PCI device config space. Use the values in the pci_dev structure
240 as the PCI "bus address" might have been remapped to a "host physical"
241 address by the arch/chip-set specific kernel support.
242
243 See Documentation/driver-api/io-mapping.rst for how to access device registers
244 or device memory.
245
246 The device driver needs to call pci_request_region() to verify
247 no other device is already using the same address resource.
248 Conversely, drivers should call pci_release_region() AFTER
249 calling pci_disable_device().
250 The idea is to prevent two devices colliding on the same address range.
251
252 .. tip::
253 See OS BUG comment above. Currently (2.6.19), The driver can only
254 determine MMIO and IO Port resource availability _after_ calling
255 pci_enable_device().
256
257 Generic flavors of pci_request_region() are request_mem_region()
258 (for MMIO ranges) and request_region() (for IO Port ranges).
259 Use these for address resources that are not described by "normal" PCI
260 BARs.
261
262 Also see pci_request_selected_regions() below.
263
264
265 Set the DMA mask size
266 ---------------------
267 .. note::
268 If anything below doesn't make sense, please refer to
269 Documentation/core-api/dma-api.rst. This section is just a reminder that
270 drivers need to indicate DMA capabilities of the device and is not
271 an authoritative source for DMA interfaces.
272
273 While all drivers should explicitly indicate the DMA capability
274 (e.g. 32 or 64 bit) of the PCI bus master, devices with more than
275 32-bit bus master capability for streaming data need the driver
276 to "register" this capability by calling dma_set_mask() with
277 appropriate parameters. In general this allows more efficient DMA
278 on systems where System RAM exists above 4G _physical_ address.
279
280 Drivers for all PCI-X and PCIe compliant devices must call
281 dma_set_mask() as they are 64-bit DMA devices.
282
283 Similarly, drivers must also "register" this capability if the device
284 can directly address "coherent memory" in System RAM above 4G physical
285 address by calling dma_set_coherent_mask().
286 Again, this includes drivers for all PCI-X and PCIe compliant devices.
287 Many 64-bit "PCI" devices (before PCI-X) and some PCI-X devices are
288 64-bit DMA capable for payload ("streaming") data but not control
289 ("coherent") data.
290
291
292 Setup shared control data
293 -------------------------
294 Once the DMA masks are set, the driver can allocate "coherent" (a.k.a. shared)
295 memory. See Documentation/core-api/dma-api.rst for a full description of
296 the DMA APIs. This section is just a reminder that it needs to be done
297 before enabling DMA on the device.
298
299
300 Initialize device registers
301 ---------------------------
302 Some drivers will need specific "capability" fields programmed
303 or other "vendor specific" register initialized or reset.
304 E.g. clearing pending interrupts.
305
306
307 Register IRQ handler
308 --------------------
309 While calling request_irq() is the last step described here,
310 this is often just another intermediate step to initialize a device.
311 This step can often be deferred until the device is opened for use.
312
313 All interrupt handlers for IRQ lines should be registered with IRQF_SHARED
314 and use the devid to map IRQs to devices (remember that all PCI IRQ lines
315 can be shared).
316
317 request_irq() will associate an interrupt handler and device handle
318 with an interrupt number. Historically interrupt numbers represent
319 IRQ lines which run from the PCI device to the Interrupt controller.
320 With MSI and MSI-X (more below) the interrupt number is a CPU "vector".
321
322 request_irq() also enables the interrupt. Make sure the device is
323 quiesced and does not have any interrupts pending before registering
324 the interrupt handler.
325
326 MSI and MSI-X are PCI capabilities. Both are "Message Signaled Interrupts"
327 which deliver interrupts to the CPU via a DMA write to a Local APIC.
328 The fundamental difference between MSI and MSI-X is how multiple
329 "vectors" get allocated. MSI requires contiguous blocks of vectors
330 while MSI-X can allocate several individual ones.
331
332 MSI capability can be enabled by calling pci_alloc_irq_vectors() with the
333 PCI_IRQ_MSI and/or PCI_IRQ_MSIX flags before calling request_irq(). This
334 causes the PCI support to program CPU vector data into the PCI device
335 capability registers. Many architectures, chip-sets, or BIOSes do NOT
336 support MSI or MSI-X and a call to pci_alloc_irq_vectors with just
337 the PCI_IRQ_MSI and PCI_IRQ_MSIX flags will fail, so try to always
338 specify PCI_IRQ_INTX as well.
339
340 Drivers that have different interrupt handlers for MSI/MSI-X and
341 legacy INTx should chose the right one based on the msi_enabled
342 and msix_enabled flags in the pci_dev structure after calling
343 pci_alloc_irq_vectors.
344
345 There are (at least) two really good reasons for using MSI:
346
347 1) MSI is an exclusive interrupt vector by definition.
348 This means the interrupt handler doesn't have to verify
349 its device caused the interrupt.
350
351 2) MSI avoids DMA/IRQ race conditions. DMA to host memory is guaranteed
352 to be visible to the host CPU(s) when the MSI is delivered. This
353 is important for both data coherency and avoiding stale control data.
354 This guarantee allows the driver to omit MMIO reads to flush
355 the DMA stream.
356
357 See drivers/infiniband/hw/mthca/ or drivers/net/tg3.c for examples
358 of MSI/MSI-X usage.
359
360
361 PCI device shutdown
362 ===================
363
364 When a PCI device driver is being unloaded, most of the following
365 steps need to be performed:
366
367 - Disable the device from generating IRQs
368 - Release the IRQ (free_irq())
369 - Stop all DMA activity
370 - Release DMA buffers (both streaming and coherent)
371 - Unregister from other subsystems (e.g. scsi or netdev)
372 - Disable device from responding to MMIO/IO Port addresses
373 - Release MMIO/IO Port resource(s)
374
375
376 Stop IRQs on the device
377 -----------------------
378 How to do this is chip/device specific. If it's not done, it opens
379 the possibility of a "screaming interrupt" if (and only if)
380 the IRQ is shared with another device.
381
382 When the shared IRQ handler is "unhooked", the remaining devices
383 using the same IRQ line will still need the IRQ enabled. Thus if the
384 "unhooked" device asserts IRQ line, the system will respond assuming
385 it was one of the remaining devices asserted the IRQ line. Since none
386 of the other devices will handle the IRQ, the system will "hang" until
387 it decides the IRQ isn't going to get handled and masks the IRQ (100,000
388 iterations later). Once the shared IRQ is masked, the remaining devices
389 will stop functioning properly. Not a nice situation.
390
391 This is another reason to use MSI or MSI-X if it's available.
392 MSI and MSI-X are defined to be exclusive interrupts and thus
393 are not susceptible to the "screaming interrupt" problem.
394
395
396 Release the IRQ
397 ---------------
398 Once the device is quiesced (no more IRQs), one can call free_irq().
399 This function will return control once any pending IRQs are handled,
400 "unhook" the drivers IRQ handler from that IRQ, and finally release
401 the IRQ if no one else is using it.
402
403
404 Stop all DMA activity
405 ---------------------
406 It's extremely important to stop all DMA operations BEFORE attempting
407 to deallocate DMA control data. Failure to do so can result in memory
408 corruption, hangs, and on some chip-sets a hard crash.
409
410 Stopping DMA after stopping the IRQs can avoid races where the
411 IRQ handler might restart DMA engines.
412
413 While this step sounds obvious and trivial, several "mature" drivers
414 didn't get this step right in the past.
415
416
417 Release DMA buffers
418 -------------------
419 Once DMA is stopped, clean up streaming DMA first.
420 I.e. unmap data buffers and return buffers to "upstream"
421 owners if there is one.
422
423 Then clean up "coherent" buffers which contain the control data.
424
425 See Documentation/core-api/dma-api.rst for details on unmapping interfaces.
426
427
428 Unregister from other subsystems
429 --------------------------------
430 Most low level PCI device drivers support some other subsystem
431 like USB, ALSA, SCSI, NetDev, Infiniband, etc. Make sure your
432 driver isn't losing resources from that other subsystem.
433 If this happens, typically the symptom is an Oops (panic) when
434 the subsystem attempts to call into a driver that has been unloaded.
435
436
437 Disable Device from responding to MMIO/IO Port addresses
438 --------------------------------------------------------
439 io_unmap() MMIO or IO Port resources and then call pci_disable_device().
440 This is the symmetric opposite of pci_enable_device().
441 Do not access device registers after calling pci_disable_device().
442
443
444 Release MMIO/IO Port Resource(s)
445 --------------------------------
446 Call pci_release_region() to mark the MMIO or IO Port range as available.
447 Failure to do so usually results in the inability to reload the driver.
448
449
450 How to access PCI config space
451 ==============================
452
453 You can use `pci_(read|write)_config_(byte|word|dword)` to access the config
454 space of a device represented by `struct pci_dev *`. All these functions return
455 0 when successful or an error code (`PCIBIOS_...`) which can be translated to a
456 text string by pcibios_strerror. Most drivers expect that accesses to valid PCI
457 devices don't fail.
458
459 If you don't have a struct pci_dev available, you can call
460 `pci_bus_(read|write)_config_(byte|word|dword)` to access a given device
461 and function on that bus.
462
463 If you access fields in the standard portion of the config header, please
464 use symbolic names of locations and bits declared in <linux/pci.h>.
465
466 If you need to access Extended PCI Capability registers, just call
467 pci_find_capability() for the particular capability and it will find the
468 corresponding register block for you.
469
470
471 Other interesting functions
472 ===========================
473
474 ============================= ================================================
475 pci_get_domain_bus_and_slot() Find pci_dev corresponding to given domain,
476 bus and slot and number. If the device is
477 found, its reference count is increased.
478 pci_set_power_state() Set PCI Power Management state (0=D0 ... 3=D3)
479 pci_find_capability() Find specified capability in device's capability
480 list.
481 pci_resource_start() Returns bus start address for a given PCI region
482 pci_resource_end() Returns bus end address for a given PCI region
483 pci_resource_len() Returns the byte length of a PCI region
484 pci_set_drvdata() Set private driver data pointer for a pci_dev
485 pci_get_drvdata() Return private driver data pointer for a pci_dev
486 pci_set_mwi() Enable Memory-Write-Invalidate transactions.
487 pci_clear_mwi() Disable Memory-Write-Invalidate transactions.
488 ============================= ================================================
489
490
491 Miscellaneous hints
492 ===================
493
494 When displaying PCI device names to the user (for example when a driver wants
495 to tell the user what card has it found), please use pci_name(pci_dev).
496
497 Always refer to the PCI devices by a pointer to the pci_dev structure.
498 All PCI layer functions use this identification and it's the only
499 reasonable one. Don't use bus/slot/function numbers except for very
500 special purposes -- on systems with multiple primary buses their semantics
501 can be pretty complex.
502
503 Don't try to turn on Fast Back to Back writes in your driver. All devices
504 on the bus need to be capable of doing it, so this is something which needs
505 to be handled by platform and generic code, not individual drivers.
506
507
508 Vendor and device identifications
509 =================================
510
511 Do not add new device or vendor IDs to include/linux/pci_ids.h unless they
512 are shared across multiple drivers. You can add private definitions in
513 your driver if they're helpful, or just use plain hex constants.
514
515 The device IDs are arbitrary hex numbers (vendor controlled) and normally used
516 only in a single location, the pci_device_id table.
517
518 Please DO submit new vendor/device IDs to https://pci-ids.ucw.cz/.
519 There's a mirror of the pci.ids file at https://github.com/pciutils/pciids.
520
521
522 Obsolete functions
523 ==================
524
525 There are several functions which you might come across when trying to
526 port an old driver to the new PCI interface. They are no longer present
527 in the kernel as they aren't compatible with hotplug or PCI domains or
528 having sane locking.
529
530 ================= ===========================================
531 pci_find_device() Superseded by pci_get_device()
532 pci_find_subsys() Superseded by pci_get_subsys()
533 pci_find_slot() Superseded by pci_get_domain_bus_and_slot()
534 pci_get_slot() Superseded by pci_get_domain_bus_and_slot()
535 ================= ===========================================
536
537 The alternative is the traditional PCI device driver that walks PCI
538 device lists. This is still possible but discouraged.
539
540
541 MMIO Space and "Write Posting"
542 ==============================
543
544 Converting a driver from using I/O Port space to using MMIO space
545 often requires some additional changes. Specifically, "write posting"
546 needs to be handled. Many drivers (e.g. tg3, acenic, sym53c8xx_2)
547 already do this. I/O Port space guarantees write transactions reach the PCI
548 device before the CPU can continue. Writes to MMIO space allow the CPU
549 to continue before the transaction reaches the PCI device. HW weenies
550 call this "Write Posting" because the write completion is "posted" to
551 the CPU before the transaction has reached its destination.
552
553 Thus, timing sensitive code should add readl() where the CPU is
554 expected to wait before doing other work. The classic "bit banging"
555 sequence works fine for I/O Port space::
556
557 for (i = 8; --i; val >>= 1) {
558 outb(val & 1, ioport_reg); /* write bit */
559 udelay(10);
560 }
561
562 The same sequence for MMIO space should be::
563
564 for (i = 8; --i; val >>= 1) {
565 writeb(val & 1, mmio_reg); /* write bit */
566 readb(safe_mmio_reg); /* flush posted write */
567 udelay(10);
568 }
569
570 It is important that "safe_mmio_reg" not have any side effects that
571 interferes with the correct operation of the device.
572
573 Another case to watch out for is when resetting a PCI device. Use PCI
574 Configuration space reads to flush the writel(). This will gracefully
575 handle the PCI master abort on all platforms if the PCI device is
576 expected to not respond to a readl(). Most x86 platforms will allow
577 MMIO reads to master abort (a.k.a. "Soft Fail") and return garbage
578 (e.g. ~0). But many RISC platforms will crash (a.k.a."Hard Fail").
579

3. 한국어 전문 번역

영어 원문의 문단 순서와 의미를 유지한 전체 번역입니다. 코드, 함수명, symbol과 URL은 원문 표기를 유지합니다.

소개와 참고 자료

1-28

저자는 Martin Mares와 Grant Grundler입니다. PCI 세계는 방대하고 대부분 불쾌한 surprise로 가득합니다. CPU architecture마다 chipset 구현이 다르고 PCI device마다 서로 다른 요구, 즉 feature가 있어 Linux kernel의 PCI 지원은 기대만큼 단순하지 않습니다.

이 문서는 PCI device driver 작성자에게 Linux PCI API를 소개합니다. 더 완전한 자료는 Jonathan Corbet, Alessandro Rubini, Greg Kroah-Hartman의 『Linux Device Drivers』 3판이며 Creative Commons license로 `https://lwn.net/Kernel/LDD3/`에서 무료로 볼 수 있습니다.

모든 문서는 시간이 지나면 낡을 수 있으므로 설명대로 동작하지 않으면 source code를 확인하십시오. Linux PCI API 질문·의견·patch는 `Linux PCI <linux-pci@atrey.karlin.mff.cuni.cz>` mailing list로 보내십시오.

.. SPDX-License-Identifier: GPL-2.0

==============================
How To Write Linux PCI Drivers
==============================

:Authors: - Martin Mares <mj@ucw.cz>
          - Grant Grundler <grundler@parisc-linux.org>

The world of PCI is vast and full of (mostly unpleasant) surprises.
Since each CPU architecture implements different chip-sets and PCI devices
have different requirements (erm, "features"), the result is the PCI support
in the Linux kernel is not as trivial as one would wish. This short paper
tries to introduce all potential driver authors to Linux APIs for
PCI device drivers.

A more complete resource is the third edition of "Linux Device Drivers"
by Jonathan Corbet, Alessandro Rubini, and Greg Kroah-Hartman.
LDD3 is available for free (under Creative Commons License) from:
https://lwn.net/Kernel/LDD3/.

However, keep in mind that all documents are subject to "bit rot".
Refer to the source code if things are not working as described here.

Please send questions/comments/patches about Linux PCI API to the
"Linux PCI" <linux-pci@atrey.karlin.mff.cuni.cz> mailing list.

PCI driver 구조와 생명주기

29-73

PCI driver는 `pci_register_driver()`로 system의 PCI device를 발견한다고 표현하지만 실제로는 generic PCI code가 새 device를 발견하고 일치하는 description을 가진 driver에 알립니다.

`pci_register_driver()`는 device probing 대부분을 PCI layer에 맡기며 hot-pluggable PCI, CardBus, ExpressCard의 online insertion/removal을 한 driver에서 지원합니다. Function pointer table을 전달하므로 driver의 상위 구조도 결정합니다.

Device를 소유한 driver는 일반적으로 device enable, MMIO/IOP resource 요청, streaming·coherent DMA mask 설정, shared control data 할당·초기화, 필요 시 config space 접근, IRQ handler 등록, chip의 LAN/SCSI 등 non-PCI 부분 초기화, DMA/processing engine enable 순으로 진행합니다.

Device 사용을 끝낼 때는 IRQ 생성 disable, IRQ 해제, DMA 중지, streaming·coherent buffer 해제, 다른 subsystem 등록 해제, MMIO/IOP resource 해제, device disable 순으로 정리합니다.

PCI driver 초기화
pci_register_driver()Device enableMMIO/IOP requestDMA masksShared control dataConfig + IRQNon-PCI initDMA engines

PCI layer가 device를 match한 뒤 driver가 가져야 할 일반 순서입니다.

PCI driver 제거
IRQ source disablefree_irq()DMA stopDMA buffers freeSubsystem unregisterMMIO/IOP releaseDevice disable

초기화의 역순으로 hardware activity와 resource를 정리합니다.

나머지 세부 사항은 LDD3 또는 `<linux/pci.h>`를 참조하십시오. `CONFIG_PCI`가 설정되지 않으면 아래 PCI function 대부분은 driver의 많은 `#ifdef`를 피하도록 빈 inline function 또는 적절한 error를 반환하는 inline function으로 정의됩니다.

Structure of PCI drivers
========================
PCI drivers "discover" PCI devices in a system via pci_register_driver().
Actually, it's the other way around. When the PCI generic code discovers
a new device, the driver with a matching "description" will be notified.
Details on this below.

pci_register_driver() leaves most of the probing for devices to
the PCI layer and supports online insertion/removal of devices [thus
supporting hot-pluggable PCI, CardBus, and Express-Card in a single driver].
pci_register_driver() call requires passing in a table of function
pointers and thus dictates the high level structure of a driver.

Once the driver knows about a PCI device and takes ownership, the
driver generally needs to perform the following initialization:

  - Enable the device
  - Request MMIO/IOP resources
  - Set the DMA mask size (for both coherent and streaming DMA)
  - Allocate and initialize shared control data (pci_allocate_coherent())
  - Access device configuration space (if needed)
  - Register IRQ handler (request_irq())
  - Initialize non-PCI (i.e. LAN/SCSI/etc parts of the chip)
  - Enable DMA/processing engines

When done using the device, and perhaps the module needs to be unloaded,
the driver needs to take the following steps:

  - Disable the device from generating IRQs
  - Release the IRQ (free_irq())
  - Stop all DMA activity
  - Release DMA buffers (both streaming and coherent)
  - Unregister from other subsystems (e.g. scsi or netdev)
  - Release MMIO/IOP resources
  - Disable the device

Most of these topics are covered in the following sections.
For the rest look at LDD3 or <linux/pci.h> .

If the PCI subsystem is not configured (CONFIG_PCI is not set), most of
the PCI functions described below are defined as inline functions either
completely empty or just returning an appropriate error codes to avoid
lots of ifdefs in the drivers.

pci_register_driver()와 ID table

74-118

PCI device driver는 초기화 중 `struct pci_driver`를 가리키는 pointer로 `pci_register_driver()`를 호출합니다. 구조체 문서는 `include/linux/pci.h`의 kernel-doc에서 생성됩니다.

ID table은 `struct pci_device_id` entry 배열이며 all-zero entry로 끝납니다. 일반적으로 `static const` 정의가 권장되고 구조체 문서는 `include/linux/mod_devicetable.h`에서 생성됩니다. 대부분 driver는 `PCI_DEVICE()` 또는 `PCI_DEVICE_CLASS()`만 있으면 table을 만들 수 있습니다.

Runtime에는 sysfs `new_id`에 hexadecimal field를 기록해 driver의 `pci_ids` table에 새 ID를 추가할 수 있습니다. Leading `0x`는 쓰지 않습니다.

echo "vendor device subvendor subdevice class class_mask driver_data" > \
/sys/bus/pci/drivers/{driver}/new_id
new_id field
Field필수기본값
vendor없음
device없음
subvendor아니오PCI_ANY_ID (FFFFFFFF)
subdevice아니오PCI_ANY_ID (FFFFFFFF)
class아니오0
class_mask아니오0
driver_data아니오0UL
override_only아니오0

Vendor와 device는 필수이고 나머지는 필요한 만큼만 전달합니다.

`driver_data`는 driver에 정의된 `pci_device_id` entry의 값과 일치해야 합니다. 모든 entry의 `driver_data`가 non-zero라면 이 field는 사실상 필수입니다.

ID가 추가되면 갱신된 `pci_ids` 목록의 아직 claim되지 않은 device마다 probe routine을 호출합니다. Driver 종료 시 `pci_unregister_driver()`만 호출하면 PCI layer가 driver가 처리하던 모든 device의 remove hook을 자동 호출합니다.

pci_register_driver() call
==========================

PCI device drivers call ``pci_register_driver()`` during their
initialization with a pointer to a structure describing the driver
(``struct pci_driver``):

.. kernel-doc:: include/linux/pci.h
   :functions: pci_driver

The ID table is an array of ``struct pci_device_id`` entries ending with an
all-zero entry.  Definitions with static const are generally preferred.

.. kernel-doc:: include/linux/mod_devicetable.h
   :functions: pci_device_id

Most drivers only need ``PCI_DEVICE()`` or ``PCI_DEVICE_CLASS()`` to set up
a pci_device_id table.

New PCI IDs may be added to a device driver pci_ids table at runtime
as shown below::

  echo "vendor device subvendor subdevice class class_mask driver_data" > \
  /sys/bus/pci/drivers/{driver}/new_id

All fields are passed in as hexadecimal values (no leading 0x).
The vendor and device fields are mandatory, the others are optional. Users
need pass only as many optional fields as necessary:

  - subvendor and subdevice fields default to PCI_ANY_ID (FFFFFFFF)
  - class and classmask fields default to 0
  - driver_data defaults to 0UL.
  - override_only field defaults to 0.

Note that driver_data must match the value used by any of the pci_device_id
entries defined in the driver. This makes the driver_data field mandatory
if all the pci_device_id entries have a non-zero driver_data value.

Once added, the driver probe routine will be invoked for any unclaimed
PCI devices listed in its (newly updated) pci_ids list.

When the driver exits, it just calls pci_unregister_driver() and the PCI layer
automatically calls the remove hook for all devices handled by the driver.

초기화·종료 attribute

119-141

`<linux/init.h>`의 `__init`은 driver 초기화 뒤 버리는 initialization code, `__exit`는 non-modular driver에서 무시하는 exit code를 표시합니다.

`module_init()`·`module_exit()` function과 오직 이들이 호출하는 initialization function에 `__init`·`__exit`를 붙입니다. `struct pci_driver`에는 붙이지 마십시오. 어떤 mark인지 확실하지 않다면 잘못 붙이는 것보다 붙이지 않는 편이 낫습니다.

Driver attribute
Attribute의미
__init초기화 뒤 제거되는 code
__exitModule exit code, built-in driver에서는 무시

Code lifetime에 따른 표시입니다.

"Attributes" for driver functions/data
--------------------------------------

Please mark the initialization and cleanup functions where appropriate
(the corresponding macros are defined in <linux/init.h>):

        ======                =================================================
        __init                Initialization code. Thrown away after the driver
                        initializes.
        __exit                Exit code. Ignored for non-modular drivers.
        ======                =================================================

Tips on when/where to use the above attributes:
        - The module_init()/module_exit() functions (and all
          initialization functions called _only_ from these)
          should be marked __init/__exit.

        - Do not mark the struct pci_driver.

        - Do NOT mark a function if you are not sure which mark to use.
          Better to not mark the function than mark the function wrong.

Device 초기화와 config access

176-196

대부분의 PCI driver 초기화는 device enable, MMIO/IOP 요청, streaming·coherent DMA mask, shared control data, config space, IRQ, non-PCI 부분, DMA engine 순입니다.

Driver는 거의 언제든 PCI config space register에 접근할 수 있습니다. 다만 BIST 실행 중에는 config space가 사라질 수 있고, 이때 PCI Bus Master Abort가 발생해 config read가 garbage를 반환합니다.

Device Initialization Steps
===========================

As noted in the introduction, most PCI drivers need the following steps
for device initialization:

  - Enable the device
  - Request MMIO/IOP resources
  - Set the DMA mask size (for both coherent and streaming DMA)
  - Allocate and initialize shared control data (pci_allocate_coherent())
  - Access device configuration space (if needed)
  - Register IRQ handler (request_irq())
  - Initialize non-PCI (i.e. LAN/SCSI/etc parts of the chip)
  - Enable DMA/processing engines.

The driver can access PCI config space registers at any time.
(Well, almost. When running BIST, config space can go away...but
that will just result in a PCI Bus Master Abort and config reads
will return garbage).

PCI device enable·bus master·MWI

197-235

Device register를 만지기 전에 `pci_enable_device()`를 호출해야 합니다. Suspended device를 깨우고 BIOS가 하지 않았다면 I/O·memory region과 IRQ를 할당합니다. 실패할 수 있으므로 반드시 return value를 확인하십시오.

알려진 OS bug로 resource allocation을 확인하기 전에 resource를 enable합니다. `pci_request_resources()`를 먼저 부르는 편이 논리적이지만 현재 driver는 두 device에 같은 range가 배정된 문제를 감지하지 못합니다. 흔하지 않고 곧 고쳐질 가능성도 낮습니다.

`pci_set_master()`는 `PCI_COMMAND`의 bus master bit를 설정해 DMA를 enable하고 BIOS가 잘못 지정한 latency timer도 고칩니다. `pci_clear_master()`는 bit를 지워 DMA를 disable합니다.

PCI Memory-Write-Invalidate transaction을 지원하면 `pci_set_mwi()`를 호출합니다. `PCI_COMMAND`의 Mem-Wr-Inval bit를 enable하고 cache line size register도 바로 설정합니다. Architecture나 chipset이 지원하지 않을 수 있으므로 return value를 확인합니다. 필수는 아니라면 `pci_try_set_mwi()`로 best effort를 요청합니다.

Device enable helper
API효과
pci_enable_device()Power·resource·IRQ 준비
pci_set_master()Bus master DMA enable
pci_clear_master()Bus master DMA disable
pci_set_mwi()MWI 필수 enable
pci_try_set_mwi()MWI best effort

초기화 단계의 hardware enable 기능입니다.

Enable the PCI device
---------------------
Before touching any device registers, the driver needs to enable
the PCI device by calling pci_enable_device(). This will:

  - wake up the device if it was in suspended state,
  - allocate I/O and memory regions of the device (if BIOS did not),
  - allocate an IRQ (if BIOS did not).

.. note::
   pci_enable_device() can fail! Check the return value.

.. warning::
   OS BUG: we don't check resource allocations before enabling those
   resources. The sequence would make more sense if we called
   pci_request_resources() before calling pci_enable_device().
   Currently, the device drivers can't detect the bug when two
   devices have been allocated the same range. This is not a common
   problem and unlikely to get fixed soon.

   This has been discussed before but not changed as of 2.6.19:
   https://lore.kernel.org/r/20060302180025.GC28895@flint.arm.linux.org.uk/


pci_set_master() will enable DMA by setting the bus master bit
in the PCI_COMMAND register. It also fixes the latency timer value if
it's set to something bogus by the BIOS.  pci_clear_master() will
disable DMA by clearing the bus master bit.

If the PCI device can use the PCI Memory-Write-Invalidate transaction,
call pci_set_mwi().  This enables the PCI_COMMAND bit for Mem-Wr-Inval
and also ensures that the cache line size register is set correctly.
Check the return value of pci_set_mwi() as not all architectures
or chip-sets may support Memory-Write-Invalidate.  Alternatively,
if Mem-Wr-Inval would be nice to have but is not required, call
pci_try_set_mwi() to have the system do its best effort at enabling
Mem-Wr-Inval.

MMIO·IOP resource 요청

236-264

MMIO memory와 I/O port address를 PCI config space에서 직접 읽지 마십시오. Architecture·chipset code가 PCI bus address를 host physical address로 remap했을 수 있으므로 `pci_dev`의 값을 사용합니다. Register와 device memory 접근법은 `Documentation/driver-api/io-mapping.rst`를 참조하십시오.

Driver는 `pci_request_region()`으로 같은 address resource를 다른 device가 사용하지 않는지 확인해야 합니다. 반대로 `pci_disable_device()` 호출 뒤 `pci_release_region()`으로 해제합니다. 두 device의 address range 충돌을 막기 위한 것입니다.

현재 driver는 `pci_enable_device()` 뒤에야 MMIO와 I/O Port resource 가용성을 판단할 수 있습니다.

Normal PCI BAR에 설명되지 않은 resource에는 MMIO용 `request_mem_region()`, I/O Port용 `request_region()`을 사용합니다. `pci_request_selected_regions()`도 참조하십시오.

Request MMIO/IOP resources
--------------------------
Memory (MMIO), and I/O port addresses should NOT be read directly
from the PCI device config space. Use the values in the pci_dev structure
as the PCI "bus address" might have been remapped to a "host physical"
address by the arch/chip-set specific kernel support.

See Documentation/driver-api/io-mapping.rst for how to access device registers
or device memory.

The device driver needs to call pci_request_region() to verify
no other device is already using the same address resource.
Conversely, drivers should call pci_release_region() AFTER
calling pci_disable_device().
The idea is to prevent two devices colliding on the same address range.

.. tip::
   See OS BUG comment above. Currently (2.6.19), The driver can only
   determine MMIO and IO Port resource availability _after_ calling
   pci_enable_device().

Generic flavors of pci_request_region() are request_mem_region()
(for MMIO ranges) and request_region() (for IO Port ranges).
Use these for address resources that are not described by "normal" PCI
BARs.

Also see pci_request_selected_regions() below.

DMA mask 설정

265-291

권위 있는 DMA interface 설명은 `Documentation/core-api/dma-api.rst`를 참조하십시오. 이 절은 driver가 device DMA capability를 표시해야 한다는 알림입니다.

모든 driver는 PCI bus master의 32-bit 또는 64-bit DMA capability를 명시해야 합니다. Streaming data에 32-bit보다 넓은 bus master capability가 있으면 적절한 parameter로 `dma_set_mask()`를 호출해 등록합니다. System RAM이 4GiB physical address 위에 있는 system에서 더 효율적인 DMA가 가능합니다.

PCI-X와 PCIe compliant device는 64-bit DMA device이므로 모든 해당 driver가 `dma_set_mask()`를 호출해야 합니다.

Device가 4GiB 위 System RAM의 coherent memory를 직접 address할 수 있으면 `dma_set_coherent_mask()`로 등록합니다. 초기 64-bit PCI device와 일부 PCI-X device는 payload streaming DMA는 64-bit지만 control coherent DMA는 그렇지 않을 수 있습니다.

DMA capability
용도API주의
Payload / streamingdma_set_mask()PCI-X·PCIe driver 필수
Control / coherentdma_set_coherent_mask()Streaming보다 좁을 수 있음

Streaming과 coherent address width를 따로 등록합니다.

Set the DMA mask size
---------------------
.. note::
   If anything below doesn't make sense, please refer to
   Documentation/core-api/dma-api.rst. This section is just a reminder that
   drivers need to indicate DMA capabilities of the device and is not
   an authoritative source for DMA interfaces.

While all drivers should explicitly indicate the DMA capability
(e.g. 32 or 64 bit) of the PCI bus master, devices with more than
32-bit bus master capability for streaming data need the driver
to "register" this capability by calling dma_set_mask() with
appropriate parameters.  In general this allows more efficient DMA
on systems where System RAM exists above 4G _physical_ address.

Drivers for all PCI-X and PCIe compliant devices must call
dma_set_mask() as they are 64-bit DMA devices.

Similarly, drivers must also "register" this capability if the device
can directly address "coherent memory" in System RAM above 4G physical
address by calling dma_set_coherent_mask().
Again, this includes drivers for all PCI-X and PCIe compliant devices.
Many 64-bit "PCI" devices (before PCI-X) and some PCI-X devices are
64-bit DMA capable for payload ("streaming") data but not control
("coherent") data.

Shared control data 설정

292-299

DMA mask를 설정한 뒤 driver는 coherent, 즉 shared memory를 할당할 수 있습니다. DMA API 전체 설명은 `Documentation/core-api/dma-api.rst`에 있습니다. Device에서 DMA를 enable하기 전에 완료해야 합니다.

Setup shared control data
-------------------------
Once the DMA masks are set, the driver can allocate "coherent" (a.k.a. shared)
memory.  See Documentation/core-api/dma-api.rst for a full description of
the DMA APIs. This section is just a reminder that it needs to be done
before enabling DMA on the device.

Device register 초기화

300-306

일부 driver는 특정 capability field를 program하거나 vendor-specific register를 초기화·reset해야 합니다. Pending interrupt를 clear하는 것이 예입니다.

Initialize device registers
---------------------------
Some drivers will need specific "capability" fields programmed
or other "vendor specific" register initialized or reset.
E.g. clearing pending interrupts.

IRQ handler와 MSI/MSI-X

307-360

`request_irq()`는 여기서 마지막 단계로 설명되지만 device 초기화의 중간 단계일 때가 많고, device를 실제 open할 때까지 미룰 수도 있습니다.

모든 PCI IRQ line을 공유할 수 있으므로 handler는 `IRQF_SHARED`로 등록하고 `devid`로 IRQ를 device에 mapping해야 합니다. `request_irq()`는 interrupt number에 handler와 device handle을 연결하고 interrupt를 enable합니다. 등록 전에 device를 quiesce하고 pending interrupt가 없어야 합니다.

전통적인 interrupt number는 PCI device에서 interrupt controller로 가는 IRQ line이지만 MSI/MSI-X에서는 CPU vector입니다.

MSI와 MSI-X는 Local APIC로 DMA write를 보내 interrupt를 전달합니다. MSI는 연속 vector block이 필요하고 MSI-X는 개별 vector 여러 개를 할당할 수 있습니다.

`request_irq()` 전에 `PCI_IRQ_MSI`·`PCI_IRQ_MSIX` flag로 `pci_alloc_irq_vectors()`를 호출합니다. 많은 architecture·chipset·BIOS가 MSI/MSI-X를 지원하지 않아 두 flag만 쓰면 실패할 수 있으므로 가능하면 `PCI_IRQ_INTX`도 지정합니다.

MSI/MSI-X와 legacy INTx handler가 다르면 할당 뒤 `pci_dev`의 `msi_enabled`와 `msix_enabled` flag로 올바른 handler를 선택합니다.

MSI는 독점 vector이므로 handler가 자기 device가 interrupt 원인인지 확인하지 않아도 됩니다. 또한 MSI 전달 시 DMA가 host CPU에 보인다는 ordering 보장으로 stale control data와 DMA/IRQ race를 피하고, DMA stream flush용 MMIO read를 생략할 수 있습니다.

사용 예는 `drivers/infiniband/hw/mthca/`와 `drivers/net/tg3.c`를 참조하십시오.

Interrupt 방식
방식할당공유Ordering
INTxIRQ line가능추가 flush 필요 가능
MSI연속 CPU vector독점DMA visibility 보장
MSI-X개별 CPU vector독점DMA visibility 보장

Legacy INTx와 message-signaled 방식의 차이입니다.

Register IRQ handler
--------------------
While calling request_irq() is the last step described here,
this is often just another intermediate step to initialize a device.
This step can often be deferred until the device is opened for use.

All interrupt handlers for IRQ lines should be registered with IRQF_SHARED
and use the devid to map IRQs to devices (remember that all PCI IRQ lines
can be shared).

request_irq() will associate an interrupt handler and device handle
with an interrupt number. Historically interrupt numbers represent
IRQ lines which run from the PCI device to the Interrupt controller.
With MSI and MSI-X (more below) the interrupt number is a CPU "vector".

request_irq() also enables the interrupt. Make sure the device is
quiesced and does not have any interrupts pending before registering
the interrupt handler.

MSI and MSI-X are PCI capabilities. Both are "Message Signaled Interrupts"
which deliver interrupts to the CPU via a DMA write to a Local APIC.
The fundamental difference between MSI and MSI-X is how multiple
"vectors" get allocated. MSI requires contiguous blocks of vectors
while MSI-X can allocate several individual ones.

MSI capability can be enabled by calling pci_alloc_irq_vectors() with the
PCI_IRQ_MSI and/or PCI_IRQ_MSIX flags before calling request_irq(). This
causes the PCI support to program CPU vector data into the PCI device
capability registers. Many architectures, chip-sets, or BIOSes do NOT
support MSI or MSI-X and a call to pci_alloc_irq_vectors with just
the PCI_IRQ_MSI and PCI_IRQ_MSIX flags will fail, so try to always
specify PCI_IRQ_INTX as well.

Drivers that have different interrupt handlers for MSI/MSI-X and
legacy INTx should chose the right one based on the msi_enabled
and msix_enabled flags in the pci_dev structure after calling
pci_alloc_irq_vectors.

There are (at least) two really good reasons for using MSI:

1) MSI is an exclusive interrupt vector by definition.
   This means the interrupt handler doesn't have to verify
   its device caused the interrupt.

2) MSI avoids DMA/IRQ race conditions. DMA to host memory is guaranteed
   to be visible to the host CPU(s) when the MSI is delivered. This
   is important for both data coherency and avoiding stale control data.
   This guarantee allows the driver to omit MMIO reads to flush
   the DMA stream.

See drivers/infiniband/hw/mthca/ or drivers/net/tg3.c for examples
of MSI/MSI-X usage.

PCI device shutdown 순서

361-375

PCI driver unload 시 device IRQ 생성 disable, `free_irq()`, 모든 DMA 정지, streaming·coherent buffer 해제, SCSI·netdev 같은 subsystem 등록 해제, MMIO/IO Port 응답 disable, resource 해제를 수행해야 합니다.

Shutdown order
IRQ source stopfree_irq()DMA stopDMA buffers releaseSubsystem unregisterMMIO/IO disableRegions release

Hardware가 더 이상 memory나 IRQ를 건드리지 못하게 한 뒤 software resource를 해제합니다.

PCI device shutdown
===================

When a PCI device driver is being unloaded, most of the following
steps need to be performed:

  - Disable the device from generating IRQs
  - Release the IRQ (free_irq())
  - Stop all DMA activity
  - Release DMA buffers (both streaming and coherent)
  - Unregister from other subsystems (e.g. scsi or netdev)
  - Disable device from responding to MMIO/IO Port addresses
  - Release MMIO/IO Port resource(s)

Device IRQ 중지

376-395

IRQ 중지 방법은 chip·device별입니다. 하지 않으면 shared IRQ에서 screaming interrupt가 발생할 수 있습니다.

Shared handler를 unhook해도 다른 device 때문에 IRQ line은 enable 상태입니다. 제거된 device가 line을 assert하면 남은 device의 interrupt로 오인하지만 아무 handler도 처리하지 않아 system이 100,000회 반복 뒤 IRQ를 mask할 때까지 hang합니다. 그러면 같은 shared IRQ의 나머지 device도 정상 동작하지 않습니다.

MSI/MSI-X는 독점 interrupt로 정의되어 screaming interrupt 문제를 피할 수 있으므로 이를 사용할 또 하나의 이유입니다.

Stop IRQs on the device
-----------------------
How to do this is chip/device specific. If it's not done, it opens
the possibility of a "screaming interrupt" if (and only if)
the IRQ is shared with another device.

When the shared IRQ handler is "unhooked", the remaining devices
using the same IRQ line will still need the IRQ enabled. Thus if the
"unhooked" device asserts IRQ line, the system will respond assuming
it was one of the remaining devices asserted the IRQ line. Since none
of the other devices will handle the IRQ, the system will "hang" until
it decides the IRQ isn't going to get handled and masks the IRQ (100,000
iterations later). Once the shared IRQ is masked, the remaining devices
will stop functioning properly. Not a nice situation.

This is another reason to use MSI or MSI-X if it's available.
MSI and MSI-X are defined to be exclusive interrupts and thus
are not susceptible to the "screaming interrupt" problem.

IRQ 해제와 DMA 중지

396-416

Device가 quiesce되어 더 이상 IRQ를 만들지 않으면 `free_irq()`를 호출합니다. Pending IRQ 처리가 끝난 뒤 handler를 unhook하고 다른 사용자가 없으면 IRQ를 해제합니다.

DMA control data를 deallocate하기 전에 모든 DMA operation을 중지하는 것은 매우 중요합니다. 실패하면 memory corruption, hang, 일부 chipset의 hard crash가 발생할 수 있습니다.

IRQ를 먼저 중지한 뒤 DMA를 멈추면 handler가 DMA engine을 다시 시작하는 race를 피할 수 있습니다. 명백해 보이지만 과거 여러 성숙한 driver도 잘못 처리했습니다.

Release the IRQ
---------------
Once the device is quiesced (no more IRQs), one can call free_irq().
This function will return control once any pending IRQs are handled,
"unhook" the drivers IRQ handler from that IRQ, and finally release
the IRQ if no one else is using it.


Stop all DMA activity
---------------------
It's extremely important to stop all DMA operations BEFORE attempting
to deallocate DMA control data. Failure to do so can result in memory
corruption, hangs, and on some chip-sets a hard crash.

Stopping DMA after stopping the IRQs can avoid races where the
IRQ handler might restart DMA engines.

While this step sounds obvious and trivial, several "mature" drivers
didn't get this step right in the past.

DMA buffer와 subsystem 정리

417-436

DMA를 멈춘 뒤 streaming DMA부터 정리합니다. Data buffer를 unmap하고 upstream owner가 있으면 반환합니다. 그 다음 control data가 든 coherent buffer를 정리합니다. Unmap interface는 `Documentation/core-api/dma-api.rst`를 참조하십시오.

대부분 low-level PCI driver는 USB, ALSA, SCSI, NetDev, Infiniband 같은 다른 subsystem을 지원합니다. 해당 subsystem resource를 빠뜨리지 마십시오. 빠뜨리면 unload된 driver를 subsystem이 호출할 때 Oops 또는 panic이 나는 경우가 많습니다.

Release DMA buffers
-------------------
Once DMA is stopped, clean up streaming DMA first.
I.e. unmap data buffers and return buffers to "upstream"
owners if there is one.

Then clean up "coherent" buffers which contain the control data.

See Documentation/core-api/dma-api.rst for details on unmapping interfaces.


Unregister from other subsystems
--------------------------------
Most low level PCI device drivers support some other subsystem
like USB, ALSA, SCSI, NetDev, Infiniband, etc. Make sure your
driver isn't losing resources from that other subsystem.
If this happens, typically the symptom is an Oops (panic) when
the subsystem attempts to call into a driver that has been unloaded.

Device disable과 region 해제

437-449

MMIO 또는 I/O Port resource를 `io_unmap()`한 뒤 `pci_disable_device()`를 호출합니다. 이는 `pci_enable_device()`의 대칭 반대이며 이후 device register에 접근하면 안 됩니다.

`pci_release_region()`으로 MMIO 또는 I/O Port range를 사용 가능 상태로 표시합니다. 하지 않으면 보통 driver를 다시 load할 수 없습니다.

Disable Device from responding to MMIO/IO Port addresses
--------------------------------------------------------
io_unmap() MMIO or IO Port resources and then call pci_disable_device().
This is the symmetric opposite of pci_enable_device().
Do not access device registers after calling pci_disable_device().


Release MMIO/IO Port Resource(s)
--------------------------------
Call pci_release_region() to mark the MMIO or IO Port range as available.
Failure to do so usually results in the inability to reload the driver.

PCI config space 접근

450-470

`struct pci_dev *`가 있으면 `pci_(read|write)_config_(byte|word|dword)`로 config space에 접근합니다. 성공 시 0, 실패 시 `PCIBIOS_...` error code를 반환하며 `pcibios_strerror`로 문자열로 바꿀 수 있습니다. 대개 driver는 유효한 PCI device 접근이 실패하지 않는다고 봅니다.

`pci_dev`가 없으면 `pci_bus_(read|write)_config_(byte|word|dword)`로 해당 bus의 device와 function에 접근합니다.

Standard config header field에는 `<linux/pci.h>`의 location·bit symbolic name을 사용하십시오. Extended PCI Capability register는 capability별 `pci_find_capability()`로 해당 register block을 찾습니다.

How to access PCI config space
==============================

You can use `pci_(read|write)_config_(byte|word|dword)` to access the config
space of a device represented by `struct pci_dev *`. All these functions return
0 when successful or an error code (`PCIBIOS_...`) which can be translated to a
text string by pcibios_strerror. Most drivers expect that accesses to valid PCI
devices don't fail.

If you don't have a struct pci_dev available, you can call
`pci_bus_(read|write)_config_(byte|word|dword)` to access a given device
and function on that bus.

If you access fields in the standard portion of the config header, please
use symbolic names of locations and bits declared in <linux/pci.h>.

If you need to access Extended PCI Capability registers, just call
pci_find_capability() for the particular capability and it will find the
corresponding register block for you.

기타 유용한 function

471-490
PCI helper
Function역할
pci_get_domain_bus_and_slot()Domain·bus·slot·number로 pci_dev를 찾고 reference 증가
pci_set_power_state()D0부터 D3까지 power state 설정
pci_find_capability()Capability list에서 지정 capability 검색
pci_resource_start()PCI region bus 시작 address
pci_resource_end()PCI region bus 끝 address
pci_resource_len()PCI region byte 길이
pci_set_drvdata()pci_dev private driver data 설정
pci_get_drvdata()Private driver data 반환
pci_set_mwi()Memory-Write-Invalidate enable
pci_clear_mwi()Memory-Write-Invalidate disable

Device lookup, power, capability, resource와 private data helper입니다.

Other interesting functions
===========================

=============================        ================================================
pci_get_domain_bus_and_slot()        Find pci_dev corresponding to given domain,
                                bus and slot and number. If the device is
                                found, its reference count is increased.
pci_set_power_state()                Set PCI Power Management state (0=D0 ... 3=D3)
pci_find_capability()                Find specified capability in device's capability
                                list.
pci_resource_start()                Returns bus start address for a given PCI region
pci_resource_end()                Returns bus end address for a given PCI region
pci_resource_len()                Returns the byte length of a PCI region
pci_set_drvdata()                Set private driver data pointer for a pci_dev
pci_get_drvdata()                Return private driver data pointer for a pci_dev
pci_set_mwi()                        Enable Memory-Write-Invalidate transactions.
pci_clear_mwi()                        Disable Memory-Write-Invalidate transactions.
=============================        ================================================

기타 힌트

491-507

사용자에게 PCI device 이름을 표시할 때는 `pci_name(pci_dev)`를 사용하십시오.

PCI device는 항상 `pci_dev` structure pointer로 참조하십시오. 모든 PCI layer function이 이 식별을 사용하며 유일하게 합리적입니다. Multi-primary-bus system에서는 의미가 복잡하므로 특별한 목적 외에는 bus/slot/function number를 쓰지 마십시오.

Driver에서 Fast Back to Back write를 켜지 마십시오. Bus의 모든 device가 지원해야 하므로 개별 driver가 아니라 platform과 generic code가 처리해야 합니다.

Miscellaneous hints
===================

When displaying PCI device names to the user (for example when a driver wants
to tell the user what card has it found), please use pci_name(pci_dev).

Always refer to the PCI devices by a pointer to the pci_dev structure.
All PCI layer functions use this identification and it's the only
reasonable one. Don't use bus/slot/function numbers except for very
special purposes -- on systems with multiple primary buses their semantics
can be pretty complex.

Don't try to turn on Fast Back to Back writes in your driver.  All devices
on the bus need to be capable of doing it, so this is something which needs
to be handled by platform and generic code, not individual drivers.

Vendor·device ID

508-521

여러 driver가 공유하지 않는 새 device 또는 vendor ID를 `include/linux/pci_ids.h`에 추가하지 마십시오. 유용하면 driver 내부 private definition을 만들거나 plain hexadecimal constant를 사용합니다.

Device ID는 vendor가 정하는 임의 hexadecimal number이며 보통 `pci_device_id` table 한 곳에서만 사용합니다.

새 vendor/device ID는 `https://pci-ids.ucw.cz/`에 제출하십시오. `pci.ids` mirror는 `https://github.com/pciutils/pciids`에 있습니다.

Vendor and device identifications
=================================

Do not add new device or vendor IDs to include/linux/pci_ids.h unless they
are shared across multiple drivers.  You can add private definitions in
your driver if they're helpful, or just use plain hex constants.

The device IDs are arbitrary hex numbers (vendor controlled) and normally used
only in a single location, the pci_device_id table.

Please DO submit new vendor/device IDs to https://pci-ids.ucw.cz/.
There's a mirror of the pci.ids file at https://github.com/pciutils/pciids.

폐기된 function

522-540

오래된 driver를 새 PCI interface로 port할 때 만날 수 있는 아래 function은 hotplug, PCI domain, 올바른 locking과 호환되지 않아 kernel에서 제거됐습니다.

Obsolete PCI API
폐기 API대체 API
pci_find_device()pci_get_device()
pci_find_subsys()pci_get_subsys()
pci_find_slot()pci_get_domain_bus_and_slot()
pci_get_slot()pci_get_domain_bus_and_slot()

현재 사용할 대체 helper입니다.

전통적으로 PCI device list를 직접 순회하는 driver도 여전히 가능하지만 권장하지 않습니다.

Obsolete functions
==================

There are several functions which you might come across when trying to
port an old driver to the new PCI interface.  They are no longer present
in the kernel as they aren't compatible with hotplug or PCI domains or
having sane locking.

=================        ===========================================
pci_find_device()        Superseded by pci_get_device()
pci_find_subsys()        Superseded by pci_get_subsys()
pci_find_slot()                Superseded by pci_get_domain_bus_and_slot()
pci_get_slot()                Superseded by pci_get_domain_bus_and_slot()
=================        ===========================================

The alternative is the traditional PCI device driver that walks PCI
device lists. This is still possible but discouraged.

MMIO와 write posting

541-572

Driver를 I/O Port space에서 MMIO space로 바꾸면 write posting을 처리해야 합니다. `tg3`, `acenic`, `sym53c8xx_2` 등 많은 driver가 이미 처리합니다.

I/O Port write는 CPU가 계속 실행하기 전에 transaction이 PCI device에 도착함을 보장합니다. MMIO write는 도착 전에 CPU가 계속할 수 있습니다. Transaction이 destination에 닿기 전에 write completion을 CPU에 post하기 때문에 write posting이라고 합니다.

따라서 timing-sensitive code에서 CPU가 기다려야 하는 지점에는 `readl()` 같은 read를 추가해야 합니다. I/O Port bit-banging은 write 후 delay만으로 충분합니다.

for (i = 8; --i; val >>= 1) {
        outb(val & 1, ioport_reg);      /* write bit */
        udelay(10);
}

MMIO에서는 write 뒤 안전한 register를 read해 posted write를 flush한 다음 delay합니다.

for (i = 8; --i; val >>= 1) {
        writeb(val & 1, mmio_reg);      /* write bit */
        readb(safe_mmio_reg);           /* flush posted write */
        udelay(10);
}

`safe_mmio_reg`의 read에는 device 정상 동작을 방해하는 side effect가 없어야 합니다.

MMIO posted write flush
writeb(mmio_reg)readb(safe_mmio_reg)Posted write 도착 보장udelay()

Timing-sensitive MMIO는 명시적 read로 write completion을 기다립니다.

MMIO Space and "Write Posting"
==============================

Converting a driver from using I/O Port space to using MMIO space
often requires some additional changes. Specifically, "write posting"
needs to be handled. Many drivers (e.g. tg3, acenic, sym53c8xx_2)
already do this. I/O Port space guarantees write transactions reach the PCI
device before the CPU can continue. Writes to MMIO space allow the CPU
to continue before the transaction reaches the PCI device. HW weenies
call this "Write Posting" because the write completion is "posted" to
the CPU before the transaction has reached its destination.

Thus, timing sensitive code should add readl() where the CPU is
expected to wait before doing other work.  The classic "bit banging"
sequence works fine for I/O Port space::

       for (i = 8; --i; val >>= 1) {
               outb(val & 1, ioport_reg);      /* write bit */
               udelay(10);
       }

The same sequence for MMIO space should be::

       for (i = 8; --i; val >>= 1) {
               writeb(val & 1, mmio_reg);      /* write bit */
               readb(safe_mmio_reg);           /* flush posted write */
               udelay(10);
       }

It is important that "safe_mmio_reg" not have any side effects that
interferes with the correct operation of the device.

PCI reset 시 write flush

573-578

PCI device reset에서도 주의해야 합니다. `writel()`을 flush할 때 PCI Configuration space read를 사용하면 device가 `readl()`에 응답하지 않아 발생하는 PCI master abort를 모든 platform에서 안전하게 처리할 수 있습니다.

대부분 x86 platform은 MMIO read master abort를 soft fail로 처리해 `~0` 같은 garbage를 반환하지만, 많은 RISC platform에서는 hard fail로 crash합니다.

Another case to watch out for is when resetting a PCI device. Use PCI
Configuration space reads to flush the writel(). This will gracefully
handle the PCI master abort on all platforms if the PCI device is
expected to not respond to a readl().  Most x86 platforms will allow
MMIO reads to master abort (a.k.a. "Soft Fail") and return garbage
(e.g. ~0). But many RISC platforms will crash (a.k.a."Hard Fail").