From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:17
The Intel(r) Dynamic Load Balancer (Intel(r) DLB) is a PCIe device and hardware
accelerator that provides load-balanced, prioritized scheduling for event based
workloads across CPU cores. It can be used to save CPU resources in high
throughput pipelines by replacing software based distribution and synchronization
schemes. Using Intel DLB has been demonstrated to reduce CPU utilization up to
2-3 CPU cores per DLB device, while also improving packet processing pipeline
performance.
In many applications running on processors with a large number of cores,
workloads must be distributed (load-balanced) across a number of cores.
In packet processing applications, for example, streams of incoming packets can
exceed the capacity of any single core. So they have to be divided between
available worker cores. The workload can be split by either breaking the
processing flow into stages and places distinct stages on separate cores in a
daisy chain fashion (a pipeline), or spraying packets across multiple workers
that may be executing the same processing stage. Many systems employ a hybrid
approach whereby each packet encounters multiple pipelined stages with
distribution across multiple workers at each individual stage.
Intel DLB distribution schemes include "parallel" (packets are load-balanced
across multiple cores and processed in parallel), "ordered" (similar to
"parallel" but packets are reordered into ingress order by the device), and
"atomic" (packet flows are scheduled to a single core at a time such that
locks are not required to access per-flow data, and dynamically migrated to
ensure load-balance).
Intel DLB consists of queues and arbiters that connect producer
cores and consumer cores. The device implements load-balanced queueing
features including:
- Lock-free multi-producer/multi-consumer operation.
- Multiple priority levels for varying traffic types.
- 'Direct' traffic (i.e. multi-producer/single-consumer)
- Simple unordered load-balanced distribution.
- Atomic lock free load balancing across multiple consumers.
- Queue element reordering feature allowing ordered load-balanced
distribution.
The fundamental unit of communication through the device is a queue element
(QE), which consists of 8B of data and 8B of metadata (destination queue,
priority, etc.). The data field can be any type that fits within 8B.
A core's interface to the device, a "port," consists of a memory-mappable
region through which the core enqueues a queue entry, and an in-memory
queue (the "consumer queue") to which the device schedules QEs. Each QE
is enqueued to a device-managed queue, and from there scheduled to a port.
Software specifies the "linking" of queues and ports; i.e. which ports the
device is allowed to schedule to for a given queue. The device uses a
credit scheme to prevent overflow of the on-device queue storage.
Applications can interface directly with the device by mapping the port's
memory and MMIO regions into the application's address space for enqueue
and dequeue operations, but call into the kernel driver for configuration
operations. An application can also be polling- or interrupt-driven;
Intel DLB supports both modes of operation.
Device resources -- i.e. ports, queues, and credits -- are contained within
a scheduling domain. Scheduling domains are isolated from one another; a
port can only enqueue to and dequeue from queues within its scheduling
domain. A scheduling domain's resources are configured through a configfs
interface. Please refer to Documentation/misc-devices/dlb.rst (in patch 01/17)
for a detailed description on the DLB configfs implementation.
Intel DLB supports SR-IOV and Scalable IOV, and allows for a flexible
division of its resources among the PF and its virtual devices. The virtual
devices are incapable of configuring the device directly; they use a
hardware mailbox to proxy configuration requests to the PF driver. This
driver supports both PF and virtual devices, as there is significant code
re-use between the two, with device-specific behavior handled through a
callback interface. Virtualization support will be added in a later patch
set.
The DLB driver uses configfs and sysfs as its primary interface While the
DLB sysfs allows users to configure DLB at the device level, the configfs
lets users to create and control DLB scheduling domains, ports and queues.
Configfs supports operations on a scheduling domain's resources
(primarily resource configuration). Scheduling domains are created
dynamically by user-space software.
[1] https://builders.intel.com/docs/networkbuilders/SKU-343247-001US-queue-management-and-load-balancing-on-intel-architecture.pdf
[2] https://doc.dpdk.org/guides/prog_guide/eventdev.html
This submission is still a work in progress. We have replaced ioctl interface
with configfs and made a few other changes based on earlier reviews (see
https://lore.kernel.org/all/BYAPR11MB30952BA538BB905331392A08D9119@BYAPR11MB3095.namprd11.prod.outlook.com/).
There are a couple of issues that we would like to get help and suggestions
from reviewers and community.
1. Before a scheduling domain is created/enabled, a set of parameters are
passed to the kernel driver via configfs attribute files in an configfs domain
directory (say $domain) created by user. Each attribute file corresponds to
a configuration parameter of the domain. After writing to all the attribute
files, user writes 1 to "create" attribute, which triggers an action (i.e.,
domain creation) in the kernel driver. Since multiple processes/users can
access the $domain directory, multiple users can write to the attribute files
at the same time. How do we guarantee an atomic update/configuration of a
domain? In other words, if user A wants to set attributes 1 and 2, how can we
prevent user B from changing attribute 1 and 2 before user A writes 1 to
"create"? A configfs directory with individual attribute files seems to not
be able to provide atomic configuration in this case. One option to solve this
issue could be write a structured data (with a set of parameters) to a single
attribute file. This would guarantee the atomic configuration, but may not be
a conventional configfs operation.
2. When a user space dlb application exits, it needs to tell the kernel driver
to reset the scheduling domain and return the associated resources back to the
resource pool for a future use. The application does so by write 0 to "create"
attribute file, and it works fine with a normal program exit. However when an
application is killed (for example, kill -9) by user, we need to a way to reset
the domain in the driver. What would be the best approach for the tear down
process in this case? In our current implementation (see patch 06/17) we create
and open an anon file descriptor when a domain is created/enabled. Since the
file is closed automatically when the user application exits or is killed, we
use the file close() operation in the driver to reset and tear down the domain.
Would this approach be acceptable with the configfs implementation?
Dan points out that configfs puts atomic update and configuration teardown
responsibilities in userspacer. We are looking for a direction check on these
issues before doing deeper reworks.
v12:
- Address Dan's following feedbacks on coding stylei issues.
-- Remove DLB_HW_ERR() and DLB_HW_DBG() macros. Use dev_err() and dev_dbg()
directly.
-- Replace FIELD_SET() macro with direct coding.
-- Remove all refererences on virt_id/phys_id. They will be introduced in
future patches.
-- Use list_move() instead of list_del() and list_add() whenever possible.
-- Remove device revision numbers.
-- Reverse the order of iosubmit_cmds512() and wmb().
-- Other cleanups based on Dan's comments.
- The following coding style changes suggested by Dan will be implemented
in the next revision
-- Replace DLB_CSR_RD() and DLB_CSR_WR() with direct ioread32() and
iowrite32() call.
-- Remove bitmap wrappers and use linux bitmap functions directly.
-- Use trace_event in configfs attribute file update.
v11:
- Change the user interface from ioctls to configfs. Provide configfs
interface for create and configure scheduling domains, queues, ports,
and link/unlink of queues and ports.
- Address all of Greg's feedback on v10, including
-- Consolidate header files. Merged dlb_main.h, dlb_hw_types.h,
dlb_resources.h and dlb_bitmap.h into dlb_main.h.
-- Removed device ops callbacks. They will be added back to when
VM support is introduced.
-- Use macros/fucntions provided by linux kernel. Replace BITS_GET()
and BITS_SET() by the existing linux kernel macros FIELD_GET()
and FIELD_PREP().
-- Revise the DLB overview document dlb.rst and provide detailed
descriptions on how DLB works.
- Remove configurations for sequeuence number and class of services.
- Remove all VF and VDEV (specially in dlb_resource.c) related code. Will
add them back in a later patches.
- Add dlb sysfs for device level control and configurtion.
- Move dynamic port and queueu linking and unlinking to future patch
set. This reduces the total patches in this set to 17 from 20 in v10.
Only static linking and unlinking is supported in this submission.
v10:
- Addressed an issue reported by kernel test robot [off-list ref]
-- Add "WITH Linux-syscall-note" to the SPDX-License-Identifier in uapi
header file dlb.h.
v9:
- Addressed all of Greg's feecback on v8, including
-- Remove function name (__func__) from dev_err() messages, that could spam log.
-- Replace list and function pointer calls in dlb_ioctl() with switch-case
and real function calls for ioctl.
-- Drop the compat_ptr_ioctl in dlb_ops (struct file_operations).
-- Change ioctl magic number for DLB to unused 0x81 (from 'h').
-- Remove all placeholder/dummy functions in the patch set.
-- Re-arrange the comments in dlb.h so that the order is consistent with that
of data structures referred.
-- Correct the comments on SPDX License and DLB versions in dlb.h.
-- Replace BIT_SET() and BITS_CLR() marcos with direct coding.
-- Remove NULL pointer checking (f->private_data) in dlb_ioctl().
-- Use whole line whenever possible and not wrapping lines unnecessarily.
-- Remove __attribute__((unused)).
-- Merge dlb_ioctl.h and dlb_file.h into dlb_main.h
v8:
- Add a functional block diagram in dlb.rst
- Modify change logs to reflect the links between patches and DPDK
eventdev library.
- Add a check of power-of-2 for CQ depth.
- Move call to INIT_WORK() to dlb_open().
- Clean dlb workqueue by calling flush_scheduled_work().
- Add unmap_mapping_range() in dlb_port_close().
v7 (Intel internal version):
- Address all of Dan's feedback, including
-- Drop DLB 2.0 throughout the patch set, use DLB only.
-- Fix license and copyright statements
-- Use pcim_enable_device() and pcim_iomap_regions(), instead of
unmanaged version.
-- Move cdev_add() to dlb_init() and add all devices at once.
-- Fix Makefile, using "+=" style.
-- Remove FLR description and mention movdir64/enqcmd usage in doc.
-- Make the permission for the domain same as that for device for
ioctl access.
-- Use idr instead of ida.
-- Add a lock in dlb_close() to prevent driver unbinding while ioctl
coomands are in progress.
-- Remove wrappers that are used for code sharing between kernel driver
and DPDK.
- Address Pierre-Louis' feedback, including
-- Clean the warinings from checkpatch
-- Fix the warnings from "make W=1"
v6 (Intel internal version):
- Change the module name to dlb(from dlb2), which currently supports Intel
DLB 2.0 only.
- Address all of Pierre-Louis' feedback on v5, including
-- Consolidate the two near-identical for loops in dlb2_release_domain_memory().
-- Remove an unnecessary "port = NULL" initialization
-- Consistently use curly braces on the *_LIST_FOR macros
when the for-loop contents spans multiple lines.
-- Add a comment to the definition of DLB2FS_MAGIC
-- Remove always true if statemnets
-- Move the get_cos_bw mutex unlock call earlier to shorten the critical
section.
- Address all of Dan's feedbacks, including
-- Replace the unions for register bits access with bitmask and shifts
-- Centralize the "to/from" user memory copies for ioctl functions.
-- Review ioctl design against Documentation/process/botching-up-ioctls.rst
-- Remove wraper functions for memory barriers.
-- Use ilog() to simplify a switch code block.
-- Add base-commit to cover letter.
v5 (Intel internal version):
- Reduce the scope of the initial patch set (drop the last 8 patches)
- Further decompose some of the remaining patches into multiple patches.
- Address all of Pierre-Louis' feedback, including:
-- Move kerneldoc to *.c files
-- Fix SPDX comment style
-- Add BAR macros
-- Improve/clarify struct dlb2_dev and struct device variable naming
-- Add const where missing
-- Clarify existing comments and add new ones in various places
-- Remove unnecessary memsets and zero-initialization
-- Remove PM abstraction, fix missing pm_runtime_allow(), and don't
update PM refcnt when port files are opened and closed.
-- Convert certain ternary operations into if-statements
-- Out-line the CQ depth valid check
-- De-duplicate the logic in dlb2_release_device_memory()
-- Limit use of devm functions to allocating/freeing struct dlb2
- Address Ira's comments on dlb2.rst and correct commit messages that
don't use the imperative voice.
v4:
- Move PCI device ID definitions into dlb2_hw_types.h, drop the VF definition
- Remove dlb2_dev_list
- Remove open/close functions and fops structure (unused)
- Remove "(char *)" cast from PCI driver name
- Unwind init failures properly
- Remove ID alloc helper functions and call IDA interfaces directly instead
v3:
- Remove DLB2_PCI_REG_READ/WRITE macros
v2:
- Change driver license to GPLv2 only
- Expand Kconfig help text and remove unnecessary (R)s
- Remove unnecessary prints
- Add a new entry in ioctl-number.rst
- Convert the ioctl handler into a switch statement
- Correct some instances of IOWR that should have been IOR
- Align macro blocks
- Don't break ioctl ABI when introducing new commands
- Remove indirect pointers from ioctl data structures
- Remove the get-sched-domain-fd ioctl command
Mike Ximing Chen (17):
dlb: add skeleton for DLB driver
dlb: initialize DLB device
dlb: add resource and device initialization
dlb: add configfs interface and scheduling domain directory
dlb: add scheduling domain configuration
dlb: add domain software reset
dlb: add low-level register reset operations
dlb: add runtime power-management support
dlb: add queue create, reset, get-depth configfs interface
dlb: add register operations for queue management
dlb: add configfs interface to configure ports
dlb: add register operations for port management
dlb: add port mmap support
dlb: add start domain configfs attribute
dlb: add queue map, unmap, and pending unmap
dlb: add static queue map register operations
dlb: add basic sysfs interfaces
Documentation/ABI/testing/sysfs-driver-dlb | 116 +
Documentation/misc-devices/dlb.rst | 323 ++
Documentation/misc-devices/index.rst | 1 +
MAINTAINERS | 7 +
drivers/misc/Kconfig | 1 +
drivers/misc/Makefile | 1 +
drivers/misc/dlb/Kconfig | 18 +
drivers/misc/dlb/Makefile | 7 +
drivers/misc/dlb/dlb_args.h | 372 ++
drivers/misc/dlb/dlb_configfs.c | 1225 ++++++
drivers/misc/dlb/dlb_configfs.h | 195 +
drivers/misc/dlb/dlb_file.c | 149 +
drivers/misc/dlb/dlb_main.c | 616 +++
drivers/misc/dlb/dlb_main.h | 653 ++++
drivers/misc/dlb/dlb_pf_ops.c | 299 ++
drivers/misc/dlb/dlb_regs.h | 3640 ++++++++++++++++++
drivers/misc/dlb/dlb_resource.c | 3960 ++++++++++++++++++++
include/uapi/linux/dlb.h | 40 +
18 files changed, 11623 insertions(+)
create mode 100644 Documentation/ABI/testing/sysfs-driver-dlb
create mode 100644 Documentation/misc-devices/dlb.rst
create mode 100644 drivers/misc/dlb/Kconfig
create mode 100644 drivers/misc/dlb/Makefile
create mode 100644 drivers/misc/dlb/dlb_args.h
create mode 100644 drivers/misc/dlb/dlb_configfs.c
create mode 100644 drivers/misc/dlb/dlb_configfs.h
create mode 100644 drivers/misc/dlb/dlb_file.c
create mode 100644 drivers/misc/dlb/dlb_main.c
create mode 100644 drivers/misc/dlb/dlb_main.h
create mode 100644 drivers/misc/dlb/dlb_pf_ops.c
create mode 100644 drivers/misc/dlb/dlb_regs.h
create mode 100644 drivers/misc/dlb/dlb_resource.c
create mode 100644 include/uapi/linux/dlb.h
base-commit: 519d81956ee277b4419c723adfb154603c2565ba
--
2.27.0
@@ -0,0 +1,323 @@+.. SPDX-License-Identifier: GPL-2.0-only++===========================================+Intel(R) Dynamic Load Balancer Overview+===========================================++:Authors: Gage Eads and Mike Ximing Chen++Contents+========++- Introduction+- Scheduling+- Queue Entry+- Port+- Queue+- Credits+- Scheduling Domain+- Interrupts+- Power Management+- User Interface+- Reset++Introduction+============++The Intel(r) Dynamic Load Balancer (Intel(r) DLB) is a PCIe device and hardware+accelerator that provides load-balanced, prioritized scheduling for event based+workloads across CPU cores. It can be used to save CPU resources in high+throughput pipelines by replacing software based distribution and synchronization+schemes. Using Intel DLB in place of software based methods has been+demonstrated to reduce CPU utilization up to 2-3 CPU cores per DLB device,+while also improving packet processing pipeline performance.++In many applications running on processors with a large number of cores,+workloads must be distributed (load-balanced) across a number of cores.+In packet processing applications, for example, streams of incoming packets can+exceed the capacity of any single core. So they have to be divided between+available worker cores. The workload can be split by either breaking the+processing flow into stages and places distinct stages on separate cores in a+daisy chain fashion (a pipeline), or spraying packets across multiple workers+that may be executing the same processing stage. Many systems employ a hybrid+approach whereby each packet encounters multiple pipelined stages with+distribution across multiple workers at each individual stage.++The following diagram shows a typical packet processing pipeline with the Intel DLB.++ WC1 WC4+ +-----+ +----+ +---+ / \ +---+ / \ +---+ +----+ +-----++ |NIC | |Rx | |DLB| / \ |DLB| / \ |DLB| |Tx | |NIC |+ |Ports|---|Core|---| |-----WC2----| |-----WC5----| |---|Core|---|Ports|+ +-----+ -----+ +---+ \ / +---+ \ / +---+ +----+ ------++ \ / \ /+ WC3 WC6++WCs are the worker cores which process packets distributed by DLB. Without+hardware accelerators (such as DLB) the distribution and load balancing are normally+carried out by software running on CPU cores. Using Intel DLB in this case not only+saves the CPU resources but also improves the system performance.++The Intel DLB consists of queues and arbiters that connect producer cores (which+enqueues events to DLB) and consumer cores (which dequeue events from DLB). The+device implements load-balanced queueing features including:++- Lock-free multi-producer/multi-consumer operation.+- Multiple priority levels for varying traffic types.+- Direct traffic (i.e. multi-producer/single-consumer)+- Simple unordered load-balanced distribution.+- Atomic lock free load balancing across multiple consumers.+- Queue element reordering feature allowing ordered load-balanced distribution.++Note: this document uses 'DLB' when discussing the device hardware and 'dlb' when+discussing the driver implementation.++Following diagram illustrates the functional blocks of an Intel DLB device.++ +----++| |+ +----------+ | | +-------++ /| IQ |---|----|--/| |+ / +----------+ | | / | CP |+ / | |/ +-------++ +--------+ / | |+| | / +----------+ | /| +-------++| PP |------| IQ |---|----|---| |+ +--------+ \ +----------+ | / | | CP |+ \ |/ | +-------++ ... \ ... | |+ +--------+ \ /| | +-------++| | \+----------+ / | | | |+| PP |------| IQ |/--|----|---| CP |+ +--------+ +----------+ | | +-------++| |+ +----+ ...+PP: Producer Port |+CP: Consumer Port |+IQ: Internal Queue DLB Scheduler+++As shown in the diagram, the high-level Intel DLB data flow is as follows:+- Software threads interact with the hardware by enqueuing and dequeuing Queue+ Elements (QEs).+- QEs are sent through a Producer Port (PP) to the Intel DLB internal QE+ storage (internal queues), optionally being reordered along the way.+- The Intel DLB schedules QEs from internal queues to a consumer according to+ a two-stage priority arbiter (DLB Scheduler).+- Once scheduled, the Intel DLB writes the QE to a memory-based Consumer Port+ (CP), which the software thread reads and processes.+++Scheduling Types+================++Intel DLB supports four types of scheduling of 'events' (i.e., queue elements),+where an event can represent any type of data (e.g. a network packet). The+first, `directed`, is multi-producer/single-consumer style scheduling. The+remaining three are multi-producer/multi-consumer, and support load-balancing+across the consumers.++-`Directed`: events are scheduled to a single consumer.++-`Unordered`: events are load-balanced across consumers without any ordering+ guarantees.++-`Ordered`: events are load-balanced across consumers, and the consumer can+ re-enqueue its events so the device re-orders them into the+ original order. This scheduling type allows software to+ parallelize ordered event processing without the synchronization+ cost of re-ordering packets.++-`Atomic`: events are load-balanced across consumers, with the guarantee that+ events from a particular 'flow' are only scheduled to a single+ consumer at a time (but can migrate over time). This allows, for+ example, packet processing applications to parallelize while+ avoiding locks on per-flow data and maintaining ordering within a+ flow.++Intel DLB provides hierarchical priority scheduling, with eight priority+levels within each. Each consumer selects up to eight queues to receive events+from, and assigns a priority to each of these 'connected' queues. To schedule+an event to a consumer, the device selects the highest priority non-empty queue+of the (up to) eight connected queues. Within that queue, the device selects+the highest priority event available (selecting a lower priority event for+starvation avoidance 1% of the time, by default).++The device also supports four load-balanced scheduler classes of service. Each+class of service receives a (user-configurable) guaranteed percentage of the+scheduler bandwidth, and any unreserved bandwidth is divided evenly among the+four classes.++Queue Element+===========++Each event is contained in a queue element (QE), the fundamental unit of+communication through the device, which consists of 8B of data and 8B of+metadata, as depicted below.++QE structure format+::+ data :64+ opaque :16+ qid :8+ sched :2+ priority :3+ msg_type :3+ lock_id :16+ rsvd :8+ cmd :8++The `data` field can be any type that fits within 8B (pointer, integer,+etc.); DLB merely copies this field from producer to consumer. The+`opaque` and `msg_type` fields behave the same way.++`qid` is set by the producer to specify to which DLB internal queue it wishes+to enqueue this QE. The ID spaces for load-balanced and directed queues are both+zero-based.++`sched` controls the scheduling type: atomic, unordered, ordered, or+directed. The first three scheduling types are only valid for load-balanced+queues, and the directed scheduling type is only valid for directed queues.+This field distinguishes whether `qid` is load-balanced or directed, since+their ID spaces overlap.++`priority` is the priority with which this QE should be scheduled.++`lock_id`, used for atomic scheduling and ignored for ordered and unordered+scheduling, identifies the atomic flow to which the QE belongs. When sending a+directed event, `lock_id` is simply copied like the `data`, `opaque`, and+`msg_type` fields.++`cmd` specifies the operation, such as:+- Enqueue a new QE+- Forward a QE that was dequeued+- Complete/terminate a QE that was dequeued+- Return one or more consumer queue tokens.+- Arm the port's consumer queue interrupt.++Port+====++A core's interface to the DLB is called a "port", and consists of an MMIO+region (producer port) through which the core enqueues a queue element, and an+in-memory queue (the "consumer queue" or consumer port) to which the device+schedules QEs. A core enqueues a QE to a device queue, then the device+schedules the event to a port. Software specifies the connection of queues+and ports; i.e. for each queue, to which ports the device is allowed to+schedule its events. The device uses a credit scheme to prevent overflow of+the on-device queue storage.++Applications interface directly with the device by mapping the port's memory+and MMIO regions into the application's address space for enqueue and dequeue+operations, but call into the kernel driver for configuration operations. An+application can be polling- or interrupt-driven; DLB supports both modes+of operation.++Internal Queue+==============++A DLB device supports an implementation specific and runtime discoverable+number of load-balanced (i.e. capable of atomic, ordered, and unordered+scheduling) and directed queues. Each internal queue supports a set of+priority levels.++A load-balanced queue is capable of scheduling its events to any combination+of load-balanced ports, whereas each directed queue can only haveone-to-one+mapping with any directed port. There is no restriction on port or queue types+when a port enqueues an event to a queue; that is, a load-balanced port can+enqueue to a directed queue and vice versa.++Credits+=======++The Intel DLB uses a credit scheme to prevent overflow of the on-device+queue storage, with separate credits for load-balanced and directed queues. A+port spends one credit when it enqueues a QE, and one credit is replenished+when a QE is dequeued from a consumer queue. Each scheduling domain has one pool+of load-balanced credits and one pool of directed credits; software is+responsible for managing the allocation and replenishment of these credits among+the scheduling domain's ports.++Scheduling Domain+=================++Device resources -- including ports, queues, and credits -- are contained+within a scheduling domain. Scheduling domains are isolated from one another; a+port can only enqueue to and dequeue from queues within its scheduling domain.++The scheduling domain with a set of resources is created through configfs, and+can be accessed/shared by multiple processes.++Consumer Queue Interrupts+=========================++Each port has its own interrupt which fires, if armed, when the consumer queue+depth becomes non-zero. Software arms an interrupt by enqueueing a special+'interrupt arm' command to the device through the port's MMIO window.++Power Management+================++The kernel driver keeps the device in D3Hot (power save mode) when not in use.+The driver transitions the device to D0 when the first device file is opened,+and keeps it there until there are no open device files or memory mappings.++User Interface+==============++The dlb driver uses configfs and sysfs as its primary user interfaces. While+the sysfs is used to configure and inquire device-wide operation and+resources, the configfs provides domain/queue/port level configuration and+resource management.++The dlb device level sysfs files are created during driver probe and is located+at /sys/class/dlb/dlb<N>/device, where N is the zero-based device ID. The+configfs directories/files can be created by user applications at+/sys/kernel/config/dlb/dlb<N> using 'mkdir'. For example, 'mkdir domain0' will+create a /domain0 directory and associated files in the configfs. Within the+domain directory, directories for queues and ports can be created. An example of+a DLB configfs structure is shown in the following diagram.++ config+ |+ dlb+ |+ +------+------+------+---+| | | |+ dlb0 dlb1 dlb2 dlb3+ |+ +-----------+--+--------+-------+| | |+ domain0 domain1 domain2+ |+ +-------+-----+------------+---------------+------------+------------+| | | | |+ num_ldb_queues port0 port1 ... queue0 queue1 ...+ num_ldb_ports | |+ ... is_ldb num_sequence_numbers+ create cq_depth num_qid_inflights+ start ... num_atomic_iflights+ enable ...+ create create+++To create a domain/queue/port in DLB, an application can configure the resources+by writing to corresponding files, and then write '1' to the 'create' file to+trigger the action in the driver.++The driver also exports an mmap interface through port files, which are+acquired through port configfs. This mmap interface is used to map+a port's memory and MMIO window into the process's address space. Once the+ports are mapped, applications may use 64-byte direct-store instructions such+as movdir64b to enqueue the events for better performance.++Reset+=====++The dlb driver currently supports scheduling domain reset.++Scheduling domain reset occurs when an application stops using its domain.+Specifically, when no more file references or memory mappings exist. At this+time, the driver resets all the domain's resources (flushes its queues and+ports) and puts them in their respective available-resource lists for later+use.
@@ -0,0 +1,156 @@+// SPDX-License-Identifier: GPL-2.0-only+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */++#include<linux/aer.h>+#include<linux/cdev.h>+#include<linux/delay.h>+#include<linux/fs.h>+#include<linux/init.h>+#include<linux/module.h>+#include<linux/pci.h>+#include<linux/uaccess.h>++#include"dlb_main.h"++MODULE_LICENSE("GPL v2");+MODULE_DESCRIPTION("Intel(R) Dynamic Load Balancer (DLB) Driver");++staticstructclass*dlb_class;+staticdev_tdlb_devt;+staticDEFINE_IDR(dlb_ids);+staticDEFINE_MUTEX(dlb_ids_lock);++/**********************************/+/****** PCI driver callbacks ******/+/**********************************/++staticintdlb_probe(structpci_dev*pdev,conststructpci_device_id*pdev_id)+{+structdlb*dlb;+intret;++dlb=devm_kzalloc(&pdev->dev,sizeof(*dlb),GFP_KERNEL);+if(!dlb)+return-ENOMEM;++pci_set_drvdata(pdev,dlb);++dlb->pdev=pdev;++mutex_lock(&dlb_ids_lock);+dlb->id=idr_alloc(&dlb_ids,(void*)dlb,0,DLB_MAX_NUM_DEVICES-1,+GFP_KERNEL);+mutex_unlock(&dlb_ids_lock);++if(dlb->id<0){+dev_err(&pdev->dev,"device ID allocation failed\n");++ret=dlb->id;+gotoalloc_id_fail;+}++ret=pcim_enable_device(pdev);+if(ret!=0){+dev_err(&pdev->dev,"failed to enable: %d\n",ret);++gotopci_enable_device_fail;+}++ret=pcim_iomap_regions(pdev,+(1U<<DLB_CSR_BAR)|(1U<<DLB_FUNC_BAR),+"dlb");+if(ret!=0){+dev_err(&pdev->dev,"failed to map: %d\n",ret);++gotopci_enable_device_fail;+}++pci_set_master(pdev);++ret=pci_enable_pcie_error_reporting(pdev);+if(ret!=0)+dev_info(&pdev->dev,"AER is not supported\n");++return0;++pci_enable_device_fail:+mutex_lock(&dlb_ids_lock);+idr_remove(&dlb_ids,dlb->id);+mutex_unlock(&dlb_ids_lock);+alloc_id_fail:+returnret;+}++staticvoiddlb_remove(structpci_dev*pdev)+{+structdlb*dlb=pci_get_drvdata(pdev);++pci_disable_pcie_error_reporting(pdev);++mutex_lock(&dlb_ids_lock);+idr_remove(&dlb_ids,dlb->id);+mutex_unlock(&dlb_ids_lock);+}++staticstructpci_device_iddlb_id_table[]={+{PCI_DEVICE_DATA(INTEL,DLB_PF,DLB_PF)},+{0}+};+MODULE_DEVICE_TABLE(pci,dlb_id_table);++staticstructpci_driverdlb_pci_driver={+.name="dlb",+.id_table=dlb_id_table,+.probe=dlb_probe,+.remove=dlb_remove,+};++staticint__initdlb_init_module(void)+{+interr;++dlb_class=class_create(THIS_MODULE,"dlb");++if(IS_ERR(dlb_class)){+pr_err("dlb: class_create() returned %ld\n",+PTR_ERR(dlb_class));++returnPTR_ERR(dlb_class);+}++err=alloc_chrdev_region(&dlb_devt,0,DLB_MAX_NUM_DEVICES,"dlb");++if(err<0){+pr_err("dlb: alloc_chrdev_region() returned %d\n",err);++gotoalloc_chrdev_fail;+}++err=pci_register_driver(&dlb_pci_driver);+if(err<0){+pr_err("dlb: pci_register_driver() returned %d\n",err);++gotopci_register_fail;+}++return0;++pci_register_fail:+unregister_chrdev_region(dlb_devt,DLB_MAX_NUM_DEVICES);+alloc_chrdev_fail:+class_destroy(dlb_class);++returnerr;+}++staticvoid__exitdlb_exit_module(void)+{+pci_unregister_driver(&dlb_pci_driver);++unregister_chrdev_region(dlb_devt,DLB_MAX_NUM_DEVICES);++class_destroy(dlb_class);+}++module_init(dlb_init_module);+module_exit(dlb_exit_module);
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:24
Add the hardware resource data structures, functions for
their initialization/teardown, and a function for device power-on. In
subsequent commits, dlb_resource.c will be expanded to hold the dlb
resource-management and configuration logic (using the data structures
defined in dlb_main.h).
There is a resource data structure for each system level: device/function,
scheduling domain, port and queue. At the device/function level, this
data structure is struct dlb_function_resources, which holds used and
avialable domains/ports/queues/history lists for the device function
(either physical or virtual).
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/Makefile | 2 +-
drivers/misc/dlb/dlb_main.c | 46 ++++++
drivers/misc/dlb/dlb_main.h | 249 ++++++++++++++++++++++++++++++++
drivers/misc/dlb/dlb_pf_ops.c | 56 +++++++
drivers/misc/dlb/dlb_regs.h | 119 +++++++++++++++
drivers/misc/dlb/dlb_resource.c | 210 +++++++++++++++++++++++++++
6 files changed, 681 insertions(+), 1 deletion(-)
create mode 100644 drivers/misc/dlb/dlb_regs.h
create mode 100644 drivers/misc/dlb/dlb_resource.c
@@ -11,11 +11,23 @@#include<linux/mutex.h>#include<linux/pci.h>#include<linux/types.h>+#include<linux/bitfield.h>/**Hardwarerelated#definesanddatastructures.**/+/* Read/write register 'reg' in the CSR BAR space */+#define DLB_CSR_REG_ADDR(a, reg) ((a)->csr_kva + (reg))+#define DLB_CSR_RD(hw, reg) ioread32(DLB_CSR_REG_ADDR((hw), (reg)))+#define DLB_CSR_WR(hw, reg, value) iowrite32((value), \+DLB_CSR_REG_ADDR((hw),(reg)))++/* Read/write register 'reg' in the func BAR space */+#define DLB_FUNC_REG_ADDR(a, reg) ((a)->func_kva + (reg))+#define DLB_FUNC_RD(hw, reg) ioread32(DLB_FUNC_REG_ADDR((hw), (reg)))+#define DLB_FUNC_WR(hw, reg, value) iowrite32((value), \+DLB_FUNC_REG_ADDR((hw),(reg)))#define DLB_MAX_NUM_VDEVS 16#define DLB_MAX_NUM_DOMAINS 32
@@ -42,6 +54,159 @@#define PCI_DEVICE_ID_INTEL_DLB_PF 0x2710+structdlb_ldb_queue{+structlist_headdomain_list;+structlist_headfunc_list;+u32id;+u32domain_id;+u32num_qid_inflights;+u32aqed_limit;+u32sn_group;/* sn == sequence number */+u32sn_slot;+u32num_mappings;+u8sn_cfg_valid;+u8num_pending_additions;+u8owned;+u8configured;+};++/*+*Directedportsandqueuesarepairedbynature,sothedrivertracksthem+*withasingledatastructure.+*/+structdlb_dir_pq_pair{+structlist_headdomain_list;+structlist_headfunc_list;+u32id;+u32domain_id;+u32ref_cnt;+u8init_tkn_cnt;+u8queue_configured;+u8port_configured;+u8owned;+u8enabled;+};++enumdlb_qid_map_state{+/* The slot doesn't contain a valid queue mapping */+DLB_QUEUE_UNMAPPED,+/* The slot contains a valid queue mapping */+DLB_QUEUE_MAPPED,+/* The driver is mapping a queue into this slot */+DLB_QUEUE_MAP_IN_PROG,+/* The driver is unmapping a queue from this slot */+DLB_QUEUE_UNMAP_IN_PROG,+/*+*Thedriverisunmappingaqueuefromthisslot,andoncecomplete+*willreplaceitwithanothermapping.+*/+DLB_QUEUE_UNMAP_IN_PROG_PENDING_MAP,+};++structdlb_ldb_port_qid_map{+enumdlb_qid_map_statestate;+u16qid;+u16pending_qid;+u8priority;+u8pending_priority;+};++structdlb_ldb_port{+structlist_headdomain_list;+structlist_headfunc_list;+u32id;+u32domain_id;+/* The qid_map represents the hardware QID mapping state. */+structdlb_ldb_port_qid_mapqid_map[DLB_MAX_NUM_QIDS_PER_LDB_CQ];+u32hist_list_entry_base;+u32hist_list_entry_limit;+u32ref_cnt;+u8init_tkn_cnt;+u8num_pending_removals;+u8num_mappings;+u8owned;+u8enabled;+u8configured;+};++structdlb_sn_group{+u32mode;+u32sequence_numbers_per_queue;+u32slot_use_bitmap;+u32id;+};++/*+*Schedulingdomainlevelresourcedatastructure.+*+*/+structdlb_hw_domain{+structdlb_function_resources*parent_func;+structlist_headfunc_list;+structlist_headused_ldb_queues;+structlist_headused_ldb_ports[DLB_NUM_COS_DOMAINS];+structlist_headused_dir_pq_pairs;+structlist_headavail_ldb_queues;+structlist_headavail_ldb_ports[DLB_NUM_COS_DOMAINS];+structlist_headavail_dir_pq_pairs;+u32total_hist_list_entries;+u32avail_hist_list_entries;+u32hist_list_entry_base;+u32hist_list_entry_offset;+u32num_ldb_credits;+u32num_dir_credits;+u32num_avail_aqed_entries;+u32num_used_aqed_entries;+u32id;+intnum_pending_removals;+intnum_pending_additions;+u8configured;+u8started;+};++/*+*Devicefunction(eitherPForVF)levelresourcedatastructure.+*+*/+structdlb_function_resources{+structlist_headavail_domains;+structlist_headused_domains;+structlist_headavail_ldb_queues;+structlist_headavail_ldb_ports[DLB_NUM_COS_DOMAINS];+structlist_headavail_dir_pq_pairs;+structdlb_bitmap*avail_hist_list_entries;+u32num_avail_domains;+u32num_avail_ldb_queues;+u32num_avail_ldb_ports[DLB_NUM_COS_DOMAINS];+u32num_avail_dir_pq_pairs;+u32num_avail_qed_entries;+u32num_avail_dqed_entries;+u32num_avail_aqed_entries;+u8locked;/* (VDEV only) */+};++/*+*Afterinitialization,eachresourceindlb_hw_resourcesislocatedinone+*ofthefollowinglists:+*--ThePF'savailableresourceslist.Theseareunconfiguredresourcesowned+*bythePFandnotallocatedtoadlbschedulingdomain.+*--AVDEV'savailableresourceslist.TheseareVDEV-ownedunconfigured+*resourcesnotallocatedtoadlbschedulingdomain.+*--Adomain'savailableresourceslist.Thesearedomain-ownedunconfigured+*resources.+*--Adomain'susedresourceslist.Thesearedomain-ownedconfigured+*resources.+*+*AresourcemovestoanewlistwhenaVDEVordomainiscreatedordestroyed,+*orwhentheresourceisconfigured.+*/+structdlb_hw_resources{+structdlb_ldb_queueldb_queues[DLB_MAX_NUM_LDB_QUEUES];+structdlb_ldb_portldb_ports[DLB_MAX_NUM_LDB_PORTS];+structdlb_dir_pq_pairdir_pq_pairs[DLB_MAX_NUM_DIR_PORTS];+structdlb_sn_groupsn_groups[DLB_MAX_NUM_SEQUENCE_NUMBER_GROUPS];+};+structdlb_hw{/* BAR 0 address */void__iomem*csr_kva;
@@ -1,7 +1,10 @@// SPDX-License-Identifier: GPL-2.0-only/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */+#include<linux/delay.h>+#include"dlb_main.h"+#include"dlb_regs.h"/********************************//****** PCI BAR management ******/
@@ -31,3 +34,56 @@ int dlb_pf_map_pci_bar_space(struct dlb *dlb, struct pci_dev *pdev)return0;}++/*******************************/+/****** Driver management ******/+/*******************************/++intdlb_pf_init_driver_state(structdlb*dlb)+{+mutex_init(&dlb->resource_mutex);++return0;+}++voiddlb_pf_enable_pm(structdlb*dlb)+{+/*+*Clearthepower-management-disableregistertopoweronthebulkof+*thedevice'shardware.+*/+dlb_clr_pmcsr_disable(&dlb->hw);+}++#define DLB_READY_RETRY_LIMIT 1000+intdlb_pf_wait_for_device_ready(structdlb*dlb,structpci_dev*pdev)+{+u32retries=DLB_READY_RETRY_LIMIT;++/* Allow at least 1s for the device to become active after power-on */+do{+u32idle,pm_st,addr;++addr=CM_CFG_PM_STATUS;++pm_st=DLB_CSR_RD(&dlb->hw,addr);++addr=CM_CFG_DIAGNOSTIC_IDLE_STATUS;++idle=DLB_CSR_RD(&dlb->hw,addr);++if(FIELD_GET(CM_CFG_PM_STATUS_PMSM,pm_st)==1&&+FIELD_GET(CM_CFG_DIAGNOSTIC_IDLE_STATUS_DLB_FUNC_IDLE,idle)+==1)+break;++usleep_range(1000,2000);+}while(--retries);++if(!retries){+dev_err(&pdev->dev,"Device idle test failed\n");+return-EIO;+}++return0;+}
@@ -0,0 +1,210 @@+// SPDX-License-Identifier: GPL-2.0-only+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */++#include"dlb_regs.h"+#include"dlb_main.h"++staticvoiddlb_init_fn_rsrc_lists(structdlb_function_resources*rsrc)+{+inti;++INIT_LIST_HEAD(&rsrc->avail_domains);+INIT_LIST_HEAD(&rsrc->used_domains);+INIT_LIST_HEAD(&rsrc->avail_ldb_queues);+INIT_LIST_HEAD(&rsrc->avail_dir_pq_pairs);++for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+INIT_LIST_HEAD(&rsrc->avail_ldb_ports[i]);+}++staticvoiddlb_init_domain_rsrc_lists(structdlb_hw_domain*domain)+{+inti;++INIT_LIST_HEAD(&domain->used_ldb_queues);+INIT_LIST_HEAD(&domain->used_dir_pq_pairs);+INIT_LIST_HEAD(&domain->avail_ldb_queues);+INIT_LIST_HEAD(&domain->avail_dir_pq_pairs);++for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+INIT_LIST_HEAD(&domain->used_ldb_ports[i]);+for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+INIT_LIST_HEAD(&domain->avail_ldb_ports[i]);+}++/**+*dlb_resource_free()-freedevicestatememory+*@hw:dlb_hwhandleforaparticulardevice.+*+*Thisfunctionfreessoftwarestatepointedtobydlb_hw.Thisfunction+*shouldbecalledwhenresettingthedeviceorunloadingthedriver.+*/+voiddlb_resource_free(structdlb_hw*hw)+{+inti;++if(hw->pf.avail_hist_list_entries)+dlb_bitmap_free(hw->pf.avail_hist_list_entries);++for(i=0;i<DLB_MAX_NUM_VDEVS;i++){+if(hw->vdev[i].avail_hist_list_entries)+dlb_bitmap_free(hw->vdev[i].avail_hist_list_entries);+}+}++/**+*dlb_resource_init()-initializethedevice+*@hw:pointertostructdlb_hw.+*+*Thisfunctioninitializesthedevice'ssoftwarestate(pointedtobythehw+*argument)andprogramsglobalschedulingQoSregisters.Thisfunctionshould+*becalledduringdriverinitialization,andthedlb_hwstructureshould+*bezero-initializedbeforecallingthefunction.+*+*Thedlb_hwstructmustbeuniqueperDLB2.0deviceandpersistuntilthe+*deviceisreset.+*+*Return:+*Returns0uponsuccess,<0otherwise.+*/+intdlb_resource_init(structdlb_hw*hw)+{+structdlb_bitmap*map;+structlist_head*list;+unsignedinti;+intret;++/*+*Foroptimalload-balancing,portsthatmaptooneormoreQIDsin+*commonshouldnotbeinnumericalsequence.Theport->QIDmappingis+*applicationdependent,butthedriverinterleavesportIDsasmuch+*aspossibletoreducethelikelihoodofsequentialportsmappingto+*thesameQID(s).ThisinitialallocationofportIDsmaximizesthe+*averagedistancebetweenanIDanditsimmediateneighbors(i.e.+*thedistancefrom1to0andto2,thedistancefrom2to1andto+*3,etc.).+*/+constu8init_ldb_port_allocation[DLB_MAX_NUM_LDB_PORTS]={+0,7,14,5,12,3,10,1,8,15,6,13,4,11,2,9,+16,23,30,21,28,19,26,17,24,31,22,29,20,27,18,25,+32,39,46,37,44,35,42,33,40,47,38,45,36,43,34,41,+48,55,62,53,60,51,58,49,56,63,54,61,52,59,50,57,+};++dlb_init_fn_rsrc_lists(&hw->pf);++for(i=0;i<DLB_MAX_NUM_VDEVS;i++)+dlb_init_fn_rsrc_lists(&hw->vdev[i]);++for(i=0;i<DLB_MAX_NUM_DOMAINS;i++){+dlb_init_domain_rsrc_lists(&hw->domains[i]);+hw->domains[i].parent_func=&hw->pf;+}++/* Give all resources to the PF driver */+hw->pf.num_avail_domains=DLB_MAX_NUM_DOMAINS;+for(i=0;i<hw->pf.num_avail_domains;i++){+list=&hw->domains[i].func_list;++list_add(list,&hw->pf.avail_domains);+}++hw->pf.num_avail_ldb_queues=DLB_MAX_NUM_LDB_QUEUES;+for(i=0;i<hw->pf.num_avail_ldb_queues;i++){+list=&hw->rsrcs.ldb_queues[i].func_list;++list_add(list,&hw->pf.avail_ldb_queues);+}++for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+hw->pf.num_avail_ldb_ports[i]=+DLB_MAX_NUM_LDB_PORTS/DLB_NUM_COS_DOMAINS;++for(i=0;i<DLB_MAX_NUM_LDB_PORTS;i++){+intcos_id=i>>DLB_NUM_COS_DOMAINS;+structdlb_ldb_port*port;++port=&hw->rsrcs.ldb_ports[init_ldb_port_allocation[i]];++list_add(&port->func_list,&hw->pf.avail_ldb_ports[cos_id]);+}++hw->pf.num_avail_dir_pq_pairs=DLB_MAX_NUM_DIR_PORTS;+for(i=0;i<hw->pf.num_avail_dir_pq_pairs;i++){+list=&hw->rsrcs.dir_pq_pairs[i].func_list;++list_add(list,&hw->pf.avail_dir_pq_pairs);+}++hw->pf.num_avail_qed_entries=DLB_MAX_NUM_LDB_CREDITS;+hw->pf.num_avail_dqed_entries=DLB_MAX_NUM_DIR_CREDITS;+hw->pf.num_avail_aqed_entries=DLB_MAX_NUM_AQED_ENTRIES;++ret=dlb_bitmap_alloc(&hw->pf.avail_hist_list_entries,+DLB_MAX_NUM_HIST_LIST_ENTRIES);+if(ret)+gotounwind;++map=hw->pf.avail_hist_list_entries;+bitmap_fill(map->map,map->len);++for(i=0;i<DLB_MAX_NUM_VDEVS;i++){+ret=dlb_bitmap_alloc(&hw->vdev[i].avail_hist_list_entries,+DLB_MAX_NUM_HIST_LIST_ENTRIES);+if(ret)+gotounwind;++map=hw->vdev[i].avail_hist_list_entries;+bitmap_zero(map->map,map->len);+}++/* Initialize the hardware resource IDs */+for(i=0;i<DLB_MAX_NUM_DOMAINS;i++)+hw->domains[i].id=i;++for(i=0;i<DLB_MAX_NUM_LDB_QUEUES;i++)+hw->rsrcs.ldb_queues[i].id=i;++for(i=0;i<DLB_MAX_NUM_LDB_PORTS;i++)+hw->rsrcs.ldb_ports[i].id=i;++for(i=0;i<DLB_MAX_NUM_DIR_PORTS;i++)+hw->rsrcs.dir_pq_pairs[i].id=i;++for(i=0;i<DLB_MAX_NUM_SEQUENCE_NUMBER_GROUPS;i++){+hw->rsrcs.sn_groups[i].id=i;+/* Default mode (0) is 64 sequence numbers per queue */+hw->rsrcs.sn_groups[i].mode=0;+hw->rsrcs.sn_groups[i].sequence_numbers_per_queue=64;+hw->rsrcs.sn_groups[i].slot_use_bitmap=0;+}++for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+hw->cos_reservation[i]=100/DLB_NUM_COS_DOMAINS;++return0;++unwind:+dlb_resource_free(hw);++returnret;+}++/**+*dlb_clr_pmcsr_disable()-poweronbulkofDLB2.0logic+*@hw:dlb_hwhandleforaparticulardevice.+*+*ClearingthePMCSRmustbedoneatinitializationtomakethedevicefully+*operational.+*/+voiddlb_clr_pmcsr_disable(structdlb_hw*hw)+{+u32pmcsr_dis;++pmcsr_dis=DLB_CSR_RD(hw,CM_CFG_PM_PMCSR_DISABLE);++/* Clear register bits */+pmcsr_dis&=~CM_CFG_PM_PMCSR_DISABLE_DISABLE;++DLB_CSR_WR(hw,CM_CFG_PM_PMCSR_DISABLE,pmcsr_dis);+}
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:27
Introduce the dlb device level configfs interface. Each dlb device has a
configfs node at /sys/kernel/config/dlb/dlb%N, which is used to control
and configure the device.
Also add the scheduling domain directory in the device node of configfs.
A scheduling domain serves as a HW container of DLB resources -- e.g.
ports, queues, and credits -- with the property that a port can only
enqueue to and dequeue from queues within its domain. A scheduling domain
is created on-demand by a user-space application, whose request includes
the DLB resource allocation.
To create a HW scheduling domain, a user creates a domain directory using
"mkdir" in the configfs, and configure the resources needed (such as number
of ports and queues, and credits, etc) by writing to the attribute files
in the directory. A final write of "1" to "create" file triggers creation
of a scheduling domain in DLB.
The hardware operation for scheduling domain creation will be added in a
subsequent commit.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/Makefile | 2 +-
drivers/misc/dlb/dlb_args.h | 60 +++++++
drivers/misc/dlb/dlb_configfs.c | 299 ++++++++++++++++++++++++++++++++
drivers/misc/dlb/dlb_configfs.h | 39 +++++
drivers/misc/dlb/dlb_main.c | 9 +
drivers/misc/dlb/dlb_main.h | 10 ++
drivers/misc/dlb/dlb_resource.c | 10 ++
7 files changed, 428 insertions(+), 1 deletion(-)
create mode 100644 drivers/misc/dlb/dlb_args.h
create mode 100644 drivers/misc/dlb/dlb_configfs.c
create mode 100644 drivers/misc/dlb/dlb_configfs.h
@@ -0,0 +1,299 @@+// SPDX-License-Identifier: GPL-2.0-only+// Copyright(c) 2017-2020 Intel Corporation++#include<linux/configfs.h>+#include"dlb_configfs.h"++structdlb_device_configfsdlb_dev_configfs[16];++staticintdlb_configfs_create_sched_domain(structdlb*dlb,+void*karg)+{+structdlb_create_sched_domain_args*arg=karg;+structdlb_cmd_responseresponse={0};+intret;++mutex_lock(&dlb->resource_mutex);++ret=dlb_hw_create_sched_domain(&dlb->hw,arg,&response);++mutex_unlock(&dlb->resource_mutex);++memcpy(karg,&response,sizeof(response));++returnret;+}++/*+*Configfsdirectorystructurefordlbdriverimplementation:+*+*config+*|+*dlb+*|+*+------+------+------+------+*||||+*dlb0dlb1dlb2dlb3...+*|+*+-----------+--+--------+-------+*|||+*domain0domain1domain2...+*|+*+-------+-----+------------+---------------+------------+----------+*|||||+*num_ldb_queuesport0port1...queue0queue1...+*num_ldb_ports||+*...is_ldbnum_sequence_numbers+*createcq_depthnum_qid_inflights+*start...num_atomic_iflights+*enable...+*...+*/++/*+*------Configfsfordlbdomains---------+*+*Thesearethetemplatesforshowandstorefunctionsindomain+*groups/directories,whichminimizesreplicationofboilerplate+*codetocopyarguments.Mostattributes,usethesimpletemplate.+*"name"istheattributenameinthegroup.+*/+#define DLB_CONFIGFS_DOMAIN_SHOW(name) \+staticssize_tdlb_cfs_domain_##name##_show(\+structconfig_item*item,\+char*page)\+{\+returnsprintf(page,"%u\n",\+to_dlb_cfs_domain(item)->name);\+}\++#define DLB_CONFIGFS_DOMAIN_STORE(name) \+staticssize_tdlb_cfs_domain_##name##_store(\+structconfig_item*item,\+constchar*page,\+size_tcount)\+{\+intret;\+structdlb_cfs_domain*dlb_cfs_domain=\+to_dlb_cfs_domain(item);\+\+ret=kstrtoint(page,10,&dlb_cfs_domain->name);\+if(ret)\+returnret;\+\+returncount;\+}\++DLB_CONFIGFS_DOMAIN_SHOW(status)+DLB_CONFIGFS_DOMAIN_SHOW(domain_id)+DLB_CONFIGFS_DOMAIN_SHOW(num_ldb_queues)+DLB_CONFIGFS_DOMAIN_SHOW(num_ldb_ports)+DLB_CONFIGFS_DOMAIN_SHOW(num_dir_ports)+DLB_CONFIGFS_DOMAIN_SHOW(num_atomic_inflights)+DLB_CONFIGFS_DOMAIN_SHOW(num_hist_list_entries)+DLB_CONFIGFS_DOMAIN_SHOW(num_ldb_credits)+DLB_CONFIGFS_DOMAIN_SHOW(num_dir_credits)+DLB_CONFIGFS_DOMAIN_SHOW(create)++DLB_CONFIGFS_DOMAIN_STORE(num_ldb_queues)+DLB_CONFIGFS_DOMAIN_STORE(num_ldb_ports)+DLB_CONFIGFS_DOMAIN_STORE(num_dir_ports)+DLB_CONFIGFS_DOMAIN_STORE(num_atomic_inflights)+DLB_CONFIGFS_DOMAIN_STORE(num_hist_list_entries)+DLB_CONFIGFS_DOMAIN_STORE(num_ldb_credits)+DLB_CONFIGFS_DOMAIN_STORE(num_dir_credits)++staticssize_tdlb_cfs_domain_create_store(structconfig_item*item,+constchar*page,size_tcount)+{+structdlb_cfs_domain*dlb_cfs_domain=to_dlb_cfs_domain(item);+structdlb_device_configfs*dlb_dev_configfs;+structdlb*dlb;+intret,create_in;++dlb_dev_configfs=container_of(dlb_cfs_domain->dev_grp,+structdlb_device_configfs,+dev_group);+dlb=dlb_dev_configfs->dlb;+if(!dlb)+return-EINVAL;++ret=kstrtoint(page,10,&create_in);+if(ret)+returnret;++/* Writing 1 to the 'create' triggers scheduling domain creation */+if(create_in==1&&dlb_cfs_domain->create==0){+structdlb_create_sched_domain_argsargs={0};++memcpy(&args.response,&dlb_cfs_domain->status,+sizeof(structdlb_create_sched_domain_args));++dev_dbg(dlb->dev,+"Create domain: %s\n",+dlb_cfs_domain->group.cg_item.ci_namebuf);++ret=dlb_configfs_create_sched_domain(dlb,&args);++dlb_cfs_domain->status=args.response.status;+dlb_cfs_domain->domain_id=args.response.id;++if(ret){+dev_err(dlb->dev,+"create sched domain failed: ret=%d\n",ret);+returnret;+}++dlb_cfs_domain->create=1;+}++returncount;+}++CONFIGFS_ATTR_RO(dlb_cfs_domain_,status);+CONFIGFS_ATTR_RO(dlb_cfs_domain_,domain_id);+CONFIGFS_ATTR(dlb_cfs_domain_,num_ldb_queues);+CONFIGFS_ATTR(dlb_cfs_domain_,num_ldb_ports);+CONFIGFS_ATTR(dlb_cfs_domain_,num_dir_ports);+CONFIGFS_ATTR(dlb_cfs_domain_,num_atomic_inflights);+CONFIGFS_ATTR(dlb_cfs_domain_,num_hist_list_entries);+CONFIGFS_ATTR(dlb_cfs_domain_,num_ldb_credits);+CONFIGFS_ATTR(dlb_cfs_domain_,num_dir_credits);+CONFIGFS_ATTR(dlb_cfs_domain_,create);++staticstructconfigfs_attribute*dlb_cfs_domain_attrs[]={+&dlb_cfs_domain_attr_status,+&dlb_cfs_domain_attr_domain_id,+&dlb_cfs_domain_attr_num_ldb_queues,+&dlb_cfs_domain_attr_num_ldb_ports,+&dlb_cfs_domain_attr_num_dir_ports,+&dlb_cfs_domain_attr_num_atomic_inflights,+&dlb_cfs_domain_attr_num_hist_list_entries,+&dlb_cfs_domain_attr_num_ldb_credits,+&dlb_cfs_domain_attr_num_dir_credits,+&dlb_cfs_domain_attr_create,++NULL,+};++staticvoiddlb_cfs_domain_release(structconfig_item*item)+{+kfree(to_dlb_cfs_domain(item));+}++staticstructconfigfs_item_operationsdlb_cfs_domain_item_ops={+.release=dlb_cfs_domain_release,+};++staticconststructconfig_item_typedlb_cfs_domain_type={+.ct_item_ops=&dlb_cfs_domain_item_ops,+.ct_attrs=dlb_cfs_domain_attrs,+.ct_owner=THIS_MODULE,+};++/*+*---------dlbdevicelevelconfigfs-----------+*+*Schedulingdomainsarecreatedinthedevice-levelconfigfsdriectory.+*/+staticstructconfig_group*dlb_cfs_device_make_domain(structconfig_group*group,+constchar*name)+{+structdlb_cfs_domain*dlb_cfs_domain;++dlb_cfs_domain=kzalloc(sizeof(*dlb_cfs_domain),GFP_KERNEL);+if(!dlb_cfs_domain)+returnERR_PTR(-ENOMEM);++dlb_cfs_domain->dev_grp=group;++config_group_init_type_name(&dlb_cfs_domain->group,name,+&dlb_cfs_domain_type);++return&dlb_cfs_domain->group;+}++staticstructconfigfs_group_operationsdlb_cfs_device_group_ops={+.make_group=dlb_cfs_device_make_domain,+};++staticconststructconfig_item_typedlb_cfs_device_type={+/* No need for _item_ops() at the device-level, and default+*attribute.+*.ct_item_ops=&dlb_cfs_device_item_ops,+*.ct_attrs=dlb_cfs_device_attrs,+*/++.ct_group_ops=&dlb_cfs_device_group_ops,+.ct_owner=THIS_MODULE,+};++/*------------------- dlb group subsystem for configfs ----------------+*+*weonlyneedasimpleconfigfsitemtypeherethatdoesnotlet+*usertocreatenewentry.Thegroupforeachdlbdevicewillbe+*generatedwhenthedeviceisdetectedindlb_probe().+*/++staticconststructconfig_item_typedlb_device_group_type={+.ct_owner=THIS_MODULE,+};++/* dlb group subsys in configfs */+staticstructconfigfs_subsystemdlb_device_group_subsys={+.su_group={+.cg_item={+.ci_namebuf="dlb",+.ci_type=&dlb_device_group_type,+},+},+};++/* Create a configfs directory dlbN for each dlb device probed+*indlb_probe()+*/+intdlb_configfs_create_device(structdlb*dlb)+{+structconfig_group*parent_group,*dev_grp;+chardevice_name[16];+intret=0;++snprintf(device_name,6,"dlb%d",dlb->id);+parent_group=&dlb_device_group_subsys.su_group;++dev_grp=&dlb_dev_configfs[dlb->id].dev_group;+config_group_init_type_name(dev_grp,+device_name,+&dlb_cfs_device_type);+ret=configfs_register_group(parent_group,dev_grp);++if(ret)+returnret;++dlb_dev_configfs[dlb->id].dlb=dlb;++returnret;+}++intconfigfs_dlb_init(void)+{+structconfigfs_subsystem*subsys;+intret;++/* set up and register configfs subsystem for dlb */+subsys=&dlb_device_group_subsys;+config_group_init(&subsys->su_group);+mutex_init(&subsys->su_mutex);+ret=configfs_register_subsystem(subsys);+if(ret){+pr_err("Error %d while registering subsystem %s\n",+ret,subsys->su_group.cg_item.ci_namebuf);+}++returnret;+}++voidconfigfs_dlb_exit(void)+{+configfs_unregister_subsystem(&dlb_device_group_subsys);+}
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:28
Add support for configuring a scheduling domain, creating the domain fd,
and reserving the domain's resources.
When a user requests to create a scheduling domain via configfs, the
requested resources are validated against the number currently available,
and then reserved for the scheduling domain. An anonymous file descriptor
for the domain is created and installed in the calling process's file
descriptor table.
The driver maintains a reference count for each scheduling domain,
incrementing it each time user-space requests a file descriptor for a dlb
port access and decrementing it in the file's release callback.
When the reference count transitions from 1->0 the driver automatically
resets the scheduling domain's resources and makes them available for use
by future applications. This ensures that applications that crash without
explicitly cleaning up do not orphan device resources. The code to perform
the domain reset will be added in subsequent commits.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_configfs.c | 46 ++-
drivers/misc/dlb/dlb_main.c | 68 ++++
drivers/misc/dlb/dlb_main.h | 128 ++++++++
drivers/misc/dlb/dlb_regs.h | 18 ++
drivers/misc/dlb/dlb_resource.c | 528 +++++++++++++++++++++++++++++++-
include/uapi/linux/dlb.h | 22 ++
6 files changed, 808 insertions(+), 2 deletions(-)
create mode 100644 include/uapi/linux/dlb.h
@@ -11,12 +14,38 @@ static int dlb_configfs_create_sched_domain(struct dlb *dlb,{structdlb_create_sched_domain_args*arg=karg;structdlb_cmd_responseresponse={0};-intret;+structdlb_domain*domain;+u32flags=O_RDONLY;+intret,fd;mutex_lock(&dlb->resource_mutex);ret=dlb_hw_create_sched_domain(&dlb->hw,arg,&response);+if(ret)+gotounlock;++ret=dlb_init_domain(dlb,response.id);+if(ret)+gotounlock;+domain=dlb->sched_domains[response.id];++if(dlb->f->f_mode&FMODE_WRITE)+flags=O_RDWR;++fd=anon_inode_getfd("[dlbdomain]",&dlb_domain_fops,+domain,flags);++if(fd<0){+dev_err(dlb->dev,"Failed to get anon fd.\n");+kref_put(&domain->refcnt,dlb_free_domain);+ret=fd;+gotounlock;+}++arg->domain_fd=fd;++unlock:mutex_unlock(&dlb->resource_mutex);memcpy(karg,&response,sizeof(response));
@@ -248,10 +249,20 @@ int dlb_pf_init_driver_state(struct dlb *dlb);voiddlb_pf_enable_pm(structdlb*dlb);intdlb_pf_wait_for_device_ready(structdlb*dlb,structpci_dev*pdev);+externconststructfile_operationsdlb_domain_fops;++structdlb_domain{+structdlb*dlb;+structkrefrefcnt;+u8id;+};+structdlb{structpci_dev*pdev;structdlb_hwhw;structdevice*dev;+structdlb_domain*sched_domains[DLB_MAX_NUM_DOMAINS];+structfile*f;/**Theresourcemutexserializesaccesstodriverdatastructuresand*hardwareregisters.
@@ -326,6 +337,123 @@ static inline void dlb_bitmap_free(struct dlb_bitmap *bitmap)kfree(bitmap);}+/**+*dlb_bitmap_clear_range()-cleararangeofbitmapentries+*@bitmap:pointertodlb_bitmapstructure.+*@bit:startingbitindex.+*@len:lengthoftherange.+*+*Return:+*Returns0uponsuccess,<0otherwise.+*+*Errors:+*EINVAL-bitmapisNULLorisuninitialized,ortherangeexceedsthebitmap+*length.+*/+staticinlineintdlb_bitmap_clear_range(structdlb_bitmap*bitmap,+unsignedintbit,+unsignedintlen)+{+if(!bitmap||!bitmap->map)+return-EINVAL;++if(bitmap->len<=bit)+return-EINVAL;++bitmap_clear(bitmap->map,bit,len);++return0;+}++/**+*dlb_bitmap_find_set_bit_range()-findanrangeofsetbits+*@bitmap:pointertodlb_bitmapstructure.+*@len:lengthoftherange.+*+*Thisfunctionlooksforarangeofsetbitsoflength@len.+*+*Return:+*Returnsthebasebitindexuponsuccess,<0otherwise.+*+*Errors:+*ENOENT-unabletofindalength*len*rangeofsetbits.+*EINVAL-bitmapisNULLorisuninitialized,orlenisinvalid.+*/+staticinlineintdlb_bitmap_find_set_bit_range(structdlb_bitmap*bitmap,+unsignedintlen)+{+structdlb_bitmap*complement_mask=NULL;+intret;++if(!bitmap||!bitmap->map||len==0)+return-EINVAL;++if(bitmap->len<len)+return-ENOENT;++ret=dlb_bitmap_alloc(&complement_mask,bitmap->len);+if(ret)+returnret;++bitmap_zero(complement_mask->map,complement_mask->len);++bitmap_complement(complement_mask->map,bitmap->map,bitmap->len);++ret=bitmap_find_next_zero_area(complement_mask->map,+complement_mask->len,+0,+len,+0);++dlb_bitmap_free(complement_mask);++/* No set bit range of length len? */+return(ret>=(int)bitmap->len)?-ENOENT:ret;+}++/**+*dlb_bitmap_longest_set_range()-returnslongestcontiguousrangeofset+*bits+*@bitmap:pointertodlb_bitmapstructure.+*+*Return:+*Returnsthebitmap'slongestcontiguousrangeofsetbitsuponsuccess,+*<0otherwise.+*+*Errors:+*EINVAL-bitmapisNULLorisuninitialized.+*/+staticinlineintdlb_bitmap_longest_set_range(structdlb_bitmap*bitmap)+{+intmax_len,len;+intstart,end;++if(!bitmap||!bitmap->map)+return-EINVAL;++if(bitmap_weight(bitmap->map,bitmap->len)==0)+return0;++max_len=0;+bitmap_for_each_set_region(bitmap->map,start,end,0,bitmap->len){+len=end-start;+if(max_len<len)+max_len=len;+}+returnmax_len;+}++intdlb_init_domain(structdlb*dlb,u32domain_id);+voiddlb_free_domain(structkref*kref);++staticinlinestructdevice*hw_to_dev(structdlb_hw*hw)+{+structdlb*dlb;++dlb=container_of(hw,structdlb,hw);+returndlb->dev;+}+/* Prototypes for dlb_resource.c */intdlb_resource_init(structdlb_hw*hw);voiddlb_resource_free(structdlb_hw*hw);
@@ -190,11 +190,537 @@ int dlb_resource_init(struct dlb_hw *hw)returnret;}+staticintdlb_attach_ldb_queues(structdlb_hw*hw,+structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,u32num_queues,+structdlb_cmd_response*resp)+{+unsignedinti;++if(rsrcs->num_avail_ldb_queues<num_queues){+resp->status=DLB_ST_LDB_QUEUES_UNAVAILABLE;+dev_dbg(hw_to_dev(hw),"[%s()] Internal error: %d\n",__func__,+resp->status);+return-EINVAL;+}++for(i=0;i<num_queues;i++){+structdlb_ldb_queue*queue;++queue=list_first_entry_or_null(&rsrcs->avail_ldb_queues,+typeof(*queue),func_list);+if(!queue){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: domain validation failed\n",+__func__);+return-EFAULT;+}++list_del(&queue->func_list);++queue->domain_id=domain->id;+queue->owned=true;++list_add(&queue->domain_list,&domain->avail_ldb_queues);+}++rsrcs->num_avail_ldb_queues-=num_queues;++return0;+}++staticstructdlb_ldb_port*+dlb_get_next_ldb_port(structdlb_hw*hw,structdlb_function_resources*rsrcs,+u32domain_id,u32cos_id)+{+structdlb_ldb_port*port;++/*+*Toreducetheoddsofconsecutiveload-balancedportsmappingtothe+*samequeue(s),thedriverattemptstoallocateportswhoseneighbors+*areownedbyadifferentdomain.+*/+list_for_each_entry(port,&rsrcs->avail_ldb_ports[cos_id],func_list){+u32next,prev;+u32phys_id;++phys_id=port->id;+next=phys_id+1;+prev=phys_id-1;++if(phys_id==DLB_MAX_NUM_LDB_PORTS-1)+next=0;+if(phys_id==0)+prev=DLB_MAX_NUM_LDB_PORTS-1;++if(!hw->rsrcs.ldb_ports[next].owned||+hw->rsrcs.ldb_ports[next].domain_id==domain_id)+continue;++if(!hw->rsrcs.ldb_ports[prev].owned||+hw->rsrcs.ldb_ports[prev].domain_id==domain_id)+continue;++returnport;+}++/*+*Failingthat,thedriverlooksforaportwithoneneighborownedby+*adifferentdomainandtheotherunallocated.+*/+list_for_each_entry(port,&rsrcs->avail_ldb_ports[cos_id],func_list){+u32next,prev;+u32phys_id;++phys_id=port->id;+next=phys_id+1;+prev=phys_id-1;++if(phys_id==DLB_MAX_NUM_LDB_PORTS-1)+next=0;+if(phys_id==0)+prev=DLB_MAX_NUM_LDB_PORTS-1;++if(!hw->rsrcs.ldb_ports[prev].owned&&+hw->rsrcs.ldb_ports[next].owned&&+hw->rsrcs.ldb_ports[next].domain_id!=domain_id)+returnport;++if(!hw->rsrcs.ldb_ports[next].owned&&+hw->rsrcs.ldb_ports[prev].owned&&+hw->rsrcs.ldb_ports[prev].domain_id!=domain_id)+returnport;+}++/*+*Failingthat,thedriverlooksforaportwithbothneighbors+*unallocated.+*/+list_for_each_entry(port,&rsrcs->avail_ldb_ports[cos_id],func_list){+u32next,prev;+u32phys_id;++phys_id=port->id;+next=phys_id+1;+prev=phys_id-1;++if(phys_id==DLB_MAX_NUM_LDB_PORTS-1)+next=0;+if(phys_id==0)+prev=DLB_MAX_NUM_LDB_PORTS-1;++if(!hw->rsrcs.ldb_ports[prev].owned&&+!hw->rsrcs.ldb_ports[next].owned)+returnport;+}++/* If all else fails, the driver returns the next available port. */+returnlist_first_entry_or_null(&rsrcs->avail_ldb_ports[cos_id],+typeof(*port),func_list);+}++staticint__dlb_attach_ldb_ports(structdlb_hw*hw,+structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,u32num_ports,+u32cos_id,structdlb_cmd_response*resp)+{+unsignedinti;++if(rsrcs->num_avail_ldb_ports[cos_id]<num_ports){+resp->status=DLB_ST_LDB_PORTS_UNAVAILABLE;+dev_dbg(hw_to_dev(hw),+"[%s()] Internal error: %d\n",+__func__,resp->status);+return-EINVAL;+}++for(i=0;i<num_ports;i++){+structdlb_ldb_port*port;++port=dlb_get_next_ldb_port(hw,rsrcs,+domain->id,cos_id);+if(!port){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: domain validation failed\n",+__func__);+return-EFAULT;+}++list_del(&port->func_list);++port->domain_id=domain->id;+port->owned=true;++list_add(&port->domain_list,+&domain->avail_ldb_ports[cos_id]);+}++rsrcs->num_avail_ldb_ports[cos_id]-=num_ports;++return0;+}++staticintdlb_attach_ldb_ports(structdlb_hw*hw,+structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,+structdlb_create_sched_domain_args*args,+structdlb_cmd_response*resp)+{+unsignedinti,j;+intret;++/* Allocate num_ldb_ports from any class-of-service */+for(i=0;i<args->num_ldb_ports;i++){+for(j=0;j<DLB_NUM_COS_DOMAINS;j++){+ret=__dlb_attach_ldb_ports(hw,rsrcs,domain,1,j,resp);+if(ret==0)+break;+}++if(ret)+returnret;+}++return0;+}++staticintdlb_attach_dir_ports(structdlb_hw*hw,+structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,u32num_ports,+structdlb_cmd_response*resp)+{+unsignedinti;++if(rsrcs->num_avail_dir_pq_pairs<num_ports){+resp->status=DLB_ST_DIR_PORTS_UNAVAILABLE;+dev_dbg(hw_to_dev(hw),+"[%s()] Internal error: %d\n",+__func__,resp->status);+return-EINVAL;+}++for(i=0;i<num_ports;i++){+structdlb_dir_pq_pair*port;++port=list_first_entry_or_null(&rsrcs->avail_dir_pq_pairs,+typeof(*port),func_list);+if(!port){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: domain validation failed\n",+__func__);+return-EFAULT;+}++list_del(&port->func_list);++port->domain_id=domain->id;+port->owned=true;++list_add(&port->domain_list,&domain->avail_dir_pq_pairs);+}++rsrcs->num_avail_dir_pq_pairs-=num_ports;++return0;+}++staticintdlb_attach_ldb_credits(structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,u32num_credits,+structdlb_cmd_response*resp)+{+if(rsrcs->num_avail_qed_entries<num_credits){+resp->status=DLB_ST_LDB_CREDITS_UNAVAILABLE;+return-EINVAL;+}++rsrcs->num_avail_qed_entries-=num_credits;+domain->num_ldb_credits+=num_credits;+return0;+}++staticintdlb_attach_dir_credits(structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,u32num_credits,+structdlb_cmd_response*resp)+{+if(rsrcs->num_avail_dqed_entries<num_credits){+resp->status=DLB_ST_DIR_CREDITS_UNAVAILABLE;+return-EINVAL;+}++rsrcs->num_avail_dqed_entries-=num_credits;+domain->num_dir_credits+=num_credits;+return0;+}++staticintdlb_attach_atomic_inflights(structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,+u32num_atomic_inflights,+structdlb_cmd_response*resp)+{+if(rsrcs->num_avail_aqed_entries<num_atomic_inflights){+resp->status=DLB_ST_ATOMIC_INFLIGHTS_UNAVAILABLE;+return-EINVAL;+}++rsrcs->num_avail_aqed_entries-=num_atomic_inflights;+domain->num_avail_aqed_entries+=num_atomic_inflights;+return0;+}++staticint+dlb_attach_domain_hist_list_entries(structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,+u32num_hist_list_entries,+structdlb_cmd_response*resp)+{+structdlb_bitmap*bitmap;+intbase;++if(num_hist_list_entries){+bitmap=rsrcs->avail_hist_list_entries;++base=dlb_bitmap_find_set_bit_range(bitmap,+num_hist_list_entries);+if(base<0)+gotoerror;++domain->total_hist_list_entries=num_hist_list_entries;+domain->avail_hist_list_entries=num_hist_list_entries;+domain->hist_list_entry_base=base;+domain->hist_list_entry_offset=0;++dlb_bitmap_clear_range(bitmap,base,num_hist_list_entries);+}+return0;++error:+resp->status=DLB_ST_HIST_LIST_ENTRIES_UNAVAILABLE;+return-EINVAL;+}++staticint+dlb_verify_create_sched_dom_args(structdlb_function_resources*rsrcs,+structdlb_create_sched_domain_args*args,+structdlb_cmd_response*resp,+structdlb_hw_domain**out_domain)+{+u32num_avail_ldb_ports,req_ldb_ports;+structdlb_bitmap*avail_hl_entries;+unsignedintmax_contig_hl_range;+structdlb_hw_domain*domain;+inti;++avail_hl_entries=rsrcs->avail_hist_list_entries;++max_contig_hl_range=dlb_bitmap_longest_set_range(avail_hl_entries);++num_avail_ldb_ports=0;+req_ldb_ports=0;+for(i=0;i<DLB_NUM_COS_DOMAINS;i++)+num_avail_ldb_ports+=rsrcs->num_avail_ldb_ports[i];++req_ldb_ports+=args->num_ldb_ports;++if(rsrcs->num_avail_domains<1){+resp->status=DLB_ST_DOMAIN_UNAVAILABLE;+return-EINVAL;+}++domain=list_first_entry_or_null(&rsrcs->avail_domains,+typeof(*domain),func_list);+if(!domain){+resp->status=DLB_ST_DOMAIN_UNAVAILABLE;+return-EFAULT;+}++if(rsrcs->num_avail_ldb_queues<args->num_ldb_queues){+resp->status=DLB_ST_LDB_QUEUES_UNAVAILABLE;+return-EINVAL;+}++if(req_ldb_ports>num_avail_ldb_ports){+resp->status=DLB_ST_LDB_PORTS_UNAVAILABLE;+return-EINVAL;+}++if(args->num_ldb_queues>0&&req_ldb_ports==0){+resp->status=DLB_ST_LDB_PORT_REQUIRED_FOR_LDB_QUEUES;+return-EINVAL;+}++if(rsrcs->num_avail_dir_pq_pairs<args->num_dir_ports){+resp->status=DLB_ST_DIR_PORTS_UNAVAILABLE;+return-EINVAL;+}++if(rsrcs->num_avail_qed_entries<args->num_ldb_credits){+resp->status=DLB_ST_LDB_CREDITS_UNAVAILABLE;+return-EINVAL;+}++if(rsrcs->num_avail_dqed_entries<args->num_dir_credits){+resp->status=DLB_ST_DIR_CREDITS_UNAVAILABLE;+return-EINVAL;+}++if(rsrcs->num_avail_aqed_entries<args->num_atomic_inflights){+resp->status=DLB_ST_ATOMIC_INFLIGHTS_UNAVAILABLE;+return-EINVAL;+}++if(max_contig_hl_range<args->num_hist_list_entries){+resp->status=DLB_ST_HIST_LIST_ENTRIES_UNAVAILABLE;+return-EINVAL;+}++*out_domain=domain;++return0;+}++staticvoiddlb_configure_domain_credits(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+u32reg;++reg=FIELD_PREP(CHP_CFG_LDB_VAS_CRD_COUNT,domain->num_ldb_credits);+DLB_CSR_WR(hw,CHP_CFG_LDB_VAS_CRD(domain->id),reg);++reg=FIELD_PREP(CHP_CFG_DIR_VAS_CRD_COUNT,domain->num_dir_credits);+DLB_CSR_WR(hw,CHP_CFG_DIR_VAS_CRD(domain->id),reg);+}++staticint+dlb_domain_attach_resources(structdlb_hw*hw,+structdlb_function_resources*rsrcs,+structdlb_hw_domain*domain,+structdlb_create_sched_domain_args*args,+structdlb_cmd_response*resp)+{+intret;++ret=dlb_attach_ldb_queues(hw,rsrcs,domain,args->num_ldb_queues,resp);+if(ret)+returnret;++ret=dlb_attach_ldb_ports(hw,rsrcs,domain,args,resp);+if(ret)+returnret;++ret=dlb_attach_dir_ports(hw,rsrcs,domain,args->num_dir_ports,resp);+if(ret)+returnret;++ret=dlb_attach_ldb_credits(rsrcs,domain,+args->num_ldb_credits,resp);+if(ret)+returnret;++ret=dlb_attach_dir_credits(rsrcs,domain,args->num_dir_credits,resp);+if(ret)+returnret;++ret=dlb_attach_domain_hist_list_entries(rsrcs,domain,+args->num_hist_list_entries,+resp);+if(ret)+returnret;++ret=dlb_attach_atomic_inflights(rsrcs,domain,+args->num_atomic_inflights,resp);+if(ret)+returnret;++dlb_configure_domain_credits(hw,domain);++domain->configured=true;++domain->started=false;++rsrcs->num_avail_domains--;++return0;+}++staticvoid+dlb_log_create_sched_domain_args(structdlb_hw*hw,+structdlb_create_sched_domain_args*args)+{+dev_dbg(hw_to_dev(hw),"DLB create sched domain arguments:\n");+dev_dbg(hw_to_dev(hw),"\tNumber of LDB queues: %d\n",+args->num_ldb_queues);+dev_dbg(hw_to_dev(hw),"\tNumber of LDB ports (any CoS): %d\n",+args->num_ldb_ports);+dev_dbg(hw_to_dev(hw),"\tNumber of DIR ports: %d\n",+args->num_dir_ports);+dev_dbg(hw_to_dev(hw),"\tNumber of ATM inflights: %d\n",+args->num_atomic_inflights);+dev_dbg(hw_to_dev(hw),"\tNumber of hist list entries: %d\n",+args->num_hist_list_entries);+dev_dbg(hw_to_dev(hw),"\tNumber of LDB credits: %d\n",+args->num_ldb_credits);+dev_dbg(hw_to_dev(hw),"\tNumber of DIR credits: %d\n",+args->num_dir_credits);+}++/**+*dlb_hw_create_sched_domain()-createaschedulingdomain+*@hw:dlb_hwhandleforaparticulardevice.+*@args:schedulingdomaincreationarguments.+*@resp:responsestructure.+*+*Thisfunctioncreatesaschedulingdomaincontainingtheresourcesspecified+*inargs.Theindividualresources(queues,ports,credits)canbeconfigured+*aftercreatingaschedulingdomain.+*+*Return:+*Returns0uponsuccess,<0otherwise.Ifanerroroccurs,resp->statusis+*assignedadetailederrorcodefromenumdlb_error.Ifsuccessful,resp->id+*containsthedomainID.+*+*Errors:+*EINVAL-Arequestedresourceisunavailable,ortherequesteddomainname+*isalreadyinuse.+*EFAULT-Internalerror(resp->statusnotset).+*/intdlb_hw_create_sched_domain(structdlb_hw*hw,structdlb_create_sched_domain_args*args,structdlb_cmd_response*resp){-resp->id=0;+structdlb_function_resources*rsrcs;+structdlb_hw_domain*domain;+intret;++rsrcs=&hw->pf;++dlb_log_create_sched_domain_args(hw,args);++/*+*Verifythathardwareresourcesareavailablebeforeattemptingto+*satisfytherequest.Thissimplifiestheerrorunwindingcode.+*/+ret=dlb_verify_create_sched_dom_args(rsrcs,args,resp,&domain);+if(ret)+returnret;++dlb_init_domain_rsrc_lists(domain);++ret=dlb_domain_attach_resources(hw,rsrcs,domain,args,resp);+if(ret){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to verify args.\n",+__func__);++returnret;+}++/*+*Configurationsucceeded,somovetheresourcefromthe'avail'to+*the'used'list(ifit'snotalreadythere).+*/+list_move(&domain->func_list,&rsrcs->used_domains);++resp->id=domain->id;resp->status=0;return0;
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:30
Implemented a power-management policy of putting the device in D0 when in
use (when there are any open device files or memory mappings, or there are
any virtual devices), and D3Hot otherwise.
Add resume/suspend callbacks; when the device resumes, reset the hardware
to a known good state.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_main.c | 98 ++++++++++++++++++++++++++++++++++-
drivers/misc/dlb/dlb_pf_ops.c | 8 +++
2 files changed, 105 insertions(+), 1 deletion(-)
@@ -254,6 +303,9 @@ static void dlb_remove(struct pci_dev *pdev){structdlb*dlb=pci_get_drvdata(pdev);+/* Undo the PM operation in dlb_probe(). */+pm_runtime_get_noresume(&pdev->dev);+dlb_resource_free(&dlb->hw);device_destroy(dlb_class,dlb->dev_number);
@@ -265,17 +317,61 @@ static void dlb_remove(struct pci_dev *pdev)mutex_unlock(&dlb_ids_lock);}+#ifdef CONFIG_PM+staticvoiddlb_reset_hardware_state(structdlb*dlb)+{+dlb_reset_device(dlb->pdev);+}++staticintdlb_runtime_suspend(structdevice*dev)+{+/* Return and let the PCI subsystem put the device in D3hot. */++return0;+}++staticintdlb_runtime_resume(structdevice*dev)+{+structpci_dev*pdev=container_of(dev,structpci_dev,dev);+structdlb*dlb=pci_get_drvdata(pdev);+intret;++/*+*ThePCIsubsystemputthedeviceinD0,butthedevicemaynothave+*completedpoweringup.Waituntilthedeviceisreadybefore+*proceeding.+*/+ret=dlb_pf_wait_for_device_ready(dlb,pdev);+if(ret)+returnret;++/* Now reinitialize the device state. */+dlb_reset_hardware_state(dlb);++return0;+}+#endif+staticstructpci_device_iddlb_id_table[]={{PCI_DEVICE_DATA(INTEL,DLB_PF,DLB_PF)},{0}};MODULE_DEVICE_TABLE(pci,dlb_id_table);+#ifdef CONFIG_PM+staticconststructdev_pm_opsdlb_pm_ops={+SET_RUNTIME_PM_OPS(dlb_runtime_suspend,dlb_runtime_resume,NULL)+};+#endif+staticstructpci_driverdlb_pci_driver={.name="dlb",.id_table=dlb_id_table,.probe=dlb_probe,.remove=dlb_remove,+#ifdef CONFIG_PM+.driver.pm=&dlb_pm_ops,+#endif};staticint__initdlb_init_module(void)
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:35
Program all registers used to configure the domain's resources back to
their reset values during scheduling domain reset. This ensures the device
is in a known good state if/when it is configured again in the future.
Additional work is required if a resource is in-use (e.g. a queue is
non-empty) at that time. Support for these cases will be added in
subsequent commits.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_regs.h | 3527 ++++++++++++++++++++++++++++++-
drivers/misc/dlb/dlb_resource.c | 387 ++++
2 files changed, 3902 insertions(+), 12 deletions(-)
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:37
Add operation to reset a domain's resource's software state when its
reference count reaches zero, and re-inserts those resources in their
respective available-resources linked lists, for use by future scheduling
domains.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_configfs.c | 10 +-
drivers/misc/dlb/dlb_main.c | 10 +-
drivers/misc/dlb/dlb_main.h | 36 ++++++
drivers/misc/dlb/dlb_resource.c | 204 ++++++++++++++++++++++++++++++++
include/uapi/linux/dlb.h | 1 +
5 files changed, 259 insertions(+), 2 deletions(-)
@@ -190,6 +190,14 @@ int dlb_resource_init(struct dlb_hw *hw)returnret;}+staticstructdlb_hw_domain*dlb_get_domain_from_id(structdlb_hw*hw,u32id)+{+if(id>=DLB_MAX_NUM_DOMAINS)+returnNULL;++return&hw->domains[id];+}+staticintdlb_attach_ldb_queues(structdlb_hw*hw,structdlb_function_resources*rsrcs,structdlb_hw_domain*domain,u32num_queues,
@@ -726,6 +734,202 @@ int dlb_hw_create_sched_domain(struct dlb_hw *hw,return0;}+/*+*dlb_domain_reset_software_state()-returnsdomain'sresources+*@hw:dlb_hwhandleforaparticulardevice.+*@domain:pointertoschedulingdomain.+*+*Thisfunctionreturnstheresourcesallocated/assignedtoadomainbackto+*thedevice/functionlevelresourcepool.Theseresourcesincludeldb/dir+*queues,ports,historylists,etc.Itiscalledbythedlb_reset_domain().+*Whenadomainiscreated/initialized,resourcesaremovedtoadomainfrom+*theresourcepool.+*+*/+staticintdlb_domain_reset_software_state(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb*dlb=container_of(hw,structdlb,hw);+structdlb_dir_pq_pair*tmp_dir_port;+structdlb_function_resources*rsrcs;+structdlb_ldb_queue*tmp_ldb_queue;+structdlb_ldb_port*tmp_ldb_port;+structdlb_dir_pq_pair*dir_port;+structdlb_ldb_queue*ldb_queue;+structdlb_ldb_port*ldb_port;+intret,i;++lockdep_assert_held(&dlb->resource_mutex);++rsrcs=domain->parent_func;++/* Move the domain's ldb queues to the function's avail list */+list_for_each_entry_safe(ldb_queue,tmp_ldb_queue,+&domain->used_ldb_queues,domain_list){+if(ldb_queue->sn_cfg_valid){+structdlb_sn_group*grp;++grp=&hw->rsrcs.sn_groups[ldb_queue->sn_group];++dlb_sn_group_free_slot(grp,ldb_queue->sn_slot);+ldb_queue->sn_cfg_valid=false;+}++ldb_queue->owned=false;+ldb_queue->num_mappings=0;+ldb_queue->num_pending_additions=0;++list_del(&ldb_queue->domain_list);+list_add(&ldb_queue->func_list,&rsrcs->avail_ldb_queues);+rsrcs->num_avail_ldb_queues++;+}++list_for_each_entry_safe(ldb_queue,tmp_ldb_queue,+&domain->avail_ldb_queues,domain_list){+ldb_queue->owned=false;++list_del(&ldb_queue->domain_list);+list_add(&ldb_queue->func_list,&rsrcs->avail_ldb_queues);+rsrcs->num_avail_ldb_queues++;+}++/* Move the domain's ldb ports to the function's avail list */+for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry_safe(ldb_port,tmp_ldb_port,+&domain->used_ldb_ports[i],domain_list){+intj;++ldb_port->owned=false;+ldb_port->configured=false;+ldb_port->num_pending_removals=0;+ldb_port->num_mappings=0;+ldb_port->init_tkn_cnt=0;+for(j=0;j<DLB_MAX_NUM_QIDS_PER_LDB_CQ;j++)+ldb_port->qid_map[j].state=+DLB_QUEUE_UNMAPPED;++list_del(&ldb_port->domain_list);+list_add(&ldb_port->func_list,+&rsrcs->avail_ldb_ports[i]);+rsrcs->num_avail_ldb_ports[i]++;+}++list_for_each_entry_safe(ldb_port,tmp_ldb_port,+&domain->avail_ldb_ports[i],domain_list){+ldb_port->owned=false;++list_del(&ldb_port->domain_list);+list_add(&ldb_port->func_list,+&rsrcs->avail_ldb_ports[i]);+rsrcs->num_avail_ldb_ports[i]++;+}+}++/* Move the domain's dir ports to the function's avail list */+list_for_each_entry_safe(dir_port,tmp_dir_port,+&domain->used_dir_pq_pairs,domain_list){+dir_port->owned=false;+dir_port->port_configured=false;+dir_port->init_tkn_cnt=0;++list_del(&dir_port->domain_list);++list_add(&dir_port->func_list,&rsrcs->avail_dir_pq_pairs);+rsrcs->num_avail_dir_pq_pairs++;+}++list_for_each_entry_safe(dir_port,tmp_dir_port,+&domain->avail_dir_pq_pairs,domain_list){+dir_port->owned=false;++list_del(&dir_port->domain_list);++list_add(&dir_port->func_list,&rsrcs->avail_dir_pq_pairs);+rsrcs->num_avail_dir_pq_pairs++;+}++/* Return hist list entries to the function */+ret=dlb_bitmap_set_range(rsrcs->avail_hist_list_entries,+domain->hist_list_entry_base,+domain->total_hist_list_entries);+if(ret){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: domain hist list base doesn't match the function's bitmap.\n",+__func__);+returnret;+}++domain->total_hist_list_entries=0;+domain->avail_hist_list_entries=0;+domain->hist_list_entry_base=0;+domain->hist_list_entry_offset=0;++rsrcs->num_avail_qed_entries+=domain->num_ldb_credits;+domain->num_ldb_credits=0;++rsrcs->num_avail_dqed_entries+=domain->num_dir_credits;+domain->num_dir_credits=0;++rsrcs->num_avail_aqed_entries+=domain->num_avail_aqed_entries;+rsrcs->num_avail_aqed_entries+=domain->num_used_aqed_entries;+domain->num_avail_aqed_entries=0;+domain->num_used_aqed_entries=0;++domain->num_pending_removals=0;+domain->num_pending_additions=0;+domain->configured=false;+domain->started=false;++/*+*Movethedomainoutoftheused_domainslistandbacktothe+*function'savail_domainslist.+*/+list_move(&domain->func_list,&rsrcs->avail_domains);+rsrcs->num_avail_domains++;++return0;+}++staticvoiddlb_log_reset_domain(structdlb_hw*hw,u32domain_id)+{+dev_dbg(hw_to_dev(hw),"DLB reset domain:\n");+dev_dbg(hw_to_dev(hw),"\tDomain ID: %d\n",domain_id);+}++/**+*dlb_reset_domain()-resetaschedulingdomain+*@hw:dlb_hwhandleforaparticulardevice.+*@domain_id:domainID.+*+*ThisfunctionresetsandfreesaDLB2.0schedulingdomainanditsassociated+*resources.+*+*Pre-condition:thedrivermustensuresoftwarehasstoppedsendingQEs+*throughthisdomain'sproducerportsbeforeinvokingthisfunction,or+*undefinedbehaviorwillresult.+*+*Return:+*Returns0uponsuccess,-1otherwise.+*+*EINVAL-InvaliddomainID,orthedomainisnotconfigured.+*EFAULT-Internalerror.(Possiblycausedifsoftwareisthepre-condition+*isnotmet.)+*ETIMEDOUT-Hardwarecomponentdidn'tresetintheexpectedtime.+*/+intdlb_reset_domain(structdlb_hw*hw,u32domain_id)+{+structdlb_hw_domain*domain;++dlb_log_reset_domain(hw,domain_id);++domain=dlb_get_domain_from_id(hw,domain_id);++if(!domain||!domain->configured)+return-EINVAL;++returndlb_domain_reset_software_state(hw,domain);+}+/***dlb_clr_pmcsr_disable()-poweronbulkofDLB2.0logic*@hw:dlb_hwhandleforaparticulardevice.
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:50:53
Add configfs interface to create DLB queues and query their depth, and the
corresponding scheduling domain reset code to drain the queues when they
are no longer in use.
When a CPU enqueues a queue entry (QE) to DLB, the QE entry is sent to
a DLB queue. These queues hold queue entries (QEs) that have not yet
been scheduled to a destination port. The queue's depth is the number of
QEs residing in a queue.
Each queue supports multiple priority levels, and while a directed queue
has a 1:1 mapping with a directed port, load-balanced queues can be
configured with a set of load-balanced ports that software desires the
queue's QEs to be scheduled to.
For ease of review, this commit is limited to higher-level code including
the configfs interface, request verification, and debug log messages. All
low level register access/configuration code will be included in a
subsequent commit.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_args.h | 110 +++++++
drivers/misc/dlb/dlb_configfs.c | 276 ++++++++++++++++
drivers/misc/dlb/dlb_configfs.h | 58 ++++
drivers/misc/dlb/dlb_main.h | 41 +++
drivers/misc/dlb/dlb_resource.c | 538 ++++++++++++++++++++++++++++++++
include/uapi/linux/dlb.h | 9 +
6 files changed, 1032 insertions(+)
@@ -586,6 +618,154 @@ dlb_verify_create_sched_dom_args(struct dlb_function_resources *rsrcs,return0;}+staticint+dlb_verify_create_ldb_queue_args(structdlb_hw*hw,u32domain_id,+structdlb_create_ldb_queue_args*args,+structdlb_cmd_response*resp,+structdlb_hw_domain**out_domain,+structdlb_ldb_queue**out_queue)+{+structdlb_hw_domain*domain;+structdlb_ldb_queue*queue;+inti;++domain=dlb_get_domain_from_id(hw,domain_id);++if(!domain){+resp->status=DLB_ST_INVALID_DOMAIN_ID;+return-EINVAL;+}++if(!domain->configured){+resp->status=DLB_ST_DOMAIN_NOT_CONFIGURED;+return-EINVAL;+}++if(domain->started){+resp->status=DLB_ST_DOMAIN_STARTED;+return-EINVAL;+}++queue=list_first_entry_or_null(&domain->avail_ldb_queues,+typeof(*queue),domain_list);+if(!queue){+resp->status=DLB_ST_LDB_QUEUES_UNAVAILABLE;+return-EINVAL;+}++if(args->num_sequence_numbers){+for(i=0;i<DLB_MAX_NUM_SEQUENCE_NUMBER_GROUPS;i++){+structdlb_sn_group*group=&hw->rsrcs.sn_groups[i];++if(group->sequence_numbers_per_queue==+args->num_sequence_numbers&&+!dlb_sn_group_full(group))+break;+}++if(i==DLB_MAX_NUM_SEQUENCE_NUMBER_GROUPS){+resp->status=DLB_ST_SEQUENCE_NUMBERS_UNAVAILABLE;+return-EINVAL;+}+}++if(args->num_qid_inflights>4096){+resp->status=DLB_ST_INVALID_QID_INFLIGHT_ALLOCATION;+return-EINVAL;+}++/* Inflights must be <= number of sequence numbers if ordered */+if(args->num_sequence_numbers!=0&&+args->num_qid_inflights>args->num_sequence_numbers){+resp->status=DLB_ST_INVALID_QID_INFLIGHT_ALLOCATION;+return-EINVAL;+}++if(domain->num_avail_aqed_entries<args->num_atomic_inflights){+resp->status=DLB_ST_ATOMIC_INFLIGHTS_UNAVAILABLE;+return-EINVAL;+}++if(args->num_atomic_inflights&&+args->lock_id_comp_level!=0&&+args->lock_id_comp_level!=64&&+args->lock_id_comp_level!=128&&+args->lock_id_comp_level!=256&&+args->lock_id_comp_level!=512&&+args->lock_id_comp_level!=1024&&+args->lock_id_comp_level!=2048&&+args->lock_id_comp_level!=4096&&+args->lock_id_comp_level!=65536){+resp->status=DLB_ST_INVALID_LOCK_ID_COMP_LEVEL;+return-EINVAL;+}++*out_domain=domain;+*out_queue=queue;++return0;+}++staticint+dlb_verify_create_dir_queue_args(structdlb_hw*hw,u32domain_id,+structdlb_create_dir_queue_args*args,+structdlb_cmd_response*resp,+structdlb_hw_domain**out_domain,+structdlb_dir_pq_pair**out_queue)+{+structdlb_hw_domain*domain;+structdlb_dir_pq_pair*pq;++domain=dlb_get_domain_from_id(hw,domain_id);++if(!domain){+resp->status=DLB_ST_INVALID_DOMAIN_ID;+return-EINVAL;+}++if(!domain->configured){+resp->status=DLB_ST_DOMAIN_NOT_CONFIGURED;+return-EINVAL;+}++if(domain->started){+resp->status=DLB_ST_DOMAIN_STARTED;+return-EINVAL;+}++/*+*Iftheuserclaimstheportisalreadyconfigured,validatetheport+*ID,itsdomain,andwhethertheportisconfigured.+*/+if(args->port_id!=-1){+pq=dlb_get_domain_used_dir_pq(args->port_id,+false,+domain);++if(!pq||pq->domain_id!=domain->id||+!pq->port_configured){+resp->status=DLB_ST_INVALID_PORT_ID;+return-EINVAL;+}+}else{+/*+*Ifthequeue'sportisnotconfigured,validatethatafree+*port-queuepairisavailable.+*/+pq=list_first_entry_or_null(&domain->avail_dir_pq_pairs,+typeof(*pq),domain_list);+if(!pq){+resp->status=DLB_ST_DIR_QUEUES_UNAVAILABLE;+return-EINVAL;+}+}++*out_domain=domain;+*out_queue=pq;++return0;+}+staticvoiddlb_configure_domain_credits(structdlb_hw*hw,structdlb_hw_domain*domain){
@@ -1311,6 +1839,16 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)if(!domain||!domain->configured)return-EINVAL;+/*+*Foreachqueueownedbythisdomain,disableitswritepermissionsto+*causeanytrafficsenttoittobedropped.Well-behavedsoftware+*shouldnotbesendingQEsatthispoint.+*/++ret=dlb_domain_verify_reset_success(hw,domain);+if(ret)+returnret;+/* Reset the QID and port state. */dlb_domain_reset_registers(hw,domain);
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:10
Add the low-level code for configuring a new queue and querying its depth.
When configuring a queue, program the device based on the user-supplied
queue configuration (from the corresponding configfs attributes).
Also add low-level code for resetting (draining) a non-empty queue during
scheduling domain reset. Draining a queue is an iterative process of
checking if the queue is empty, and if not then selecting a linked 'victim'
port and dequeueing the queue's events through this port. A port can only
receive a small number of events at a time, usually much fewer than the
queue depth, so draining a queue typically takes multiple iterations. This
process is finite since software cannot enqueue new events to the DLB's
(finite) on-device storage.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_main.h | 46 +++
drivers/misc/dlb/dlb_resource.c | 574 +++++++++++++++++++++++++++++++-
2 files changed, 618 insertions(+), 2 deletions(-)
@@ -1,9 +1,21 @@// SPDX-License-Identifier: GPL-2.0-only/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */+#include<linux/log2.h>#include"dlb_regs.h"#include"dlb_main.h"+/*+*ThePFdrivercannotassumethataregisterwritewillaffectsubsequentHCW+*writes.Toensureawritecompletes,thedrivermustreadbackaCSR.This+*functiononlyneedbecalledforconfigurationthatcanoccurafterthe+*domainhasstarted;priortostarting,applicationscan'tsendHCWs.+*/+staticinlinevoiddlb_flush_csr(structdlb_hw*hw)+{+DLB_CSR_RD(hw,SYS_TOTAL_VAS);+}+staticvoiddlb_init_fn_rsrc_lists(structdlb_function_resources*rsrc){inti;
@@ -766,6 +778,112 @@ dlb_verify_create_dir_queue_args(struct dlb_hw *hw, u32 domain_id,return0;}+staticvoiddlb_configure_ldb_queue(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_ldb_queue*queue,+structdlb_create_ldb_queue_args*args)+{+structdlb_sn_group*sn_group;+unsignedintoffs;+u32reg=0;+u32alimit;+u32level;++/* QID write permissions are turned on when the domain is started */+offs=domain->id*DLB_MAX_NUM_LDB_QUEUES+queue->id;++DLB_CSR_WR(hw,SYS_LDB_VASQID_V(offs),reg);++/*+*UnorderedQIDsget4Kinflights,orderedgetasmanyasthenumber+*ofsequencenumbers.+*/+reg=FIELD_PREP(LSP_QID_LDB_INFL_LIM_LIMIT,args->num_qid_inflights);+DLB_CSR_WR(hw,LSP_QID_LDB_INFL_LIM(queue->id),reg);++alimit=queue->aqed_limit;++if(alimit>DLB_MAX_NUM_AQED_ENTRIES)+alimit=DLB_MAX_NUM_AQED_ENTRIES;++reg=FIELD_PREP(LSP_QID_AQED_ACTIVE_LIM_LIMIT,alimit);+DLB_CSR_WR(hw,LSP_QID_AQED_ACTIVE_LIM(queue->id),reg);++level=args->lock_id_comp_level;+if(level>=64&&level<=4096&&is_power_of_2(level)){+reg&=~AQED_QID_HID_WIDTH_COMPRESS_CODE;+reg|=FIELD_PREP(AQED_QID_HID_WIDTH_COMPRESS_CODE,ilog2(level)-5);+}else{+reg=0;+}++DLB_CSR_WR(hw,AQED_QID_HID_WIDTH(queue->id),reg);++reg=0;+/* Don't timestamp QEs that pass through this queue */+DLB_CSR_WR(hw,SYS_LDB_QID_ITS(queue->id),reg);++reg=FIELD_PREP(LSP_QID_ATM_DEPTH_THRSH_THRESH,args->depth_threshold);+DLB_CSR_WR(hw,LSP_QID_ATM_DEPTH_THRSH(queue->id),reg);++reg=FIELD_PREP(LSP_QID_NALDB_DEPTH_THRSH_THRESH,args->depth_threshold);+DLB_CSR_WR(hw,LSP_QID_NALDB_DEPTH_THRSH(queue->id),reg);++/*+*Thisregisterlimitsthenumberofinflightflowsaqueuecanhave+*atonetime.Ithasanupperboundof2048,butcanbe+*over-subscribed.512ischosensothatasinglequeuedoesn'tuse+*theentireatomicstorage,butcanuseasubstantialportionif+*needed.+*/+reg=FIELD_PREP(AQED_QID_FID_LIM_QID_FID_LIMIT,512);+DLB_CSR_WR(hw,AQED_QID_FID_LIM(queue->id),reg);++/* Configure SNs */+sn_group=&hw->rsrcs.sn_groups[queue->sn_group];+reg=FIELD_PREP(CHP_ORD_QID_SN_MAP_MODE,sn_group->mode);+reg|=FIELD_PREP(CHP_ORD_QID_SN_MAP_SLOT,queue->sn_slot);+reg|=FIELD_PREP(CHP_ORD_QID_SN_MAP_GRP,sn_group->id);++DLB_CSR_WR(hw,CHP_ORD_QID_SN_MAP(queue->id),reg);++reg=FIELD_PREP(SYS_LDB_QID_CFG_V_SN_CFG_V,+(u32)(args->num_sequence_numbers!=0));+reg|=FIELD_PREP(SYS_LDB_QID_CFG_V_FID_CFG_V,+(u32)(args->num_atomic_inflights!=0));++DLB_CSR_WR(hw,SYS_LDB_QID_CFG_V(queue->id),reg);++reg=SYS_LDB_QID_V_QID_V;+DLB_CSR_WR(hw,SYS_LDB_QID_V(queue->id),reg);+}++staticvoiddlb_configure_dir_queue(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_dir_pq_pair*queue,+structdlb_create_dir_queue_args*args)+{+unsignedintoffs;+u32reg=0;++/* QID write permissions are turned on when the domain is started */+offs=domain->id*DLB_MAX_NUM_DIR_QUEUES++queue->id;++DLB_CSR_WR(hw,SYS_DIR_VASQID_V(offs),reg);++/* Don't timestamp QEs that pass through this queue */+DLB_CSR_WR(hw,SYS_DIR_QID_ITS(queue->id),reg);++reg=FIELD_PREP(LSP_QID_DIR_DEPTH_THRSH_THRESH,args->depth_threshold);+DLB_CSR_WR(hw,LSP_QID_DIR_DEPTH_THRSH(queue->id),reg);++reg=SYS_DIR_QID_V_QID_V;+DLB_CSR_WR(hw,SYS_DIR_QID_V(queue->id),reg);++queue->queue_configured=true;+}+staticvoiddlb_configure_domain_credits(structdlb_hw*hw,structdlb_hw_domain*domain){
@@ -1038,6 +1206,8 @@ int dlb_hw_create_ldb_queue(struct dlb_hw *hw, u32 domain_id,returnret;}+dlb_configure_ldb_queue(hw,domain,queue,args);+queue->num_mappings=0;queue->configured=true;
@@ -1101,6 +1271,8 @@ int dlb_hw_create_dir_queue(struct dlb_hw *hw, u32 domain_id,if(ret)returnret;+dlb_configure_dir_queue(hw,domain,queue,args);+/**Configurationsucceeded,somovetheresourcefromthe'avail'to*the'used'list(ifit'snotalreadythere).
@@ -1115,6 +1287,92 @@ int dlb_hw_create_dir_queue(struct dlb_hw *hw, u32 domain_id,return0;}+staticu32dlb_ldb_cq_inflight_count(structdlb_hw*hw,+structdlb_ldb_port*port)+{+u32cnt;++cnt=DLB_CSR_RD(hw,LSP_CQ_LDB_INFL_CNT(port->id));++returnFIELD_GET(LSP_CQ_LDB_INFL_CNT_COUNT,cnt);+}++staticu32dlb_ldb_cq_token_count(structdlb_hw*hw,structdlb_ldb_port*port)+{+u32cnt;++cnt=DLB_CSR_RD(hw,LSP_CQ_LDB_TKN_CNT(port->id));++/*+*Accountfortheinitialtokencount,whichisusedinorderto+*provideaCQwithdepthlessthan8.+*/++returnFIELD_GET(LSP_CQ_LDB_TKN_CNT_TOKEN_COUNT,cnt)-port->init_tkn_cnt;+}++staticvoid__iomem*dlb_producer_port_addr(structdlb_hw*hw,u8port_id,+boolis_ldb)+{+structdlb*dlb=container_of(hw,structdlb,hw);+uintptr_taddress=(uintptr_t)dlb->hw.func_kva;+unsignedlongsize;++if(is_ldb){+size=DLB_LDB_PP_STRIDE;+address+=DLB_DRV_LDB_PP_BASE+size*port_id;+}else{+size=DLB_DIR_PP_STRIDE;+address+=DLB_DRV_DIR_PP_BASE+size*port_id;+}++return(void__iomem*)address;+}++staticvoiddlb_drain_ldb_cq(structdlb_hw*hw,structdlb_ldb_port*port)+{+u32infl_cnt,tkn_cnt;+unsignedinti;++infl_cnt=dlb_ldb_cq_inflight_count(hw,port);+tkn_cnt=dlb_ldb_cq_token_count(hw,port);++if(infl_cnt||tkn_cnt){+structdlb_hcwhcw_mem[8],*hcw;+void__iomem*pp_addr;++pp_addr=dlb_producer_port_addr(hw,port->id,true);++/* Point hcw to a 64B-aligned location */+hcw=(structdlb_hcw*)((uintptr_t)&hcw_mem[4]&~0x3F);++/*+*ProgramthefirstHCWforacompletionandtokenreturnand+*theotherHCWsasNOOPS+*/++memset(hcw,0,4*sizeof(*hcw));+hcw->qe_comp=(infl_cnt>0);+hcw->cq_token=(tkn_cnt>0);+hcw->lock_id=tkn_cnt-1;++/*+*ToensureoutstandingHCWsreachthedevicebeforesubsequent+*deviceaccesses,fencethem.+*/+wmb();++/* Return tokens in the first HCW */+iosubmit_cmds512(pp_addr,hcw,1);++hcw->cq_token=0;++/* Issue remaining completions (if any) */+for(i=1;i<infl_cnt;i++)+iosubmit_cmds512(pp_addr,hcw,1);+}+}+/**dlb_domain_reset_software_state()-returnsdomain'sresources*@hw:dlb_hwhandleforaparticulardevice.
@@ -1271,6 +1529,21 @@ static int dlb_domain_reset_software_state(struct dlb_hw *hw,return0;}+staticu32dlb_dir_queue_depth(structdlb_hw*hw,structdlb_dir_pq_pair*queue)+{+u32cnt;++cnt=DLB_CSR_RD(hw,LSP_QID_DIR_ENQUEUE_CNT(queue->id));++returnFIELD_GET(LSP_QID_DIR_ENQUEUE_CNT_COUNT,cnt);+}++staticbooldlb_dir_queue_is_empty(structdlb_hw*hw,+structdlb_dir_pq_pair*queue)+{+returndlb_dir_queue_depth(hw,queue)==0;+}+staticvoiddlb_log_get_dir_queue_depth(structdlb_hw*hw,u32domain_id,u32queue_id){
@@ -1322,7 +1595,7 @@ int dlb_hw_get_dir_queue_depth(struct dlb_hw *hw, u32 domain_id,return-EINVAL;}-resp->id=0;+resp->id=dlb_dir_queue_depth(hw,queue);return0;}
@@ -1391,7 +1664,7 @@ int dlb_hw_get_ldb_queue_depth(struct dlb_hw *hw, u32 domain_id,return-EINVAL;}-resp->id=0;+resp->id=dlb_ldb_queue_depth(hw,queue);return0;}
@@ -1801,6 +2089,270 @@ static void dlb_domain_reset_registers(struct dlb_hw *hw,CHP_CFG_DIR_VAS_CRD_RST);}+staticvoiddlb_domain_drain_ldb_cqs(structdlb_hw*hw,+structdlb_hw_domain*domain,+booltoggle_port)+{+structdlb_ldb_port*port;+inti;++/* If the domain hasn't been started, there's no traffic to drain */+if(!domain->started)+return;++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(port,&domain->used_ldb_ports[i],domain_list){+if(toggle_port)+dlb_ldb_port_cq_disable(hw,port);++dlb_drain_ldb_cq(hw,port);++if(toggle_port)+dlb_ldb_port_cq_enable(hw,port);+}+}+}++staticbooldlb_domain_mapped_queues_empty(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_ldb_queue*queue;++list_for_each_entry(queue,&domain->used_ldb_queues,domain_list){+if(queue->num_mappings==0)+continue;++if(!dlb_ldb_queue_is_empty(hw,queue))+returnfalse;+}++returntrue;+}++staticintdlb_domain_drain_mapped_queues(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+inti;++/* If the domain hasn't been started, there's no traffic to drain */+if(!domain->started)+return0;++if(domain->num_pending_removals>0){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to unmap domain queues\n",+__func__);+return-EFAULT;+}++for(i=0;i<DLB_MAX_QID_EMPTY_CHECK_LOOPS;i++){+dlb_domain_drain_ldb_cqs(hw,domain,true);++if(dlb_domain_mapped_queues_empty(hw,domain))+break;+}++if(i==DLB_MAX_QID_EMPTY_CHECK_LOOPS){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to empty queues\n",+__func__);+return-EFAULT;+}++/*+*DraintheCQsonemoretime.Forthequeuestogoempty,theywould+*havescheduledoneormoreQEs.+*/+dlb_domain_drain_ldb_cqs(hw,domain,true);++return0;+}++staticintdlb_drain_dir_cq(structdlb_hw*hw,structdlb_dir_pq_pair*port)+{+unsignedintport_id=port->id;+u32cnt;++/* Return any outstanding tokens */+cnt=dlb_dir_cq_token_count(hw,port);++if(cnt!=0){+structdlb_hcwhcw_mem[8],*hcw;+void__iomem*pp_addr;++pp_addr=dlb_producer_port_addr(hw,port_id,false);++/* Point hcw to a 64B-aligned location */+hcw=(structdlb_hcw*)((uintptr_t)&hcw_mem[4]&~0x3F);++/*+*ProgramthefirstHCWforabatchtokenreturnand+*therestasNOOPS+*/+memset(hcw,0,4*sizeof(*hcw));+hcw->cq_token=1;+hcw->lock_id=cnt-1;++/*+*ToensureoutstandingHCWsreachthedevicebeforesubsequent+*deviceaccesses,fencethem.+*/+wmb();++iosubmit_cmds512(pp_addr,hcw,1);+}++return0;+}++staticintdlb_domain_drain_dir_cqs(structdlb_hw*hw,+structdlb_hw_domain*domain,+booltoggle_port)+{+structdlb_dir_pq_pair*port;+intret;++list_for_each_entry(port,&domain->used_dir_pq_pairs,domain_list){+/*+*Can'tdrainaportifit'snotconfigured,andthere's+*nothingtodrainifitsqueueisunconfigured.+*/+if(!port->port_configured||!port->queue_configured)+continue;++if(toggle_port)+dlb_dir_port_cq_disable(hw,port);++ret=dlb_drain_dir_cq(hw,port);+if(ret)+returnret;++if(toggle_port)+dlb_dir_port_cq_enable(hw,port);+}++return0;+}++staticbooldlb_domain_dir_queues_empty(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_dir_pq_pair*queue;++list_for_each_entry(queue,&domain->used_dir_pq_pairs,domain_list){+if(!dlb_dir_queue_is_empty(hw,queue))+returnfalse;+}++returntrue;+}++staticintdlb_domain_drain_dir_queues(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+inti,ret;++/* If the domain hasn't been started, there's no traffic to drain */+if(!domain->started)+return0;++for(i=0;i<DLB_MAX_QID_EMPTY_CHECK_LOOPS;i++){+ret=dlb_domain_drain_dir_cqs(hw,domain,true);+if(ret)+returnret;++if(dlb_domain_dir_queues_empty(hw,domain))+break;+}++if(i==DLB_MAX_QID_EMPTY_CHECK_LOOPS){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to empty queues\n",+__func__);+return-EFAULT;+}++/*+*DraintheCQsonemoretime.Forthequeuestogoempty,theywould+*havescheduledoneormoreQEs.+*/+ret=dlb_domain_drain_dir_cqs(hw,domain,true);+if(ret)+returnret;++return0;+}++staticvoid+dlb_domain_disable_ldb_queue_write_perms(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+intdomain_offset=domain->id*DLB_MAX_NUM_LDB_QUEUES;+structdlb_ldb_queue*queue;++list_for_each_entry(queue,&domain->used_ldb_queues,domain_list){+intidx=domain_offset+queue->id;++DLB_CSR_WR(hw,SYS_LDB_VASQID_V(idx),0);+}+}++staticvoid+dlb_domain_disable_dir_queue_write_perms(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+intdomain_offset=domain->id*DLB_MAX_NUM_DIR_PORTS;+structdlb_dir_pq_pair*queue;++list_for_each_entry(queue,&domain->used_dir_pq_pairs,domain_list){+intidx=domain_offset+queue->id;++DLB_CSR_WR(hw,SYS_DIR_VASQID_V(idx),0);+}+}++staticvoiddlb_domain_disable_dir_cqs(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_dir_pq_pair*port;++list_for_each_entry(port,&domain->used_dir_pq_pairs,domain_list){+port->enabled=false;++dlb_dir_port_cq_disable(hw,port);+}+}++staticvoiddlb_domain_disable_ldb_cqs(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_ldb_port*port;+inti;++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(port,&domain->used_ldb_ports[i],domain_list){+port->enabled=false;++dlb_ldb_port_cq_disable(hw,port);+}+}+}++staticvoiddlb_domain_enable_ldb_cqs(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_ldb_port*port;+inti;++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(port,&domain->used_ldb_ports[i],domain_list){+port->enabled=true;++dlb_ldb_port_cq_enable(hw,port);+}+}+}+staticvoiddlb_log_reset_domain(structdlb_hw*hw,u32domain_id){dev_dbg(hw_to_dev(hw),"DLB reset domain:\n");
@@ -1844,6 +2396,24 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)*causeanytrafficsenttoittobedropped.Well-behavedsoftware*shouldnotbesendingQEsatthispoint.*/+dlb_domain_disable_dir_queue_write_perms(hw,domain);++dlb_domain_disable_ldb_queue_write_perms(hw,domain);++/* Re-enable the CQs in order to drain the mapped queues. */+dlb_domain_enable_ldb_cqs(hw,domain);++ret=dlb_domain_drain_mapped_queues(hw,domain);+if(ret)+returnret;++/* Done draining LDB QEs, so disable the CQs. */+dlb_domain_disable_ldb_cqs(hw,domain);++dlb_domain_drain_dir_queues(hw,domain);++/* Done draining DIR QEs, so disable the CQs. */+dlb_domain_disable_dir_cqs(hw,domain);ret=dlb_domain_verify_reset_success(hw,domain);if(ret)
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:13
Add high-level code for port configuration and poll-mode query configfs
interface, argument verification, and placeholder functions for the
low-level register accesses. A subsequent commit will add the low-level
logic.
The port is a core's interface to the DLB, and it consists of an MMIO page
(the "producer port" (PP)) through which the core enqueues queue entries
and an in-memory queue (the "consumer queue" (CQ)) to which the device
schedules QEs. A subsequent commit will add the mmap interface for an
application to directly access the PP and CQ regions.
The driver allocates DMA memory for each port's CQ, and frees this memory
during domain reset or driver removal. Domain reset will also drains each
port's CQ and disables them from further scheduling.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_args.h | 59 +++++-
drivers/misc/dlb/dlb_configfs.c | 295 ++++++++++++++++++++++++++-
drivers/misc/dlb/dlb_configfs.h | 24 +++
drivers/misc/dlb/dlb_main.c | 51 +++++
drivers/misc/dlb/dlb_main.h | 22 +++
drivers/misc/dlb/dlb_pf_ops.c | 7 +
drivers/misc/dlb/dlb_resource.c | 341 ++++++++++++++++++++++++++++++++
include/uapi/linux/dlb.h | 6 +
8 files changed, 803 insertions(+), 2 deletions(-)
@@ -45,6 +45,109 @@ DLB_DOMAIN_CONFIGFS_CALLBACK_TEMPLATE(create_dir_queue)DLB_DOMAIN_CONFIGFS_CALLBACK_TEMPLATE(get_ldb_queue_depth)DLB_DOMAIN_CONFIGFS_CALLBACK_TEMPLATE(get_dir_queue_depth)+staticintdlb_domain_configfs_create_ldb_port(structdlb*dlb,+structdlb_domain*domain,+void*karg)+{+structdlb_cmd_responseresponse={0};+structdlb_create_ldb_port_args*arg=karg;+dma_addr_tcq_dma_base=0;+void*cq_base;+intret;++mutex_lock(&dlb->resource_mutex);++cq_base=dma_alloc_coherent(&dlb->pdev->dev,+DLB_CQ_SIZE,+&cq_dma_base,+GFP_KERNEL);+if(!cq_base){+response.status=DLB_ST_NO_MEMORY;+ret=-ENOMEM;+gotounlock;+}++ret=dlb_hw_create_ldb_port(&dlb->hw,domain->id,+arg,(uintptr_t)cq_dma_base,+&response);+if(ret)+gotounlock;++/* Fill out the per-port data structure */+dlb->ldb_port[response.id].id=response.id;+dlb->ldb_port[response.id].is_ldb=true;+dlb->ldb_port[response.id].domain=domain;+dlb->ldb_port[response.id].cq_base=cq_base;+dlb->ldb_port[response.id].cq_dma_base=cq_dma_base;+dlb->ldb_port[response.id].valid=true;++unlock:+if(ret&&cq_dma_base)+dma_free_coherent(&dlb->pdev->dev,+DLB_CQ_SIZE,+cq_base,+cq_dma_base);++mutex_unlock(&dlb->resource_mutex);++BUILD_BUG_ON(offsetof(typeof(*arg),response)!=0);++memcpy(karg,&response,sizeof(response));++returnret;+}++staticintdlb_domain_configfs_create_dir_port(structdlb*dlb,+structdlb_domain*domain,+void*karg)+{+structdlb_cmd_responseresponse={0};+structdlb_create_dir_port_args*arg=karg;+dma_addr_tcq_dma_base=0;+void*cq_base;+intret;++mutex_lock(&dlb->resource_mutex);++cq_base=dma_alloc_coherent(&dlb->pdev->dev,+DLB_CQ_SIZE,+&cq_dma_base,+GFP_KERNEL);+if(!cq_base){+response.status=DLB_ST_NO_MEMORY;+ret=-ENOMEM;+gotounlock;+}++ret=dlb_hw_create_dir_port(&dlb->hw,domain->id,+arg,(uintptr_t)cq_dma_base,+&response);+if(ret)+gotounlock;++/* Fill out the per-port data structure */+dlb->dir_port[response.id].id=response.id;+dlb->dir_port[response.id].is_ldb=false;+dlb->dir_port[response.id].domain=domain;+dlb->dir_port[response.id].cq_base=cq_base;+dlb->dir_port[response.id].cq_dma_base=cq_dma_base;+dlb->dir_port[response.id].valid=true;+unlock:+if(ret&&cq_dma_base)+dma_free_coherent(&dlb->pdev->dev,+DLB_CQ_SIZE,+cq_base,+cq_dma_base);++mutex_unlock(&dlb->resource_mutex);++BUILD_BUG_ON(offsetof(typeof(*arg),response)!=0);++memcpy(karg,&response,sizeof(response));++returnret;+}+staticintdlb_configfs_create_sched_domain(structdlb*dlb,void*karg){
@@ -125,12 +125,56 @@ int dlb_init_domain(struct dlb *dlb, u32 domain_id)return0;}+staticvoiddlb_release_port_memory(structdlb*dlb,+structdlb_port*port,+boolcheck_domain,+u32domain_id)+{+if(port->valid&&+(!check_domain||port->domain->id==domain_id))+dma_free_coherent(&dlb->pdev->dev,+DLB_CQ_SIZE,+port->cq_base,+port->cq_dma_base);++port->valid=false;+}++staticvoiddlb_release_domain_memory(structdlb*dlb,+boolcheck_domain,+u32domain_id)+{+structdlb_port*port;+inti;++for(i=0;i<DLB_MAX_NUM_LDB_PORTS;i++){+port=&dlb->ldb_port[i];++dlb_release_port_memory(dlb,port,check_domain,domain_id);+}++for(i=0;i<DLB_MAX_NUM_DIR_PORTS;i++){+port=&dlb->dir_port[i];++dlb_release_port_memory(dlb,port,check_domain,domain_id);+}+}++staticvoiddlb_release_device_memory(structdlb*dlb)+{+dlb_release_domain_memory(dlb,false,0);+}+staticint__dlb_free_domain(structdlb_domain*domain){structdlb*dlb=domain->dlb;intret;ret=dlb_reset_domain(&dlb->hw,domain->id);++/* Unpin and free all memory pages associated with the domain */+dlb_release_domain_memory(dlb,true,domain->id);+if(ret){dlb->domain_reset_failed=true;dev_err(dlb->dev,
@@ -321,6 +369,9 @@ static void dlb_remove(struct pci_dev *pdev)staticvoiddlb_reset_hardware_state(structdlb*dlb){dlb_reset_device(dlb->pdev);++/* Reinitialize any other hardware state */+dlb_pf_init_hardware(dlb);}staticintdlb_runtime_suspend(structdevice*dev)
@@ -884,6 +884,163 @@ static void dlb_configure_dir_queue(struct dlb_hw *hw,queue->queue_configured=true;}+staticbool+dlb_cq_depth_is_valid(u32depth)+{+/* Valid values for depth are+*1,2,4,8,16,32,64,128,256,512,and1024.+*/+if(!is_power_of_2(depth)||depth>1024)+returnfalse;++returntrue;+}++staticint+dlb_verify_create_ldb_port_args(structdlb_hw*hw,u32domain_id,+uintptr_tcq_dma_base,+structdlb_create_ldb_port_args*args,+structdlb_cmd_response*resp,+structdlb_hw_domain**out_domain,+structdlb_ldb_port**out_port,int*out_cos_id)+{+structdlb_ldb_port*port=NULL;+structdlb_hw_domain*domain;+inti,id;++domain=dlb_get_domain_from_id(hw,domain_id);++if(!domain){+resp->status=DLB_ST_INVALID_DOMAIN_ID;+return-EINVAL;+}++if(!domain->configured){+resp->status=DLB_ST_DOMAIN_NOT_CONFIGURED;+return-EINVAL;+}++if(domain->started){+resp->status=DLB_ST_DOMAIN_STARTED;+return-EINVAL;+}++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+id=i%DLB_NUM_COS_DOMAINS;++port=list_first_entry_or_null(&domain->avail_ldb_ports[id],+typeof(*port),domain_list);+if(port)+break;+}++if(!port){+resp->status=DLB_ST_LDB_PORTS_UNAVAILABLE;+return-EINVAL;+}++/* DLB requires 64B alignment */+if(!IS_ALIGNED(cq_dma_base,64)){+resp->status=DLB_ST_INVALID_CQ_VIRT_ADDR;+return-EINVAL;+}++if(!dlb_cq_depth_is_valid(args->cq_depth)){+resp->status=DLB_ST_INVALID_CQ_DEPTH;+return-EINVAL;+}++/* The history list size must be >= 1 */+if(!args->cq_history_list_size){+resp->status=DLB_ST_INVALID_HIST_LIST_DEPTH;+return-EINVAL;+}++if(args->cq_history_list_size>domain->avail_hist_list_entries){+resp->status=DLB_ST_HIST_LIST_ENTRIES_UNAVAILABLE;+return-EINVAL;+}++*out_domain=domain;+*out_port=port;+*out_cos_id=id;++return0;+}++staticint+dlb_verify_create_dir_port_args(structdlb_hw*hw,u32domain_id,+uintptr_tcq_dma_base,+structdlb_create_dir_port_args*args,+structdlb_cmd_response*resp,+structdlb_hw_domain**out_domain,+structdlb_dir_pq_pair**out_port)+{+structdlb_hw_domain*domain;+structdlb_dir_pq_pair*pq;++domain=dlb_get_domain_from_id(hw,domain_id);++if(!domain){+resp->status=DLB_ST_INVALID_DOMAIN_ID;+return-EINVAL;+}++if(!domain->configured){+resp->status=DLB_ST_DOMAIN_NOT_CONFIGURED;+return-EINVAL;+}++if(domain->started){+resp->status=DLB_ST_DOMAIN_STARTED;+return-EINVAL;+}++if(args->queue_id!=-1){+/*+*Iftheuserclaimsthequeueisalreadyconfigured,validate+*thequeueID,itsdomain,andwhetherthequeueis+*configured.+*/+pq=dlb_get_domain_used_dir_pq(args->queue_id,+false,+domain);++if(!pq||pq->domain_id!=domain->id||+!pq->queue_configured){+resp->status=DLB_ST_INVALID_DIR_QUEUE_ID;+return-EINVAL;+}+}else{+/*+*Iftheport'squeueisnotconfigured,validatethatafree+*port-queuepairisavailable.+*/+pq=list_first_entry_or_null(&domain->avail_dir_pq_pairs,+typeof(*pq),domain_list);+if(!pq){+resp->status=DLB_ST_DIR_PORTS_UNAVAILABLE;+return-EINVAL;+}+}++/* DLB requires 64B alignment */+if(!IS_ALIGNED(cq_dma_base,64)){+resp->status=DLB_ST_INVALID_CQ_VIRT_ADDR;+return-EINVAL;+}++if(!dlb_cq_depth_is_valid(args->cq_depth)){+resp->status=DLB_ST_INVALID_CQ_DEPTH;+return-EINVAL;+}++*out_domain=domain;+*out_port=pq;++return0;+}+staticvoiddlb_configure_domain_credits(structdlb_hw*hw,structdlb_hw_domain*domain){
@@ -1287,6 +1444,146 @@ int dlb_hw_create_dir_queue(struct dlb_hw *hw, u32 domain_id,return0;}+staticvoid+dlb_log_create_ldb_port_args(structdlb_hw*hw,u32domain_id,+uintptr_tcq_dma_base,+structdlb_create_ldb_port_args*args)+{+dev_dbg(hw_to_dev(hw),"DLB create load-balanced port arguments:\n");+dev_dbg(hw_to_dev(hw),"\tDomain ID: %d\n",+domain_id);+dev_dbg(hw_to_dev(hw),"\tCQ depth: %d\n",+args->cq_depth);+dev_dbg(hw_to_dev(hw),"\tCQ hist list size: %d\n",+args->cq_history_list_size);+dev_dbg(hw_to_dev(hw),"\tCQ base address: 0x%lx\n",+cq_dma_base);+}++/**+*dlb_hw_create_ldb_port()-createaload-balancedport+*@hw:dlb_hwhandleforaparticulardevice.+*@domain_id:domainID.+*@args:portcreationarguments.+*@cq_dma_base:baseaddressoftheCQmemory.ThiscanbeaPAoranIOVA.+*@resp:responsestructure.+*+*Thisfunctioncreatesaload-balancedport.+*+*Return:+*Returns0uponsuccess,<0otherwise.Ifanerroroccurs,resp->statusis+*assignedadetailederrorcodefromenumdlb_error.Ifsuccessful,resp->id+*containstheportID.+*+*Errors:+*EINVAL-Arequestedresourceisunavailable,acreditsettingisinvalid,a+*pointeraddressisnotproperlyaligned,thedomainisnot+*configured,orthedomainhasalreadybeenstarted.+*EFAULT-Internalerror(resp->statusnotset).+*/+intdlb_hw_create_ldb_port(structdlb_hw*hw,u32domain_id,+structdlb_create_ldb_port_args*args,+uintptr_tcq_dma_base,+structdlb_cmd_response*resp)+{+structdlb_hw_domain*domain;+structdlb_ldb_port*port;+intret,cos_id;++dlb_log_create_ldb_port_args(hw,domain_id,cq_dma_base,+args);++/*+*Verifythathardwareresourcesareavailablebeforeattemptingto+*satisfytherequest.Thissimplifiestheerrorunwindingcode.+*/+ret=dlb_verify_create_ldb_port_args(hw,domain_id,cq_dma_base,args,+resp,&domain,+&port,&cos_id);+if(ret)+returnret;++/*+*Configurationsucceeded,somovetheresourcefromthe'avail'to+*the'used'list.+*/+list_move(&port->domain_list,&domain->used_ldb_ports[cos_id]);++resp->status=0;+resp->id=port->id;++return0;+}++staticvoid+dlb_log_create_dir_port_args(structdlb_hw*hw,+u32domain_id,uintptr_tcq_dma_base,+structdlb_create_dir_port_args*args)+{+dev_dbg(hw_to_dev(hw),"DLB create directed port arguments:\n");+dev_dbg(hw_to_dev(hw),"\tDomain ID: %d\n",+domain_id);+dev_dbg(hw_to_dev(hw),"\tCQ depth: %d\n",+args->cq_depth);+dev_dbg(hw_to_dev(hw),"\tCQ base address: 0x%lx\n",+cq_dma_base);+}++/**+*dlb_hw_create_dir_port()-createadirectedport+*@hw:dlb_hwhandleforaparticulardevice.+*@domain_id:domainID.+*@args:portcreationarguments.+*@cq_dma_base:baseaddressoftheCQmemory.ThiscanbeaPAoranIOVA.+*@resp:responsestructure.+*+*Thisfunctioncreatesadirectedport.+*+*Return:+*Returns0uponsuccess,<0otherwise.Ifanerroroccurs,resp->statusis+*assignedadetailederrorcodefromenumdlb_error.Ifsuccessful,resp->id+*containstheportID.+*+*Errors:+*EINVAL-Arequestedresourceisunavailable,acreditsettingisinvalid,a+*pointeraddressisnotproperlyaligned,thedomainisnot+*configured,orthedomainhasalreadybeenstarted.+*EFAULT-Internalerror(resp->statusnotset).+*/+intdlb_hw_create_dir_port(structdlb_hw*hw,u32domain_id,+structdlb_create_dir_port_args*args,+uintptr_tcq_dma_base,+structdlb_cmd_response*resp)+{+structdlb_dir_pq_pair*port;+structdlb_hw_domain*domain;+intret;++dlb_log_create_dir_port_args(hw,domain_id,cq_dma_base,args);++/*+*Verifythathardwareresourcesareavailablebeforeattemptingto+*satisfytherequest.Thissimplifiestheerrorunwindingcode.+*/+ret=dlb_verify_create_dir_port_args(hw,domain_id,cq_dma_base,+args,resp,+&domain,&port);+if(ret)+returnret;++/*+*Configurationsucceeded,somovetheresourcefromthe'avail'to+*the'used'list(ifit'snotalreadythere).+*/+if(args->queue_id==-1)+list_move(&port->domain_list,&domain->used_dir_pq_pairs);++resp->status=0;+resp->id=port->id;++return0;+}+staticu32dlb_ldb_cq_inflight_count(structdlb_hw*hw,structdlb_ldb_port*port){
@@ -2400,6 +2697,15 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)dlb_domain_disable_ldb_queue_write_perms(hw,domain);+/*+*DisabletheLDBCQsanddraintheminordertocompletethemapand+*unmapprocedures,whichrequirezeroCQinflightsandzeroQID+*inflightsrespectively.+*/+dlb_domain_disable_ldb_cqs(hw,domain);++dlb_domain_drain_ldb_cqs(hw,domain,false);+/* Re-enable the CQs in order to drain the mapped queues. */dlb_domain_enable_ldb_cqs(hw,domain);
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:27
Add the low-level code for configuring a new port, programming the
device-wide poll mode setting, and resetting a port.
The low-level port configuration functions program the device based on the
user-supplied parameters (from configfs attributes). These parameter are
first verified, e.g. to ensure that the port's CQ base address is properly
cache-line aligned.
During domain reset, each port is drained until its inflight count and
owed-token count reaches 0, reflecting an empty CQ. Once the ports are
drained, the domain reset operation disables them from being candidates
for future scheduling decisions -- until they are re-assigned to a new
scheduling domain in the future and re-enabled.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_resource.c | 418 +++++++++++++++++++++++++++++++-
1 file changed, 417 insertions(+), 1 deletion(-)
@@ -1217,6 +1217,296 @@ static void dlb_dir_port_cq_disable(struct dlb_hw *hw,dlb_flush_csr(hw);}+staticvoiddlb_ldb_port_configure_pp(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_ldb_port*port)+{+u32reg;++reg=FIELD_PREP(SYS_LDB_PP2VAS_VAS,domain->id);+DLB_CSR_WR(hw,SYS_LDB_PP2VAS(port->id),reg);++reg=0;+reg|=SYS_LDB_PP_V_PP_V;+DLB_CSR_WR(hw,SYS_LDB_PP_V(port->id),reg);+}++staticintdlb_ldb_port_configure_cq(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_ldb_port*port,+uintptr_tcq_dma_base,+structdlb_create_ldb_port_args*args)+{+u32hl_base=0;+u32reg=0;+u32ds,n;++/* The CQ address is 64B-aligned, and the DLB only wants bits [63:6] */+reg=FIELD_PREP(SYS_LDB_CQ_ADDR_L_ADDR_L,cq_dma_base>>6);+DLB_CSR_WR(hw,SYS_LDB_CQ_ADDR_L(port->id),reg);++reg=cq_dma_base>>32;+DLB_CSR_WR(hw,SYS_LDB_CQ_ADDR_U(port->id),reg);++/*+*'ro'==relaxedordering.ThissettingallowsDLBtowrite+*cachelinesout-of-order(butQEswithinacachelinearealways+*updatedin-order).+*/+reg=FIELD_PREP(SYS_LDB_CQ2VF_PF_RO_IS_PF,1);+reg|=SYS_LDB_CQ2VF_PF_RO_RO;++DLB_CSR_WR(hw,SYS_LDB_CQ2VF_PF_RO(port->id),reg);++if(!dlb_cq_depth_is_valid(args->cq_depth)){+dev_err(hw_to_dev(hw),+"[%s():%d] Internal error: invalid CQ depth\n",+__func__,__LINE__);+return-EINVAL;+}++if(args->cq_depth<=8){+ds=1;+}else{+n=ilog2(args->cq_depth);+ds=(n-2)&0x0f;+}++reg=FIELD_PREP(CHP_LDB_CQ_TKN_DEPTH_SEL_TOKEN_DEPTH_SELECT,ds);+DLB_CSR_WR(hw,CHP_LDB_CQ_TKN_DEPTH_SEL(port->id),reg);++/*+*TosupportCQswithdepthlessthan8,programthetokencount+*registerwithanon-zeroinitialvalue.Operationssuchasdomain+*resetmusttakethisinitialvalueintoaccountwhenquiescingthe+*CQ.+*/+port->init_tkn_cnt=0;++if(args->cq_depth<8){+port->init_tkn_cnt=8-args->cq_depth;++reg=FIELD_PREP(LSP_CQ_LDB_TKN_CNT_TOKEN_COUNT,port->init_tkn_cnt);+DLB_CSR_WR(hw,LSP_CQ_LDB_TKN_CNT(port->id),reg);+}else{+DLB_CSR_WR(hw,+LSP_CQ_LDB_TKN_CNT(port->id),+LSP_CQ_LDB_TKN_CNT_RST);+}++reg=FIELD_PREP(LSP_CQ_LDB_TKN_DEPTH_SEL_TOKEN_DEPTH_SELECT,ds);+DLB_CSR_WR(hw,LSP_CQ_LDB_TKN_DEPTH_SEL(port->id),reg);++/* Reset the CQ write pointer */+DLB_CSR_WR(hw,+CHP_LDB_CQ_WPTR(port->id),+CHP_LDB_CQ_WPTR_RST);++reg=FIELD_PREP(CHP_HIST_LIST_LIM_LIMIT,port->hist_list_entry_limit-1);+DLB_CSR_WR(hw,CHP_HIST_LIST_LIM(port->id),reg);++hl_base=FIELD_PREP(CHP_HIST_LIST_BASE_BASE,port->hist_list_entry_base);+DLB_CSR_WR(hw,CHP_HIST_LIST_BASE(port->id),hl_base);++/*+*TheinflightlimitsetsacaponthenumberofQEsforwhichthisCQ+*canowecompletionsatonetime.+*/+reg=FIELD_PREP(LSP_CQ_LDB_INFL_LIM_LIMIT,args->cq_history_list_size);+DLB_CSR_WR(hw,LSP_CQ_LDB_INFL_LIM(port->id),reg);++reg=FIELD_PREP(CHP_HIST_LIST_PUSH_PTR_PUSH_PTR,+FIELD_GET(CHP_HIST_LIST_BASE_BASE,hl_base));+DLB_CSR_WR(hw,CHP_HIST_LIST_PUSH_PTR(port->id),reg);++reg=FIELD_PREP(CHP_HIST_LIST_POP_PTR_POP_PTR,+FIELD_GET(CHP_HIST_LIST_BASE_BASE,hl_base));+DLB_CSR_WR(hw,CHP_HIST_LIST_POP_PTR(port->id),reg);++/*+*Addresstranslation(AT)settings:0:untranslated,2:translated+*(seeATSspecregardingAddressTypefieldformoredetails)+*/++reg=0;+DLB_CSR_WR(hw,SYS_LDB_CQ_AT(port->id),reg);+DLB_CSR_WR(hw,SYS_LDB_CQ_PASID(port->id),reg);++reg=FIELD_PREP(CHP_LDB_CQ2VAS_CQ2VAS,domain->id);+DLB_CSR_WR(hw,CHP_LDB_CQ2VAS(port->id),reg);++/* Disable the port's QID mappings */+reg=0;+DLB_CSR_WR(hw,LSP_CQ2PRIOV(port->id),reg);++return0;+}++staticintdlb_configure_ldb_port(structdlb_hw*hw,structdlb_hw_domain*domain,+structdlb_ldb_port*port,+uintptr_tcq_dma_base,+structdlb_create_ldb_port_args*args)+{+intret,i;++port->hist_list_entry_base=domain->hist_list_entry_base++domain->hist_list_entry_offset;+port->hist_list_entry_limit=port->hist_list_entry_base++args->cq_history_list_size;++domain->hist_list_entry_offset+=args->cq_history_list_size;+domain->avail_hist_list_entries-=args->cq_history_list_size;++ret=dlb_ldb_port_configure_cq(hw,+domain,+port,+cq_dma_base,+args);+if(ret)+returnret;++dlb_ldb_port_configure_pp(hw,domain,port);++dlb_ldb_port_cq_enable(hw,port);++for(i=0;i<DLB_MAX_NUM_QIDS_PER_LDB_CQ;i++)+port->qid_map[i].state=DLB_QUEUE_UNMAPPED;+port->num_mappings=0;++port->enabled=true;++port->configured=true;++return0;+}++staticvoiddlb_dir_port_configure_pp(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_dir_pq_pair*port)+{+u32reg;++reg=FIELD_PREP(SYS_DIR_PP2VAS_VAS,domain->id);+DLB_CSR_WR(hw,SYS_DIR_PP2VAS(port->id),reg);++reg=0;+reg|=SYS_DIR_PP_V_PP_V;+DLB_CSR_WR(hw,SYS_DIR_PP_V(port->id),reg);+}++staticintdlb_dir_port_configure_cq(structdlb_hw*hw,+structdlb_hw_domain*domain,+structdlb_dir_pq_pair*port,+uintptr_tcq_dma_base,+structdlb_create_dir_port_args*args)+{+u32reg;+u32ds,n;++/* The CQ address is 64B-aligned, and the DLB only wants bits [63:6] */+reg=FIELD_PREP(SYS_DIR_CQ_ADDR_L_ADDR_L,cq_dma_base>>6);+DLB_CSR_WR(hw,SYS_DIR_CQ_ADDR_L(port->id),reg);++reg=cq_dma_base>>32;+DLB_CSR_WR(hw,SYS_DIR_CQ_ADDR_U(port->id),reg);++/*+*'ro'==relaxedordering.ThissettingallowsDLBtowrite+*cachelinesout-of-order(butQEswithinacachelinearealways+*updatedin-order).+*/+reg=FIELD_PREP(SYS_DIR_CQ2VF_PF_RO_IS_PF,1);+reg|=SYS_DIR_CQ2VF_PF_RO_RO;++DLB_CSR_WR(hw,SYS_DIR_CQ2VF_PF_RO(port->id),reg);++if(!dlb_cq_depth_is_valid(args->cq_depth)){+dev_err(hw_to_dev(hw),+"[%s():%d] Internal error: invalid CQ depth\n",+__func__,__LINE__);+return-EINVAL;+}++if(args->cq_depth<=8){+ds=1;+}else{+n=ilog2(args->cq_depth);+ds=(n-2)&0x0f;+}++reg=FIELD_PREP(CHP_DIR_CQ_TKN_DEPTH_SEL_TOKEN_DEPTH_SELECT,ds);+DLB_CSR_WR(hw,CHP_DIR_CQ_TKN_DEPTH_SEL(port->id),reg);++/*+*TosupportCQswithdepthlessthan8,programthetokencount+*registerwithanon-zeroinitialvalue.Operationssuchasdomain+*resetmusttakethisinitialvalueintoaccountwhenquiescingthe+*CQ.+*/+port->init_tkn_cnt=0;++if(args->cq_depth<8){+port->init_tkn_cnt=8-args->cq_depth;++reg=FIELD_PREP(LSP_CQ_DIR_TKN_CNT_COUNT,port->init_tkn_cnt);+DLB_CSR_WR(hw,LSP_CQ_DIR_TKN_CNT(port->id),reg);+}else{+DLB_CSR_WR(hw,+LSP_CQ_DIR_TKN_CNT(port->id),+LSP_CQ_DIR_TKN_CNT_RST);+}++reg=FIELD_PREP(LSP_CQ_DIR_TKN_DEPTH_SEL_DSI_TOKEN_DEPTH_SELECT,ds);+DLB_CSR_WR(hw,LSP_CQ_DIR_TKN_DEPTH_SEL_DSI(port->id),reg);++/* Reset the CQ write pointer */+DLB_CSR_WR(hw,+CHP_DIR_CQ_WPTR(port->id),+CHP_DIR_CQ_WPTR_RST);++/* Virtualize the PPID */+reg=0;+DLB_CSR_WR(hw,SYS_DIR_CQ_FMT(port->id),reg);++/*+*Addresstranslation(AT)settings:0:untranslated,2:translated+*(seeATSspecregardingAddressTypefieldformoredetails)+*/+reg=0;+DLB_CSR_WR(hw,SYS_DIR_CQ_AT(port->id),reg);++DLB_CSR_WR(hw,SYS_DIR_CQ_PASID(port->id),reg);++reg=FIELD_PREP(CHP_DIR_CQ2VAS_CQ2VAS,domain->id);+DLB_CSR_WR(hw,CHP_DIR_CQ2VAS(port->id),reg);++return0;+}++staticintdlb_configure_dir_port(structdlb_hw*hw,structdlb_hw_domain*domain,+structdlb_dir_pq_pair*port,+uintptr_tcq_dma_base,+structdlb_create_dir_port_args*args)+{+intret;++ret=dlb_dir_port_configure_cq(hw,domain,port,cq_dma_base,+args);++if(ret)+returnret;++dlb_dir_port_configure_pp(hw,domain,port);++dlb_dir_port_cq_enable(hw,port);++port->enabled=true;++port->port_configured=true;++return0;+}+staticvoiddlb_log_create_sched_domain_args(structdlb_hw*hw,structdlb_create_sched_domain_args*args)
@@ -1503,6 +1793,11 @@ int dlb_hw_create_ldb_port(struct dlb_hw *hw, u32 domain_id,if(ret)returnret;+ret=dlb_configure_ldb_port(hw,domain,port,cq_dma_base,+args);+if(ret)+returnret;+/**Configurationsucceeded,somovetheresourcefromthe'avail'to*the'used'list.
@@ -1571,6 +1866,11 @@ int dlb_hw_create_dir_port(struct dlb_hw *hw, u32 domain_id,if(ret)returnret;+ret=dlb_configure_dir_port(hw,domain,port,cq_dma_base,+args);+if(ret)+returnret;+/**Configurationsucceeded,somovetheresourcefromthe'avail'to*the'used'list(ifit'snotalreadythere).
@@ -2363,6 +2693,35 @@ static int dlb_domain_verify_reset_success(struct dlb_hw *hw,}}+/* Confirm that all the domain's CQs inflight and token counts are 0. */+for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(ldb_port,&domain->used_ldb_ports[i],domain_list){+if(dlb_ldb_cq_inflight_count(hw,ldb_port)||+dlb_ldb_cq_token_count(hw,ldb_port)){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to empty ldb port %d\n",+__func__,ldb_port->id);+return-EFAULT;+}+}+}++list_for_each_entry(dir_port,&domain->used_dir_pq_pairs,domain_list){+if(!dlb_dir_queue_is_empty(hw,dir_port)){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to empty dir queue %d\n",+__func__,dir_port->id);+return-EFAULT;+}++if(dlb_dir_cq_token_count(hw,dir_port)){+dev_err(hw_to_dev(hw),+"[%s()] Internal error: failed to empty dir port %d\n",+__func__,dir_port->id);+return-EFAULT;+}+}+return0;}
@@ -2580,6 +2939,51 @@ static int dlb_domain_drain_dir_queues(struct dlb_hw *hw,return0;}+staticvoid+dlb_domain_disable_dir_producer_ports(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_dir_pq_pair*port;+u32pp_v=0;++list_for_each_entry(port,&domain->used_dir_pq_pairs,domain_list){+DLB_CSR_WR(hw,SYS_DIR_PP_V(port->id),pp_v);+}+}++staticvoid+dlb_domain_disable_ldb_producer_ports(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_ldb_port*port;+u32pp_v=0;+inti;++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(port,&domain->used_ldb_ports[i],domain_list){+DLB_CSR_WR(hw,+SYS_LDB_PP_V(port->id),+pp_v);+}+}+}++staticvoiddlb_domain_disable_ldb_seq_checks(structdlb_hw*hw,+structdlb_hw_domain*domain)+{+structdlb_ldb_port*port;+u32chk_en=0;+inti;++for(i=0;i<DLB_NUM_COS_DOMAINS;i++){+list_for_each_entry(port,&domain->used_ldb_ports[i],domain_list){+DLB_CSR_WR(hw,+CHP_SN_CHK_ENBL(port->id),+chk_en);+}+}+}+staticvoiddlb_domain_disable_ldb_queue_write_perms(structdlb_hw*hw,structdlb_hw_domain*domain)
@@ -2697,6 +3101,9 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)dlb_domain_disable_ldb_queue_write_perms(hw,domain);+/* Turn off completion tracking on all the domain's PPs. */+dlb_domain_disable_ldb_seq_checks(hw,domain);+/**DisabletheLDBCQsanddraintheminordertocompletethemapand*unmapprocedures,whichrequirezeroCQinflightsandzeroQID
@@ -2706,6 +3113,10 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)dlb_domain_drain_ldb_cqs(hw,domain,false);+ret=dlb_domain_wait_for_ldb_cqs_to_empty(hw,domain);+if(ret)+returnret;+/* Re-enable the CQs in order to drain the mapped queues. */dlb_domain_enable_ldb_cqs(hw,domain);
@@ -2721,6 +3132,11 @@ int dlb_reset_domain(struct dlb_hw *hw, u32 domain_id)/* Done draining DIR QEs, so disable the CQs. */dlb_domain_disable_dir_cqs(hw,domain);+/* Disable PPs */+dlb_domain_disable_dir_producer_ports(hw,domain);++dlb_domain_disable_ldb_producer_ports(hw,domain);+ret=dlb_domain_verify_reset_success(hw,domain);if(ret)returnret;
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:35
Once a port is created, the application can mmap the corresponding DMA
memory and MMIO into user-space. This allows user-space applications to
do (performance-sensitive) enqueue and dequeue independent of the kernel
driver.
The mmap callback is only available through special port files: a producer
port (PP) file and a consumer queue (CQ) file. User-space gets an fd for
these files by calling a new ioctl, DLB_DOMAIN_CMD_GET_{LDB,
DIR}_PORT_{PP, CQ}_FD, and passing in a port ID. If the ioctl succeeds, the
returned fd can be used to mmap that port's PP/CQ.
Device reset requires first unmapping all user-space mappings, to prevent
applications from interfering with the reset operation. To this end, the
driver uses a single inode -- allocated when the first PP/CQ file is
created, and freed when the last such file is closed -- and attaches all
port files to this common inode, as done elsewhere in Linux (e.g. cxl,
dax).
Allocating this inode requires creating a pseudo-filesystem. The driver
initializes this FS when the inode is allocated, and frees the FS after the
inode is freed.
The driver doesn't use anon_inode_getfd() for these port mmap files because
the anon inode layer uses a single inode that is shared with other kernel
components -- calling unmap_mapping_range() on that shared inode would
likely break the kernel.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/Makefile | 1 +
drivers/misc/dlb/dlb_args.h | 31 ++++++
drivers/misc/dlb/dlb_configfs.c | 162 ++++++++++++++++++++++++++++++++
drivers/misc/dlb/dlb_configfs.h | 70 ++++++++++++++
drivers/misc/dlb/dlb_file.c | 149 +++++++++++++++++++++++++++++
drivers/misc/dlb/dlb_main.c | 120 +++++++++++++++++++++++
drivers/misc/dlb/dlb_main.h | 24 +++++
drivers/misc/dlb/dlb_resource.c | 109 +++++++++++++++++++++
8 files changed, 666 insertions(+)
create mode 100644 drivers/misc/dlb/dlb_file.c
@@ -0,0 +1,149 @@+// SPDX-License-Identifier: GPL-2.0-only+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */++#include<linux/anon_inodes.h>+#include<linux/file.h>+#include<linux/module.h>+#include<linux/mount.h>+#include<linux/pseudo_fs.h>++#include"dlb_main.h"++/*+*dlbtracksitsmemorymappingssoitcanrevokethemwhenanFLRis+*requestedanduser-spacecannotbeallowedtoaccessthedevice.Toachieve+*that,thedrivercreatesasingleinodethroughwhichalldriver-created+*filescanshareastructaddress_space,andunmapstheinode'saddressspace+*duringtheresetpreparationphase.Sincetheanoninodelayersharesits+*inodewithmultiplekernelcomponents,wecannotusethathere.+*+*Doingsorequiresacustompseudo-filesystemtoallocatetheinode.TheFS+*andtheinodeareallocatedondemandwhenafileiscreated,andbothare+*freedwhenthelastsuchfileisclosed.+*+*Thisisinspiredbyotherdrivers(cxl,dax,mem)andtheanoninodelayer.+*/+staticintdlb_fs_cnt;+staticstructvfsmount*dlb_vfs_mount;++#define DLBFS_MAGIC 0x444C4232 /* ASCII for DLB */+staticintdlb_init_fs_context(structfs_context*fc)+{+returninit_pseudo(fc,DLBFS_MAGIC)?0:-ENOMEM;+}++staticstructfile_system_typedlb_fs_type={+.name="dlb",+.owner=THIS_MODULE,+.init_fs_context=dlb_init_fs_context,+.kill_sb=kill_anon_super,+};++/* Allocate an anonymous inode. Must hold the resource mutex while calling. */+staticstructinode*dlb_alloc_inode(structdlb*dlb)+{+structinode*inode;+intret;++/* Increment the pseudo-FS's refcnt and (if not already) mount it. */+ret=simple_pin_fs(&dlb_fs_type,&dlb_vfs_mount,&dlb_fs_cnt);+if(ret<0){+dev_err(dlb->dev,+"[%s()] Cannot mount pseudo filesystem: %d\n",+__func__,ret);+returnERR_PTR(ret);+}++dlb->inode_cnt++;++if(dlb->inode_cnt>1){+/*+*Returnthepreviouslyallocatedinode.Inthiscase,there+*isguaranteed>=1referenceandsoihold()issafetocall.+*/+ihold(dlb->inode);+returndlb->inode;+}++inode=alloc_anon_inode(dlb_vfs_mount->mnt_sb);+if(IS_ERR(inode)){+dev_err(dlb->dev,+"[%s()] Cannot allocate inode: %ld\n",+__func__,PTR_ERR(inode));+dlb->inode_cnt=0;+simple_release_fs(&dlb_vfs_mount,&dlb_fs_cnt);+}++dlb->inode=inode;++returninode;+}++/*+*DecrementtheinodereferencecountandreleasetheFS.Intendedfor+*unwindingdlb_alloc_inode().Mustholdtheresourcemutexwhilecalling.+*/+staticvoiddlb_free_inode(structinode*inode)+{+iput(inode);+simple_release_fs(&dlb_vfs_mount,&dlb_fs_cnt);+}++/*+*ReleasetheFS.Intendedforuseinafile_operationsreleasecallback,+*whichdecrementstheinodereferencecountseparately.Mustholdthe+*resourcemutexwhilecalling.+*/+voiddlb_release_fs(structdlb*dlb)+{+mutex_lock(&dlb_driver_mutex);++simple_release_fs(&dlb_vfs_mount,&dlb_fs_cnt);++dlb->inode_cnt--;++/* When the fs refcnt reaches zero, the inode has been freed */+if(dlb->inode_cnt==0)+dlb->inode=NULL;++mutex_unlock(&dlb_driver_mutex);+}++/*+*Allocateafilewiththerequestedflags,fileoperations,andnamethat+*usesthedevice'ssharedinode.Mustholdtheresourcemutexwhilecalling.+*+*Callermustseparatelyallocateanfdandinstallthefileinthatfd.+*/+structfile*dlb_getfile(structdlb*dlb,+intflags,+conststructfile_operations*fops,+constchar*name)+{+structinode*inode;+structfile*f;++if(!try_module_get(THIS_MODULE))+returnERR_PTR(-ENOENT);++mutex_lock(&dlb_driver_mutex);++inode=dlb_alloc_inode(dlb);+if(IS_ERR(inode)){+mutex_unlock(&dlb_driver_mutex);+module_put(THIS_MODULE);+returnERR_CAST(inode);+}++f=alloc_file_pseudo(inode,dlb_vfs_mount,name,flags,fops);+if(IS_ERR(f)){+dlb_free_inode(inode);+mutex_unlock(&dlb_driver_mutex);+module_put(THIS_MODULE);+returnf;+}++mutex_unlock(&dlb_driver_mutex);++returnf;+}
@@ -16,6 +16,9 @@MODULE_LICENSE("GPL v2");MODULE_DESCRIPTION("Intel(R) Dynamic Load Balancer (DLB) Driver");+/* The driver mutex protects data structures that used by multiple devices. */+DEFINE_MUTEX(dlb_driver_mutex);+staticstructclass*dlb_class;staticstructcdevdlb_cdev;staticdev_tdlb_devt;
@@ -226,6 +229,123 @@ const struct file_operations dlb_domain_fops = {.release=dlb_domain_close,};+staticunsignedlongdlb_get_pp_addr(structdlb*dlb,structdlb_port*port)+{+unsignedlongpgoff=dlb->hw.func_phys_addr;++if(port->is_ldb)+pgoff+=DLB_LDB_PP_OFFSET(port->id);+else+pgoff+=DLB_DIR_PP_OFFSET(port->id);++returnpgoff;+}++staticintdlb_pp_mmap(structfile*f,structvm_area_struct*vma)+{+structdlb_port*port=f->private_data;+structdlb_domain*domain=port->domain;+structdlb*dlb=domain->dlb;+unsignedlongpgoff;+pgprot_tpgprot;+intret;++mutex_lock(&dlb->resource_mutex);++if((vma->vm_end-vma->vm_start)!=DLB_PP_SIZE){+ret=-EINVAL;+gotoend;+}++pgprot=pgprot_noncached(vma->vm_page_prot);++pgoff=dlb_get_pp_addr(dlb,port);+ret=io_remap_pfn_range(vma,+vma->vm_start,+pgoff>>PAGE_SHIFT,+vma->vm_end-vma->vm_start,+pgprot);++end:+mutex_unlock(&dlb->resource_mutex);++returnret;+}++staticintdlb_cq_mmap(structfile*f,structvm_area_struct*vma)+{+structdlb_port*port=f->private_data;+structdlb_domain*domain=port->domain;+structdlb*dlb=domain->dlb;+structpage*page;+intret;++mutex_lock(&dlb->resource_mutex);++if((vma->vm_end-vma->vm_start)!=DLB_CQ_SIZE){+ret=-EINVAL;+gotoend;+}++page=virt_to_page(port->cq_base);++ret=remap_pfn_range(vma,+vma->vm_start,+page_to_pfn(page),+vma->vm_end-vma->vm_start,+vma->vm_page_prot);+end:+mutex_unlock(&dlb->resource_mutex);++returnret;+}++staticvoiddlb_port_unmap(structdlb*dlb,structdlb_port*port)+{+if(!port->cq_base){+unmap_mapping_range(dlb->inode->i_mapping,+(unsignedlong)port->cq_base,+DLB_CQ_SIZE,1);+}else{+unmap_mapping_range(dlb->inode->i_mapping,+dlb_get_pp_addr(dlb,port),+DLB_PP_SIZE,1);+}+}++staticintdlb_port_close(structinode*i,structfile*f)+{+structdlb_port*port=f->private_data;+structdlb_domain*domain=port->domain;+structdlb*dlb=domain->dlb;++mutex_lock(&dlb->resource_mutex);++kref_put(&domain->refcnt,dlb_free_domain);++dlb_port_unmap(dlb,port);+dlb_configfs_reset_port_fd(dlb,domain,port->id);++/* Decrement the refcnt of the pseudo-FS used to allocate the inode */+dlb_release_fs(dlb);++mutex_unlock(&dlb->resource_mutex);++return0;+}++conststructfile_operationsdlb_pp_fops={+.owner=THIS_MODULE,+.release=dlb_port_close,+.mmap=dlb_pp_mmap,+};++conststructfile_operationsdlb_cq_fops={+.owner=THIS_MODULE,+.release=dlb_port_close,+.mmap=dlb_cq_mmap,+};+/**********************************//****** PCI driver callbacks ******//**********************************/
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:50
Add configfs interface to start a domain. Once a scheduling domain and its
resources have been configured, this ioctl is called to allow the domain's
ports to begin enqueueing to the device. Once started, the domain's
resources cannot be configured again until after the domain is reset.
A write to "start" configfs file in a domain directory instructs the DLB
device to start load-balancing operations.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_args.h | 17 +++++
drivers/misc/dlb/dlb_configfs.c | 42 +++++++++++++
drivers/misc/dlb/dlb_configfs.h | 1 +
drivers/misc/dlb/dlb_main.h | 2 +
drivers/misc/dlb/dlb_resource.c | 106 ++++++++++++++++++++++++++++++++
5 files changed, 168 insertions(+)
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:52
Add the high-level code for queue map, unmap, and pending unmap query
configfs interface and argument verification -- with stubs for the
low-level register accesses and the queue map/unmap state machine, to be
filled in a later commit.
The queue map/unmap in this commit refers to link/unlink between DLB's
load-balanced queues (internal) and consumer ports.See Documentation/
misc-devices/dlb.rst for details.
Load-balanced queues can be "mapped" to any number of load-balanced ports.
Once mapped, the port becomes a candidate to which the device can schedule
queue entries from the queue. If a port is unmapped from a queue, it is no
longer a candidate for scheduling from that queue.
The pending unmaps function queries how many unmap operations are
in-progress for a given port. These operations are asynchronous, so
multiple may be in-flight at any given time.
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_args.h | 63 ++++++
drivers/misc/dlb/dlb_configfs.c | 101 ++++++++++
drivers/misc/dlb/dlb_configfs.h | 3 +
drivers/misc/dlb/dlb_main.h | 9 +
drivers/misc/dlb/dlb_resource.c | 330 ++++++++++++++++++++++++++++++++
include/uapi/linux/dlb.h | 2 +
6 files changed, 508 insertions(+)
@@ -68,6 +68,9 @@ struct dlb_cfs_port {unsignedintcq_history_list_size;unsignedintcreate;+/* For LDB port only */+unsignedintqueue_link[DLB_MAX_NUM_QIDS_PER_LDB_CQ];+/* For DIR port only, default = 0xffffffff */unsignedintqueue_id;
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:51:55
Add the register accesses that implement the static queue map operation and
handle an unmap request when a queue map operation is in progress.
If a queue map operation is requested before the domain is started, it is a
synchronous procedure on "static"/unchanging hardware. (The "dynamic"
operation, when traffic is flowing in the device, will be added in a later
commit.)
Signed-off-by: Mike Ximing Chen <redacted>
---
drivers/misc/dlb/dlb_resource.c | 163 ++++++++++++++++++++++++++++++++
1 file changed, 163 insertions(+)
@@ -1720,6 +1752,125 @@ static int dlb_configure_dir_port(struct dlb_hw *hw, struct dlb_hw_domain *domaireturn0;}+staticintdlb_ldb_port_map_qid_static(structdlb_hw*hw,structdlb_ldb_port*p,+structdlb_ldb_queue*q,u8priority)+{+u32lsp_qid2cq2;+u32lsp_qid2cq;+u32atm_qid2cq;+u32cq2priov;+u32cq2qid;+inti;++/* Look for a pending or already mapped slot, else an unused slot */+if(!dlb_port_find_slot_queue(p,DLB_QUEUE_MAP_IN_PROG,q,&i)&&+!dlb_port_find_slot_queue(p,DLB_QUEUE_MAPPED,q,&i)&&+!dlb_port_find_slot(p,DLB_QUEUE_UNMAPPED,&i)){+dev_err(hw_to_dev(hw),+"[%s():%d] Internal error: CQ has no available QID mapping slots\n",+__func__,__LINE__);+return-EFAULT;+}++/* Read-modify-write the priority and valid bit register */+cq2priov=DLB_CSR_RD(hw,LSP_CQ2PRIOV(p->id));++cq2priov|=(1U<<(i+LSP_CQ2PRIOV_V_LOC))&LSP_CQ2PRIOV_V;+cq2priov|=((priority&0x7)<<(i+LSP_CQ2PRIOV_PRIO_LOC)*3)+&LSP_CQ2PRIOV_PRIO;++DLB_CSR_WR(hw,LSP_CQ2PRIOV(p->id),cq2priov);++/* Read-modify-write the QID map register */+if(i<4)+cq2qid=DLB_CSR_RD(hw,LSP_CQ2QID0(p->id));+else+cq2qid=DLB_CSR_RD(hw,LSP_CQ2QID1(p->id));++if(i==0||i==4){+cq2qid&=~LSP_CQ2QID0_QID_P0;+cq2qid|=FIELD_PREP(LSP_CQ2QID0_QID_P0,q->id);+}elseif(i==1||i==5){+cq2qid&=~LSP_CQ2QID0_QID_P1;+cq2qid|=FIELD_PREP(LSP_CQ2QID0_QID_P1,q->id);+}elseif(i==2||i==6){+cq2qid&=~LSP_CQ2QID0_QID_P2;+cq2qid|=FIELD_PREP(LSP_CQ2QID0_QID_P2,q->id);+}elseif(i==3||i==7){+cq2qid&=~LSP_CQ2QID0_QID_P3;+cq2qid|=FIELD_PREP(LSP_CQ2QID0_QID_P3,q->id);+}++if(i<4)+DLB_CSR_WR(hw,LSP_CQ2QID0(p->id),cq2qid);+else+DLB_CSR_WR(hw,LSP_CQ2QID1(p->id),cq2qid);++atm_qid2cq=DLB_CSR_RD(hw,+ATM_QID2CQIDIX(q->id,+p->id/4));++lsp_qid2cq=DLB_CSR_RD(hw,+LSP_QID2CQIDIX(q->id,+p->id/4));++lsp_qid2cq2=DLB_CSR_RD(hw,+LSP_QID2CQIDIX2(q->id,+p->id/4));++switch(p->id%4){+case0:+atm_qid2cq|=(1<<(i+ATM_QID2CQIDIX_00_CQ_P0_LOC));+lsp_qid2cq|=(1<<(i+LSP_QID2CQIDIX_00_CQ_P0_LOC));+lsp_qid2cq2|=(1<<(i+LSP_QID2CQIDIX2_00_CQ_P0_LOC));+break;++case1:+atm_qid2cq|=(1<<(i+ATM_QID2CQIDIX_00_CQ_P1_LOC));+lsp_qid2cq|=(1<<(i+LSP_QID2CQIDIX_00_CQ_P1_LOC));+lsp_qid2cq2|=(1<<(i+LSP_QID2CQIDIX2_00_CQ_P1_LOC));+break;++case2:+atm_qid2cq|=(1<<(i+ATM_QID2CQIDIX_00_CQ_P2_LOC));+lsp_qid2cq|=(1<<(i+LSP_QID2CQIDIX_00_CQ_P2_LOC));+lsp_qid2cq2|=(1<<(i+LSP_QID2CQIDIX2_00_CQ_P2_LOC));+break;++case3:+atm_qid2cq|=(1<<(i+ATM_QID2CQIDIX_00_CQ_P3_LOC));+lsp_qid2cq|=(1<<(i+LSP_QID2CQIDIX_00_CQ_P3_LOC));+lsp_qid2cq2|=(1<<(i+LSP_QID2CQIDIX2_00_CQ_P3_LOC));+break;+}++DLB_CSR_WR(hw,+ATM_QID2CQIDIX(q->id,p->id/4),+atm_qid2cq);++DLB_CSR_WR(hw,+LSP_QID2CQIDIX(q->id,p->id/4),+lsp_qid2cq);++DLB_CSR_WR(hw,+LSP_QID2CQIDIX2(q->id,p->id/4),+lsp_qid2cq2);++dlb_flush_csr(hw);++p->qid_map[i].qid=q->id;+p->qid_map[i].priority=priority;++return0;+}++staticintdlb_ldb_port_map_qid(structdlb_hw*hw,structdlb_hw_domain*domain,+structdlb_ldb_port*port,+structdlb_ldb_queue*queue,u8prio)+{+returndlb_ldb_port_map_qid_static(hw,port,queue,prio);+}+staticvoiddlb_log_create_sched_domain_args(structdlb_hw*hw,structdlb_create_sched_domain_args*args)
@@ -2155,6 +2306,7 @@ int dlb_hw_map_qid(struct dlb_hw *hw, u32 domain_id,structdlb_ldb_queue*queue;structdlb_ldb_port*port;intret;+u8prio;dlb_log_map_qid(hw,domain_id,args);
@@ -2167,6 +2319,17 @@ int dlb_hw_map_qid(struct dlb_hw *hw, u32 domain_id,if(ret)returnret;+prio=args->priority;++ret=dlb_ldb_port_map_qid(hw,domain,port,queue,prio);++/* If ret is less than zero, it's due to an internal error */+if(ret<0)+returnret;++if(port->enabled)+dlb_ldb_port_cq_enable(hw,port);+resp->status=0;return0;
From: Mike Ximing Chen <hidden> Date: 2021-12-21 06:52:00
The dlb sysfs interfaces include files for reading the total and
available device resources, and reading the device ID and version. The
interfaces are used for device level configurations and resource
inquiries.
Signed-off-by: Mike Ximing Chen <redacted>
---
Documentation/ABI/testing/sysfs-driver-dlb | 116 ++++++++++++
drivers/misc/dlb/dlb_args.h | 34 ++++
drivers/misc/dlb/dlb_main.c | 5 +
drivers/misc/dlb/dlb_main.h | 3 +
drivers/misc/dlb/dlb_pf_ops.c | 195 +++++++++++++++++++++
drivers/misc/dlb/dlb_resource.c | 50 ++++++
6 files changed, 403 insertions(+)
create mode 100644 Documentation/ABI/testing/sysfs-driver-dlb
@@ -0,0 +1,116 @@+What: /sys/bus/pci/devices/.../total_resources/num_atomic_inflights+What: /sys/bus/pci/devices/.../total_resources/num_dir_credits+What: /sys/bus/pci/devices/.../total_resources/num_dir_ports+What: /sys/bus/pci/devices/.../total_resources/num_hist_list_entries+What: /sys/bus/pci/devices/.../total_resources/num_ldb_credits+What: /sys/bus/pci/devices/.../total_resources/num_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos0_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos1_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos2_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos3_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_ldb_queues+What: /sys/bus/pci/devices/.../total_resources/num_sched_domains+Date: Oct 15, 2021+KernelVersion: 5.15+Contact: mike.ximing.chen@intel.com+Description:+ The total_resources subdirectory contains read-only files that+ indicate the total number of resources in the device.++ num_atomic_inflights: Total number of atomic inflights in the+ device. Atomic inflights refers to the+ on-device storage used by the atomic+ scheduler.++ num_dir_credits: Total number of directed credits in the+ device.++ num_dir_ports: Total number of directed ports (and+ queues) in the device.++ num_hist_list_entries: Total number of history list entries in+ the device.++ num_ldb_credits: Total number of load-balanced credits in+ the device.++ num_ldb_ports: Total number of load-balanced ports in+ the device.++ num_cos<M>_ldb_ports: Total number of load-balanced ports+ belonging to class-of-service M in the+ device.++ num_ldb_queues: Total number of load-balanced queues in+ the device.++ num_sched_domains: Total number of scheduling domains in the+ device.++What: /sys/bus/pci/devices/.../avail_resources/num_atomic_inflights+What: /sys/bus/pci/devices/.../avail_resources/num_dir_credits+What: /sys/bus/pci/devices/.../avail_resources/num_dir_ports+What: /sys/bus/pci/devices/.../avail_resources/num_hist_list_entries+What: /sys/bus/pci/devices/.../avail_resources/num_ldb_credits+What: /sys/bus/pci/devices/.../avail_resources/num_ldb_ports+What: /sys/bus/pci/devices/.../avail_resources/num_cos0_ldb_ports+What: /sys/bus/pci/devices/.../avail_resources/num_cos1_ldb_ports+What: /sys/bus/pci/devices/.../avail_resources/num_cos2_ldb_ports+What: /sys/bus/pci/devices/.../avail_resources/num_cos3_ldb_ports+What: /sys/bus/pci/devices/.../avail_resources/num_ldb_queues+What: /sys/bus/pci/devices/.../avail_resources/num_sched_domains+What: /sys/bus/pci/devices/.../avail_resources/max_ctg_hl_entries+Date: Oct 15, 2021+KernelVersion: 5.15+Contact: mike.ximing.chen@intel.com+Description:+ The avail_resources subdirectory contains read-only files that+ indicate the available number of resources in the device.+ "Available" here means resources that are not currently in use+ by an application or, in the case of a physical function+ device, assigned to a virtual function.++ num_atomic_inflights: Available number of atomic inflights in+ the device.++ num_dir_ports: Available number of directed ports (and+ queues) in the device.++ num_hist_list_entries: Available number of history list entries+ in the device.++ num_ldb_credits: Available number of load-balanced credits+ in the device.++ num_ldb_ports: Available number of load-balanced ports+ in the device.++ num_cos<M>_ldb_ports: Available number of load-balanced ports+ belonging to class-of-service M in the+ device.++ num_ldb_queues: Available number of load-balanced queues+ in the device.++ num_sched_domains: Available number of scheduling domains in+ the device.++ max_ctg_hl_entries: Maximum contiguous history list entries+ available in the device.++ Each scheduling domain is created with+ an allocation of history list entries,+ and each domain's allocation of entries+ must be contiguous.++What: /sys/bus/pci/devices/.../dev_id+Date: Oct 15, 2021+KernelVersion: 5.15+Contact: mike.ximing.chen@intel.com+Description: Device ID used in /dev, i.e. /dev/dlb<device ID>++ Each DLB 2.0 PF and VF device is granted a unique ID by the+ kernel driver, and this ID is used to construct the device's+ /dev directory: /dev/dlb<device ID>. This sysfs file can be read+ to determine a device's ID, which allows the user to map a+ device file to a PCI BDF.
On Tue, Dec 21, 2021 at 12:50:30AM -0600, Mike Ximing Chen wrote:
v12:
<snip>
How is a "RFC" series on version 12? "RFC" means "I do not think this
should be merged, please give me some comments on how this is all
structured" which I think is not the case here.
quoted hunk
- The following coding style changes suggested by Dan will be implemented in the next revision-- Replace DLB_CSR_RD() and DLB_CSR_WR() with direct ioread32() and iowrite32() call.-- Remove bitmap wrappers and use linux bitmap functions directly.-- Use trace_event in configfs attribute file update.
Why submit a patch series that you know will be changed? Just do the
work, don't ask anyone to review stuff you know is incorrect, that just
wastes our time and ensures that we never want to review it again.
greg k-h
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */
So you did not touch this at all in 2021? And it had a copyrightable
changed added to it for every year, inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
greg k-h
From: Joe Perches <joe@perches.com> Date: 2021-12-21 07:20:34
On Tue, 2021-12-21 at 00:50 -0600, Mike Ximing Chen wrote:
The dlb sysfs interfaces include files for reading the total and
available device resources, and reading the device ID and version. The
interfaces are used for device level configurations and resource
inquiries.
@@ -58,6 +58,40 @@ struct dlb_create_sched_domain_args { __u32 num_dir_credits; };+/*+ * dlb_get_num_resources_args: Used to get the number of available resources+ * (queues, ports, etc.) that this device owns.+ *+ * Output parameters:+ * @response.status: Detailed error code. In certain cases, such as if the+ * request arg is invalid, the driver won't set status.+ * @num_domains: Number of available scheduling domains.+ * @num_ldb_queues: Number of available load-balanced queues.+ * @num_ldb_ports: Total number of available load-balanced ports.+ * @num_dir_ports: Number of available directed ports. There is one directed+ * queue for every directed port.+ * @num_atomic_inflights: Amount of available temporary atomic QE storage.+ * @num_hist_list_entries: Amount of history list storage.+ * @max_contiguous_hist_list_entries: History list storage is allocated in+ * a contiguous chunk, and this return value is the longest available+ * contiguous range of history list entries.+ * @num_ldb_credits: Amount of available load-balanced QE storage.+ * @num_dir_credits: Amount of available directed QE storage.+ */
Is this supposed to be kernel-doc format with /** as the comment initiator ?
On Tue, Dec 21, 2021 at 08:12:00AM +0100, Greg KH wrote:
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
quoted
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */
So you did not touch this at all in 2021? And it had a copyrightable
changed added to it for every year, inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
Ah, this was a "throw it over the fence at the community to handle for
me before I go on vacation" type of posting, based on your autoresponse
email that happened when I sent this.
That too isn't the most kind thing, would you want to be the reviewer of
this if it were sent to you? Please take some time and start doing
patch reviews for the char/misc drivers on the mailing list before
submitting any more new code.
Also, this patch series goes agains the internal rules that I know your
company has, why is that? Those rules are there for a good reason, and
by ignoring them, it's going to make it much harder to get patches to be
reviewed.
best of luck!
greg k-h
From: Andrew Lunn <andrew@lunn.ch> Date: 2021-12-21 09:40:52
1. Before a scheduling domain is created/enabled, a set of parameters are
passed to the kernel driver via configfs attribute files in an configfs domain
directory (say $domain) created by user. Each attribute file corresponds to
a configuration parameter of the domain. After writing to all the attribute
files, user writes 1 to "create" attribute, which triggers an action (i.e.,
domain creation) in the kernel driver. Since multiple processes/users can
access the $domain directory, multiple users can write to the attribute files
at the same time. How do we guarantee an atomic update/configuration of a
domain? In other words, if user A wants to set attributes 1 and 2, how can we
prevent user B from changing attribute 1 and 2 before user A writes 1 to
"create"? A configfs directory with individual attribute files seems to not
be able to provide atomic configuration in this case. One option to solve this
issue could be write a structured data (with a set of parameters) to a single
attribute file. This would guarantee the atomic configuration, but may not be
a conventional configfs operation.
How about throw away configfs and use netlink? Messages are atomic,
and you can add an arbitrary number of attributes to a single netlink
message. It will also make your code more network like, since nothing
else in the network stack uses configfs, as far as i know.
Andrew
This is the only mention of NIC here. Does the application interface
to the network stack in the usual way to receive packets from the
TCP/IP stack up into user space and then copy it back down into the
MMIO block for it to enter the DLB for the first time? And at the end
of the path, does the application copy it from the MMIO into a
standard socket for TCP/IP processing to be send out the NIC?
Do you even needs NICs here? Could the data be coming of a video
camera and you are distributing image processing over a number of
cores?
Andrew
From: Chen, Mike Ximing <hidden> Date: 2021-12-21 14:03:57
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 2:10 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 00/17] dlb: introduce DLB device driver
On Tue, Dec 21, 2021 at 12:50:30AM -0600, Mike Ximing Chen wrote:
quoted
v12:
<snip>
How is a "RFC" series on version 12? "RFC" means "I do not think this should be merged, please give me
some comments on how this is all structured" which I think is not the case here.
Hi Greg,
"RFC" here means exactly what you referred to. As you know we have made many changes since your
last review of the patch set (which was v10). At this point we are not sure if we are on the right track in
terms of some configfs implementation, and would like some comments from the community. I stated
this in the cover letter before the change log: " This submission is still a work in progress.... , a couple of
issues that we would like to get help and suggestions from reviewers and community". I presented two
issues/questions we are facing, and would like to get comments.
The code on the other hand are tested and validated on our hardware platforms. I kept the version number
in series (using v12, instead v1) so that reviewers can track the old submissions and have a better
understanding of the patch set's history.
quoted
- The following coding style changes suggested by Dan will be implemented in the next revision-- Replace DLB_CSR_RD() and DLB_CSR_WR() with direct ioread32() and iowrite32() call.-- Remove bitmap wrappers and use linux bitmap functions directly.-- Use trace_event in configfs attribute file update.
Why submit a patch series that you know will be changed? Just do the work, don't ask anyone to review
stuff you know is incorrect, that just wastes our time and ensures that we never want to review it again.
Since this is a RFC, and is not for merging or a full review, we though it was OK to log the pending coding
style changes. The patch set was submitted and reviewed by the community before, and there was no
complains on using macros like DLB_CSR_RD(), etc, but we think we can replace them for better
readability of the code.
Thanks
Mike
From: Chen, Mike Ximing <hidden> Date: 2021-12-21 14:05:42
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 2:12 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
quoted
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved. */
So you did not touch this at all in 2021? And it had a copyrightable changed added to it for every year,
inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
greg k-h
From: Chen, Mike Ximing <hidden> Date: 2021-12-21 14:26:04
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 3:57 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
On Tue, Dec 21, 2021 at 08:12:00AM +0100, Greg KH wrote:
quoted
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
quoted
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved.
+*/
So you did not touch this at all in 2021? And it had a copyrightable
changed added to it for every year, inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
Ah, this was a "throw it over the fence at the community to handle for me before I go on vacation" type of
posting, based on your autoresponse email that happened when I sent this.
That too isn't the most kind thing, would you want to be the reviewer of this if it were sent to you? Please
take some time and start doing patch reviews for the char/misc drivers on the mailing list before
submitting any more new code.
Also, this patch series goes agains the internal rules that I know your company has, why is that? Those
rules are there for a good reason, and by ignoring them, it's going to make it much harder to get patches
to be reviewed.
I assume that you referred to the "Reviewed-by" rule from Intel. Since this is a RFC and we are seeking for
comments and guidance on our code structure, we thought it was appropriate to sent out patch set out
with a full endorsement from our internal reviewers. The questions I posted in the cover letter
(patch 00/17) are from the discussions with our internal reviewers.
I will take some days off as many people would do during this time of the year 😊, but will check mails daily
and response to questions/comments on the submission.
Thanks for your help.
Mike
From: Chen, Mike Ximing <hidden> Date: 2021-12-21 14:42:23
-----Original Message-----
From: Chen, Mike Ximing <redacted>
Sent: Tuesday, December 21, 2021 9:26 AM
To: Greg KH <gregkh@linuxfoundation.org>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: RE: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
quoted
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 3:57 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J
[off-list ref]; pierre- louis.bossart@linux.intel.com;
netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
On Tue, Dec 21, 2021 at 08:12:00AM +0100, Greg KH wrote:
quoted
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
quoted
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved.
+*/
So you did not touch this at all in 2021? And it had a
copyrightable changed added to it for every year, inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
Ah, this was a "throw it over the fence at the community to handle for
me before I go on vacation" type of posting, based on your autoresponse email that happened when I
sent this.
quoted
That too isn't the most kind thing, would you want to be the reviewer
of this if it were sent to you? Please take some time and start doing
patch reviews for the char/misc drivers on the mailing list before submitting any more new code.
Also, this patch series goes agains the internal rules that I know
your company has, why is that? Those rules are there for a good
reason, and by ignoring them, it's going to make it much harder to get patches to be reviewed.
I assume that you referred to the "Reviewed-by" rule from Intel. Since this is a RFC and we are seeking for
comments and guidance on our code structure, we thought it was appropriate to send out patch set out
with a full endorsement from our internal reviewers. The questions I posted in the cover letter (patch
00/17) are from the discussions with our internal reviewers.
.
"we thought it was appropriate to send out the patch set out without* a full endorsement from our
Internal reviewers" --- sorry for misspelling.
I will take some days off as many people would do during this time of the year 😊, but will check mails
daily and response to questions/comments on the submission.
Thanks for your help.
Mike
On Tue, Dec 21, 2021 at 02:03:38PM +0000, Chen, Mike Ximing wrote:
quoted
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 2:10 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 00/17] dlb: introduce DLB device driver
On Tue, Dec 21, 2021 at 12:50:30AM -0600, Mike Ximing Chen wrote:
quoted
v12:
<snip>
How is a "RFC" series on version 12? "RFC" means "I do not think this should be merged, please give me
some comments on how this is all structured" which I think is not the case here.
Hi Greg,
"RFC" here means exactly what you referred to. As you know we have made many changes since your
last review of the patch set (which was v10). At this point we are not sure if we are on the right track in
terms of some configfs implementation, and would like some comments from the community. I stated
this in the cover letter before the change log: " This submission is still a work in progress.... , a couple of
issues that we would like to get help and suggestions from reviewers and community". I presented two
issues/questions we are facing, and would like to get comments.
The code on the other hand are tested and validated on our hardware platforms. I kept the version number
in series (using v12, instead v1) so that reviewers can track the old submissions and have a better
understanding of the patch set's history.
"RFC" means "I have no idea if this is correct, I am throwing it out
there and anyone who also cares about this type of thing, please
comment".
A patch that is on "RFC 12" means, "We all have no clue how to do this,
we give up and hope you all will do it for us."
I almost never comment on RFC patch series, except for portions of the
kernel that I really care about. For a brand-new subsystem like this,
that I still do not understand who needs it, that is not the case.
I'm going to stop reviewing this patch series until you at least follow
the Intel required rules for sending kernel patches like this out. To
not do so would be unfair to your coworkers who _DO_ follow the rules.
quoted
quoted
- The following coding style changes suggested by Dan will be implemented in the next revision-- Replace DLB_CSR_RD() and DLB_CSR_WR() with direct ioread32() and iowrite32() call.-- Remove bitmap wrappers and use linux bitmap functions directly.-- Use trace_event in configfs attribute file update.
Why submit a patch series that you know will be changed? Just do the work, don't ask anyone to review
stuff you know is incorrect, that just wastes our time and ensures that we never want to review it again.
Since this is a RFC, and is not for merging or a full review, we though it was OK to log the pending coding
style changes. The patch set was submitted and reviewed by the community before, and there was no
complains on using macros like DLB_CSR_RD(), etc, but we think we can replace them for better
readability of the code.
Coding style changes should NEVER be ignored and put off for later.
To do so means you do not care about the brains of anyone who you are
wanting to read this code. We have a coding style because of brains and
pattern matching, not because we are being mean.
good luck,
greg k-h
On Tue, Dec 21, 2021 at 02:42:10PM +0000, Chen, Mike Ximing wrote:
quoted
-----Original Message-----
From: Chen, Mike Ximing <redacted>
Sent: Tuesday, December 21, 2021 9:26 AM
To: Greg KH <gregkh@linuxfoundation.org>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: RE: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
quoted
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 3:57 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J
[off-list ref]; pierre- louis.bossart@linux.intel.com;
netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
On Tue, Dec 21, 2021 at 08:12:00AM +0100, Greg KH wrote:
quoted
On Tue, Dec 21, 2021 at 12:50:31AM -0600, Mike Ximing Chen wrote:
quoted
+/* Copyright(C) 2016-2020 Intel Corporation. All rights reserved.
+*/
So you did not touch this at all in 2021? And it had a
copyrightable changed added to it for every year, inclusive, from 2016-2020?
Please run this past your lawyers on how to do this properly.
Ah, this was a "throw it over the fence at the community to handle for
me before I go on vacation" type of posting, based on your autoresponse email that happened when I
sent this.
quoted
That too isn't the most kind thing, would you want to be the reviewer
of this if it were sent to you? Please take some time and start doing
patch reviews for the char/misc drivers on the mailing list before submitting any more new code.
Also, this patch series goes agains the internal rules that I know
your company has, why is that? Those rules are there for a good
reason, and by ignoring them, it's going to make it much harder to get patches to be reviewed.
I assume that you referred to the "Reviewed-by" rule from Intel. Since this is a RFC and we are seeking for
comments and guidance on our code structure, we thought it was appropriate to send out patch set out
with a full endorsement from our internal reviewers. The questions I posted in the cover letter (patch
00/17) are from the discussions with our internal reviewers.
.
"we thought it was appropriate to send out the patch set out without* a full endorsement from our
Internal reviewers" --- sorry for misspelling.
I think you mean something like "we had an internal deadline to meet so
we are willing to ignore our rules and throw it over the wall and let
the community worry about it."
Congratulations, I now feel like the hours I spent last week talking to
hundreds of Intel developers about this very topic were totally wasted.
{sigh}
Intel now owes me another bottle of liquor...
greg k-h
From: Dan Williams <hidden> Date: 2021-12-21 18:44:23
[ add Christoph for configfs feedback ]
On Tue, Dec 21, 2021 at 7:03 AM Greg KH [off-list ref] wrote:
On Tue, Dec 21, 2021 at 02:03:38PM +0000, Chen, Mike Ximing wrote:
quoted
quoted
-----Original Message-----
From: Greg KH <gregkh@linuxfoundation.org>
Sent: Tuesday, December 21, 2021 2:10 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 00/17] dlb: introduce DLB device driver
On Tue, Dec 21, 2021 at 12:50:30AM -0600, Mike Ximing Chen wrote:
quoted
v12:
<snip>
How is a "RFC" series on version 12? "RFC" means "I do not think this should be merged, please give me
some comments on how this is all structured" which I think is not the case here.
Hi Greg,
"RFC" here means exactly what you referred to. As you know we have made many changes since your
last review of the patch set (which was v10). At this point we are not sure if we are on the right track in
terms of some configfs implementation, and would like some comments from the community. I stated
this in the cover letter before the change log: " This submission is still a work in progress.... , a couple of
issues that we would like to get help and suggestions from reviewers and community". I presented two
issues/questions we are facing, and would like to get comments.
The code on the other hand are tested and validated on our hardware platforms. I kept the version number
in series (using v12, instead v1) so that reviewers can track the old submissions and have a better
understanding of the patch set's history.
"RFC" means "I have no idea if this is correct, I am throwing it out
there and anyone who also cares about this type of thing, please
comment".
A patch that is on "RFC 12" means, "We all have no clue how to do this,
we give up and hope you all will do it for us."
I almost never comment on RFC patch series, except for portions of the
kernel that I really care about. For a brand-new subsystem like this,
that I still do not understand who needs it, that is not the case.
I'm going to stop reviewing this patch series until you at least follow
the Intel required rules for sending kernel patches like this out. To
not do so would be unfair to your coworkers who _DO_ follow the rules.
quoted
quoted
quoted
- The following coding style changes suggested by Dan will be implemented in the next revision-- Replace DLB_CSR_RD() and DLB_CSR_WR() with direct ioread32() and iowrite32() call.-- Remove bitmap wrappers and use linux bitmap functions directly.-- Use trace_event in configfs attribute file update.
Why submit a patch series that you know will be changed? Just do the work, don't ask anyone to review
stuff you know is incorrect, that just wastes our time and ensures that we never want to review it again.
Since this is a RFC, and is not for merging or a full review, we though it was OK to log the pending coding
style changes. The patch set was submitted and reviewed by the community before, and there was no
complains on using macros like DLB_CSR_RD(), etc, but we think we can replace them for better
readability of the code.
Coding style changes should NEVER be ignored and put off for later.
To do so means you do not care about the brains of anyone who you are
wanting to read this code. We have a coding style because of brains and
pattern matching, not because we are being mean.
Hey Greg,
This is my fault.
To date Mike has been patiently and diligently following my review
feedback to continue to make the driver smaller and more Linux
idiomatic. Primarily this has been ripping and replacing a pile of
object configuration ioctls with configfs. While my confidence in that
review feedback was high, my confidence in the current round of deeper
architecture reworks is lower and they seemed to raise questions that
are likely FAQs with using configfs. Specifically the observation that
configfs, like sysfs, lacks an "atomically update multiple attributes"
capability. To my knowledge that's just the expected tradeoff with
pseudo-fs based configuration and it is up to userspace to coordinate
multiple configuration writers.
The other question is the use of anon_inode_getfd(). To me that
mechanism is reserved for syscall and ioctl based architectures, and
in this case it was only being used as a mechanism to get an automatic
teardown action at process exit. Again, my inclination is that configs
requires userspace to clean up anything it created. If "tear down on
last close" behavior is needed that would either need to come from a
userspace daemon to watch clients, or another character device that
clients could open to represent the active users of the configuration.
My preference is for the former.
I green-lighted the work-in-progress / RFC posting (with the known
style warts) to get momentum on just those questions. I thought it
better to not polish this driver to a shine and get some mid-rework
feedback. Mike continues to be a pleasure to work with, please take
any frustrations on how this was presented out on me, I'll do better
next time for these types of questions. DLB is unlike anything I have
reviewed previously.
From: Andrew Lunn <andrew@lunn.ch> Date: 2021-12-21 19:57:46
Hey Greg,
This is my fault.
To date Mike has been patiently and diligently following my review
feedback to continue to make the driver smaller and more Linux
idiomatic. Primarily this has been ripping and replacing a pile of
object configuration ioctls with configfs. While my confidence in that
review feedback was high, my confidence in the current round of deeper
architecture reworks is lower and they seemed to raise questions that
are likely FAQs with using configfs. Specifically the observation that
configfs, like sysfs, lacks an "atomically update multiple attributes"
capability. To my knowledge that's just the expected tradeoff with
pseudo-fs based configuration and it is up to userspace to coordinate
multiple configuration writers.
Hi Dan
If this is considered a network accelerator, it probably should use
the same configuration mechanisms all of networking uses, netlink. I'm
not aware of anything network related using configfs, but it could
exist. netlink messages should also solve your atomisity problem.
But it does not really help with cleanup when the userspace user goes
away. Is there anything from GPU drivers which can be reused? They
must have some sort of cleanup when the user space DRM driver exits.
Andrew
This is the only mention of NIC here. Does the application interface to the network stack in the usual way
to receive packets from the TCP/IP stack up into user space and then copy it back down into the MMIO
block for it to enter the DLB for the first time? And at the end of the path, does the application copy it
from the MMIO into a standard socket for TCP/IP processing to be send out the NIC?
For load balancing and distribution purposes, we do not handle packets directly in DLB. Instead, we only
send QEs (queue events) to MMIO for DLB to process. In an network application, QEs (64 bytes each) can
contain pointers to the actual packets. The worker cores can use these pointers to process packets and
forward them to the next stage. At the end of the path, the last work core can send the packets out to NIC.
Do you even needs NICs here? Could the data be coming of a video camera and you are distributing image
processing over a number of cores?
No, the diagram is just an example for packet processing applications. The data can come from other sources
such video cameras. The DLB can schedule up to 100 million packets/events per seconds. The frame rate from
a single camera is normally much, much lower than that.
This is the only mention of NIC here. Does the application interface to the network stack in the usual way
to receive packets from the TCP/IP stack up into user space and then copy it back down into the MMIO
block for it to enter the DLB for the first time? And at the end of the path, does the application copy it
from the MMIO into a standard socket for TCP/IP processing to be send out the NIC?
For load balancing and distribution purposes, we do not handle packets directly in DLB. Instead, we only
send QEs (queue events) to MMIO for DLB to process. In an network application, QEs (64 bytes each) can
contain pointers to the actual packets. The worker cores can use these pointers to process packets and
forward them to the next stage. At the end of the path, the last work core can send the packets out to NIC.
Sorry for asking so many questions, but i'm trying to understand the
architecture. As a network maintainer, and somebody who reviews
network drivers, i was trying to be sure there is not an actual
network MAC and PHY driver hiding in this code.
So you talk about packets. Do you actually mean frames? As in Ethernet
frames? TCP/IP processing has not occurred? Or does this plug into the
network stack at some level? After TCP reassembly has occurred? Are
these pointers to skbufs?
quoted
Do you even needs NICs here? Could the data be coming of a video camera and you are distributing image
processing over a number of cores?
No, the diagram is just an example for packet processing applications. The data can come from other sources
such video cameras. The DLB can schedule up to 100 million packets/events per seconds. The frame rate from
a single camera is normally much, much lower than that.
So i'm trying to understand the scope of this accelerator. Is it just
a network accelerator? If so, are you pointing to skbufs? How are the
lifetimes of skbufs managed? How do you get skbufs out of the NIC? Are
you using XDP?
Andrew
This is the only mention of NIC here. Does the application interface
to the network stack in the usual way to receive packets from the
TCP/IP stack up into user space and then copy it back down into the
MMIO block for it to enter the DLB for the first time? And at the end of the path, does the application
copy it from the MMIO into a standard socket for TCP/IP processing to be send out the NIC?
quoted
quoted
For load balancing and distribution purposes, we do not handle packets
directly in DLB. Instead, we only send QEs (queue events) to MMIO for
DLB to process. In an network application, QEs (64 bytes each) can
contain pointers to the actual packets. The worker cores can use these pointers to process packets and
forward them to the next stage. At the end of the path, the last work core can send the packets out to NIC.
Sorry for asking so many questions, but i'm trying to understand the architecture. As a network maintainer,
and somebody who reviews network drivers, i was trying to be sure there is not an actual network MAC
and PHY driver hiding in this code.
So you talk about packets. Do you actually mean frames? As in Ethernet frames? TCP/IP processing has not
occurred? Or does this plug into the network stack at some level? After TCP reassembly has occurred? Are
these pointers to skbufs?
There is no network MAC or PHY driver in the code. Actually DLB and the driver does not have any direct access to
the network ports/sockets. In the above diagram, the Rx/Tx CPU core receives/transmits packet (or frames)
from/to the NIC. These can be either L2 or L3 packets/frames. The Rx CPU core sends corresponding QEs with
proper meta data (such as pointers to packets/frames) to DLB, which distributes QEs to a set of worker cores.
the worker cores receive QEs, process the corresponding packets/frames, and send QEs back to DLB for
the next stage processing. After several stages of processing, the worker cores in the last stage send the QEs
to Tx core, which then transmits the packets/frames to NIC ports. So between the Rx core and Tx core is where
DLB and the driver operates. The DLB operation itself does not involve any network access.
I am not very familiar with skbufs, but they sound like queue buffers in the kernel. Most of the DLB applications
are in user space. So these pointers can be for any buffers that an application uses. DLB does not process any
packets/frames, it distributes QEs to worker cores which process the corresponding packets/frames.
quoted
quoted
Do you even needs NICs here? Could the data be coming of a video
camera and you are distributing image processing over a number of cores?
No, the diagram is just an example for packet processing applications.
The data can come from other sources such video cameras. The DLB can
schedule up to 100 million packets/events per seconds. The frame rate from a single camera is normally
much, much lower than that.
So i'm trying to understand the scope of this accelerator. Is it just a network accelerator? If so, are you
pointing to skbufs? How are the lifetimes of skbufs managed? How do you get skbufs out of the NIC? Are
you using XDP?
This is not a network accelerator in the sense that it does not have direct access to the network sockets/ports. We do not use XDP.
What it does is to effectively distribute workloads (such as packet processing) among CPU cores and therefore
increases the total packet/frame processing throughput of the CPU processors (such as Intel's Xeon processors).
Imagine, for example, that the Rx core receives 1000 packets/frames in a burst with random payloads, how to
distribute the packet processing to (say) 16 CPU cores is the job of the DLB hardware. The driver is responsible
for the resource management, system configuration and reset, multiple user/application support, and virtualization
enablement.
Thanks
Mike
From: Chen, Mike Ximing <hidden> Date: 2021-12-21 23:18:29
-----Original Message-----
From: Joe Perches <joe@perches.com>
Sent: Tuesday, December 21, 2021 2:20 AM
To: Chen, Mike Ximing <redacted>; linux-kernel@vger.kernel.org
Cc: arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J <redacted>; pierre-
louis.bossart@linux.intel.com; netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 17/17] dlb: add basic sysfs interfaces
On Tue, 2021-12-21 at 00:50 -0600, Mike Ximing Chen wrote:
quoted
The dlb sysfs interfaces include files for reading the total and
available device resources, and reading the device ID and version. The
interfaces are used for device level configurations and resource
inquiries.
@@ -58,6 +58,40 @@ struct dlb_create_sched_domain_args { __u32 num_dir_credits; };+/*+ * dlb_get_num_resources_args: Used to get the number of available resources+ * (queues, ports, etc.) that this device owns.+ *+ * Output parameters:+ * @response.status: Detailed error code. In certain cases, such as if the+ * request arg is invalid, the driver won't set status.+ * @num_domains: Number of available scheduling domains.+ * @num_ldb_queues: Number of available load-balanced queues.+ * @num_ldb_ports: Total number of available load-balanced ports.+ * @num_dir_ports: Number of available directed ports. There is one directed+ * queue for every directed port.+ * @num_atomic_inflights: Amount of available temporary atomic QE storage.+ * @num_hist_list_entries: Amount of history list storage.+ * @max_contiguous_hist_list_entries: History list storage is allocated in+ * a contiguous chunk, and this return value is the longest available+ * contiguous range of history list entries.+ * @num_ldb_credits: Amount of available load-balanced QE storage.+ * @num_dir_credits: Amount of available directed QE storage.+ */
Is this supposed to be kernel-doc format with /** as the comment initiator ?
From: Stephen Hemminger <stephen@networkplumber.org> Date: 2021-12-21 23:35:03
On Tue, 21 Dec 2021 00:50:47 -0600
Mike Ximing Chen [off-list ref] wrote:
quoted hunk
The dlb sysfs interfaces include files for reading the total and
available device resources, and reading the device ID and version. The
interfaces are used for device level configurations and resource
inquiries.
Signed-off-by: Mike Ximing Chen <redacted>
---
Documentation/ABI/testing/sysfs-driver-dlb | 116 ++++++++++++
drivers/misc/dlb/dlb_args.h | 34 ++++
drivers/misc/dlb/dlb_main.c | 5 +
drivers/misc/dlb/dlb_main.h | 3 +
drivers/misc/dlb/dlb_pf_ops.c | 195 +++++++++++++++++++++
drivers/misc/dlb/dlb_resource.c | 50 ++++++
6 files changed, 403 insertions(+)
create mode 100644 Documentation/ABI/testing/sysfs-driver-dlb
@@ -0,0 +1,116 @@+What: /sys/bus/pci/devices/.../total_resources/num_atomic_inflights+What: /sys/bus/pci/devices/.../total_resources/num_dir_credits+What: /sys/bus/pci/devices/.../total_resources/num_dir_ports+What: /sys/bus/pci/devices/.../total_resources/num_hist_list_entries+What: /sys/bus/pci/devices/.../total_resources/num_ldb_credits+What: /sys/bus/pci/devices/.../total_resources/num_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos0_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos1_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos2_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos3_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_ldb_queues+What: /sys/bus/pci/devices/.../total_resources/num_sched_domains+Date: Oct 15, 2021+KernelVersion: 5.15+Contact: mike.ximing.chen@intel.com+Description:+ The total_resources subdirectory contains read-only files that+ indicate the total number of resources in the device.++ num_atomic_inflights: Total number of atomic inflights in the+ device. Atomic inflights refers to the+ on-device storage used by the atomic+ scheduler.++ num_dir_credits: Total number of directed credits in the+ device.++ num_dir_ports: Total number of directed ports (and+ queues) in the device.++ num_hist_list_entries: Total number of history list entries in+ the device.++ num_ldb_credits: Total number of load-balanced credits in+ the device.++ num_ldb_ports: Total number of load-balanced ports in+ the device.++ num_cos<M>_ldb_ports: Total number of load-balanced ports+ belonging to class-of-service M in the+ device.++ num_ldb_queues: Total number of load-balanced queues in+ the device.++ num_sched_domains: Total number of scheduling domains in the+ device.+
Sysfs is only slightly better than /proc as an API.
If it is just for testing than debugfs might be better.
Could this be done with a real netlink interface?
Maybe as part of devlink?
From: Chen, Mike Ximing <hidden> Date: 2021-12-22 04:21:13
-----Original Message-----
From: Stephen Hemminger <stephen@networkplumber.org>
Sent: Tuesday, December 21, 2021 6:35 PM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com; netdev@vger.kernel.org;
davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 17/17] dlb: add basic sysfs interfaces
On Tue, 21 Dec 2021 00:50:47 -0600
Mike Ximing Chen [off-list ref] wrote:
quoted
The dlb sysfs interfaces include files for reading the total and
available device resources, and reading the device ID and version. The
interfaces are used for device level configurations and resource
inquiries.
Signed-off-by: Mike Ximing Chen <redacted>
---
Documentation/ABI/testing/sysfs-driver-dlb | 116 ++++++++++++
drivers/misc/dlb/dlb_args.h | 34 ++++
drivers/misc/dlb/dlb_main.c | 5 +
drivers/misc/dlb/dlb_main.h | 3 +
drivers/misc/dlb/dlb_pf_ops.c | 195 +++++++++++++++++++++
drivers/misc/dlb/dlb_resource.c | 50 ++++++
6 files changed, 403 insertions(+)
create mode 100644 Documentation/ABI/testing/sysfs-driver-dlb
@@ -0,0 +1,116 @@+What: /sys/bus/pci/devices/.../total_resources/num_atomic_inflights+What: /sys/bus/pci/devices/.../total_resources/num_dir_credits+What: /sys/bus/pci/devices/.../total_resources/num_dir_ports+What: /sys/bus/pci/devices/.../total_resources/num_hist_list_entries+What: /sys/bus/pci/devices/.../total_resources/num_ldb_credits+What: /sys/bus/pci/devices/.../total_resources/num_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos0_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos1_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos2_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_cos3_ldb_ports+What: /sys/bus/pci/devices/.../total_resources/num_ldb_queues+What: /sys/bus/pci/devices/.../total_resources/num_sched_domains+Date: Oct 15, 2021+KernelVersion: 5.15+Contact: mike.ximing.chen@intel.com+Description:+ The total_resources subdirectory contains read-only files that+ indicate the total number of resources in the device.++ num_atomic_inflights: Total number of atomic inflights in the+ device. Atomic inflights refers to the+ on-device storage used by the atomic+ scheduler.++ num_dir_credits: Total number of directed credits in the+ device.++ num_dir_ports: Total number of directed ports (and+ queues) in the device.++ num_hist_list_entries: Total number of history list entries in+ the device.++ num_ldb_credits: Total number of load-balanced credits in+ the device.++ num_ldb_ports: Total number of load-balanced ports in+ the device.++ num_cos<M>_ldb_ports: Total number of load-balanced ports+ belonging to class-of-service M in the+ device.++ num_ldb_queues: Total number of load-balanced queues in+ the device.++ num_sched_domains: Total number of scheduling domains in the+ device.+
Sysfs is only slightly better than /proc as an API.
If it is just for testing than debugfs might be better.
Sysfs in our driver is not only for testing. It is used for the system configuration
at run time.
Could this be done with a real netlink interface?
Maybe as part of devlink?
Thanks for the suggestion. I will look into some sample implementations of
devlink/netlink. Our current plan is to stay with the configfs interface, and
find ways to resolve issues related to atomic update and resource reset at
tear-down time.
Mike
From: Chen, Mike Ximing <hidden> Date: 2021-12-22 04:37:37
-----Original Message-----
From: Andrew Lunn <andrew@lunn.ch>
Sent: Tuesday, December 21, 2021 4:41 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com; netdev@vger.kernel.org;
davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 00/17] dlb: introduce DLB device driver
quoted
1. Before a scheduling domain is created/enabled, a set of parameters
are passed to the kernel driver via configfs attribute files in an
configfs domain directory (say $domain) created by user. Each
attribute file corresponds to a configuration parameter of the domain.
After writing to all the attribute files, user writes 1 to "create"
attribute, which triggers an action (i.e., domain creation) in the
kernel driver. Since multiple processes/users can access the $domain
directory, multiple users can write to the attribute files at the same
time. How do we guarantee an atomic update/configuration of a domain?
In other words, if user A wants to set attributes 1 and 2, how can we
prevent user B from changing attribute 1 and 2 before user A writes 1
to "create"? A configfs directory with individual attribute files
seems to not be able to provide atomic configuration in this case. One
option to solve this issue could be write a structured data (with a
set of parameters) to a single attribute file. This would guarantee the atomic configuration, but may not
be a conventional configfs operation.
How about throw away configfs and use netlink? Messages are atomic, and you can add an arbitrary
number of attributes to a single netlink message. It will also make your code more network like, since
nothing else in the network stack uses configfs, as far as i know.
Hi Andrew,
As I explained in my other response, DLB is not a network accelerator and DLB
driver is not a part of network stack. We would obviously prefer to resolve the
atomic update and resource reset at tear-down Issues within the configfs
framework if possible. But I will take a look at the netlink implementations.
Thanks for the suggestion
Mike
From: Christoph Hellwig <hch@lst.de> Date: 2021-12-22 08:01:07
On Tue, Dec 21, 2021 at 10:44:11AM -0800, Dan Williams wrote:
are likely FAQs with using configfs. Specifically the observation that
configfs, like sysfs, lacks an "atomically update multiple attributes"
capability. To my knowledge that's just the expected tradeoff with
pseudo-fs based configuration and it is up to userspace to coordinate
multiple configuration writers.
Yes. For the SCSI and nvme targets we do a required attributes must
be set before something can be enabled, but that might not work
everywhere.
The other question is the use of anon_inode_getfd(). To me that
mechanism is reserved for syscall and ioctl based architectures, and
It is.
in this case it was only being used as a mechanism to get an automatic
teardown action at process exit. Again, my inclination is that configs
requires userspace to clean up anything it created. If "tear down on
last close" behavior is needed that would either need to come from a
userspace daemon to watch clients, or another character device that
clients could open to represent the active users of the configuration.
My preference is for the former.
This really sounds like configfs is the wrong interface. But I'd have
to find time to see what dlb actually is before commenting on what might
be a better interface.
From: Andrew Lunn <andrew@lunn.ch> Date: 2021-12-22 21:26:48
quoted
pointing to skbufs? How are the lifetimes of skbufs managed? How do you get skbufs out of the NIC? Are
you using XDP?
This is not a network accelerator in the sense that it does not have
direct access to the network sockets/ports. We do not use XDP.
So not using XDP is a problem. I looked at previous versions of this
patch, and it is all DPDK. But DPDK is not in mainline, XDP is. In
order for this to be merged into mainline you need a mainline user of
it.
Maybe you should abandon mainline, and just get this driver merged
into the DPDK fork of Linux?
Andrew
From: Chen, Mike Ximing <hidden> Date: 2021-12-23 05:15:48
-----Original Message-----
From: Andrew Lunn <andrew@lunn.ch>
Sent: Wednesday, December 22, 2021 4:27 PM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com; netdev@vger.kernel.org;
davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
quoted
quoted
pointing to skbufs? How are the lifetimes of skbufs managed? How do
you get skbufs out of the NIC? Are you using XDP?
This is not a network accelerator in the sense that it does not have
direct access to the network sockets/ports. We do not use XDP.
So not using XDP is a problem. I looked at previous versions of this patch, and it is all DPDK. But DPDK is
not in mainline, XDP is. In order for this to be merged into mainline you need a mainline user of it.
Maybe you should abandon mainline, and just get this driver merged into the DPDK fork of Linux?
Hi Andrew,
I am not sure why not using XDP is a problem. As mentioned earlier, the
DLB driver is not a part of network stack.
DPDK is one of applications that can make a good use of DLB, but is not the
only one. We have applications that access DLB directly via the kernel driver API
without using DPDK. Even in a DPDK application, the only part that involves
DLB is the eventdev poll mod driver module, in which we can replace a
traditional software base queue manager with DLB. The network access and
receiving/transmitting data from/to NIC is, on the other hand, handled by
the ethdev in DPDK. The eventdev (and DLB) distributes packet processing
over multiple worker cores.
Thanks
Mike
From: Andrew Lunn <andrew@lunn.ch> Date: 2021-12-23 10:23:02
On Thu, Dec 23, 2021 at 05:15:34AM +0000, Chen, Mike Ximing wrote:
quoted
-----Original Message-----
From: Andrew Lunn <andrew@lunn.ch>
Sent: Wednesday, December 22, 2021 4:27 PM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com; netdev@vger.kernel.org;
davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
quoted
quoted
pointing to skbufs? How are the lifetimes of skbufs managed? How do
you get skbufs out of the NIC? Are you using XDP?
This is not a network accelerator in the sense that it does not have
direct access to the network sockets/ports. We do not use XDP.
So not using XDP is a problem. I looked at previous versions of this patch, and it is all DPDK. But DPDK is
not in mainline, XDP is. In order for this to be merged into mainline you need a mainline user of it.
Maybe you should abandon mainline, and just get this driver merged into the DPDK fork of Linux?
Hi Andrew,
I am not sure why not using XDP is a problem. As mentioned earlier, the
DLB driver is not a part of network stack.
DPDK is one of applications that can make a good use of DLB, but is not the
only one. We have applications that access DLB directly via the kernel driver API
without using DPDK.
Cool. Please can you point at a repo for the code? As i said, we just
need a userspace user, which gives us a good idea how the hardware is
supposed to be used, how the kAPI is to be used, and act as a good
test case for when kernel modifications are made. But it needs to be
pure mainline.
There have been a few good discussion on LWN about accelerators
recently. Worth reading.
Andrew
From: Chen, Mike Ximing <hidden> Date: 2021-12-27 00:45:38
-----Original Message-----
From: Andrew Lunn <andrew@lunn.ch>
Sent: Thursday, December 23, 2021 5:23 AM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de; gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com; netdev@vger.kernel.org;
davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
On Thu, Dec 23, 2021 at 05:15:34AM +0000, Chen, Mike Ximing wrote:
quoted
quoted
-----Original Message-----
From: Andrew Lunn <andrew@lunn.ch>
Sent: Wednesday, December 22, 2021 4:27 PM
To: Chen, Mike Ximing <redacted>
Cc: linux-kernel@vger.kernel.org; arnd@arndb.de;
gregkh@linuxfoundation.org; Williams, Dan J
[off-list ref]; pierre-louis.bossart@linux.intel.com;
netdev@vger.kernel.org; davem@davemloft.net; kuba@kernel.org
Subject: Re: [RFC PATCH v12 01/17] dlb: add skeleton for DLB driver
quoted
quoted
pointing to skbufs? How are the lifetimes of skbufs managed? How
do you get skbufs out of the NIC? Are you using XDP?
This is not a network accelerator in the sense that it does not
have direct access to the network sockets/ports. We do not use XDP.
So not using XDP is a problem. I looked at previous versions of this
patch, and it is all DPDK. But DPDK is not in mainline, XDP is. In order for this to be merged into
mainline you need a mainline user of it.
quoted
quoted
Maybe you should abandon mainline, and just get this driver merged into the DPDK fork of Linux?
Hi Andrew,
I am not sure why not using XDP is a problem. As mentioned earlier,
the DLB driver is not a part of network stack.
DPDK is one of applications that can make a good use of DLB, but is
not the only one. We have applications that access DLB directly via
the kernel driver API without using DPDK.
Cool. Please can you point at a repo for the code? As i said, we just need a userspace user, which gives us
a good idea how the hardware is supposed to be used, how the kAPI is to be used, and act as a good test
case for when kernel modifications are made. But it needs to be pure mainline.
There have been a few good discussion on LWN about accelerators recently. Worth reading.