linux_dsm_epyc7002

mirror of https://github.com/AuxXxilium/linux_dsm_epyc7002.git synced 2024-11-24 02:20:54 +07:00

History

Feng Tang 103c4a08ba mm: relocate 'write_protect_seq' in struct mm_struct [ Upstream commit 2e3025434a6ba090c85871a1d4080ff784109e1f ] 0day robot reported a 9.2% regression for will-it-scale mmap1 test case[1], caused by commit 57efa1fe5957 ("mm/gup: prevent gup_fast from racing with COW during fork"). Further debug shows the regression is due to that commit changes the offset of hot fields 'mmap_lock' inside structure 'mm_struct', thus some cache alignment changes. From the perf data, the contention for 'mmap_lock' is very severe and takes around 95% cpu cycles, and it is a rw_semaphore struct rw_semaphore { atomic_long_t count; /* 8 bytes / atomic_long_t owner; / 8 bytes / struct optimistic_spin_queue osq; / spinner MCS lock / ... Before commit 57efa1fe5957 adds the 'write_protect_seq', it happens to have a very optimal cache alignment layout, as Linus explained: "and before the addition of the 'write_protect_seq' field, the mmap_sem was at offset 120 in 'struct mm_struct'. Which meant that count and owner were in two different cachelines, and then when you have contention and spend time in rwsem_down_write_slowpath(), this is probably exactly* the kind of layout you want. Because first the rwsem_write_trylock() will do a cmpxchg on the first cacheline (for the optimistic fast-path), and then in the case of contention, rwsem_down_write_slowpath() will just access the second cacheline. Which is probably just optimal for a load that spends a lot of time contended - new waiters touch that first cacheline, and then they queue themselves up on the second cacheline." After the commit, the rw_semaphore is at offset 128, which means the 'count' and 'owner' fields are now in the same cacheline, and causes more cache bouncing. Currently there are 3 "#ifdef CONFIG_XXX" before 'mmap_lock' which will affect its offset: CONFIG_MMU CONFIG_MEMBARRIER CONFIG_HAVE_ARCH_COMPAT_MMAP_BASES The layout above is on 64 bits system with 0day's default kernel config (similar to RHEL-8.3's config), in which all these 3 options are 'y'. And the layout can vary with different kernel configs. Relayouting a structure is usually a double-edged sword, as sometimes it can helps one case, but hurt other cases. For this case, one solution is, as the newly added 'write_protect_seq' is a 4 bytes long seqcount_t (when CONFIG_DEBUG_LOCK_ALLOC=n), placing it into an existing 4 bytes hole in 'mm_struct' will not change other fields' alignment, while restoring the regression. Link: https://lore.kernel.org/lkml/20210525031636.GB7744@xsang-OptiPlex-9020/ [1] Reported-by: kernel test robot <oliver.sang@intel.com> Signed-off-by: Feng Tang <feng.tang@intel.com> Reviewed-by: John Hubbard <jhubbard@nvidia.com> Reviewed-by: Jason Gunthorpe <jgg@nvidia.com> Cc: Peter Xu <peterx@redhat.com> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org> Signed-off-by: Sasha Levin <sashal@kernel.org>		2021-06-23 14:42:49 +02:00
..
acpi	ACPI: scan: Use unique number for instance_no	2021-03-30 14:32:06 +02:00
asm-generic	vmlinux.lds.h: Avoid orphan section with !SMP	2021-06-16 12:01:45 +02:00
clocksource
crypto	crypto: poly1305 - fix poly1305_core_setkey() declaration	2021-05-14 09:50:13 +02:00
drm	drm/dp/mst: Export drm_dp_get_vc_payload_bw()	2021-02-10 09:29:18 +01:00
dt-bindings	ASoC: dt-bindings: lpass: Fix and common up lpass dai ids	2021-02-03 23:28:46 +01:00
keys	security: keys: trusted: fix TPM2 authorizations	2021-05-14 09:50:20 +02:00
kunit	kunit: fix display of failed expectations for strings	2020-11-10 13:45:15 -07:00
kvm	ARM:	2020-10-23 11:17:56 -07:00
linux	mm: relocate 'write_protect_seq' in struct mm_struct	2021-06-23 14:42:49 +02:00
math-emu
media	media: v4l2-ctrls: fix reference to freed memory	2021-05-11 14:47:39 +02:00
memory
misc
net	net: make get_net_ns return error if NET_NS is disabled	2021-06-23 14:42:44 +02:00
pcmcia
ras	mm,hwpoison: introduce MF_MSG_UNSPLIT_THP	2020-10-16 11:11:17 -07:00
rdma	RDMA: Lift ibdev_to_node from rds to common code	2021-02-26 10:12:59 +01:00
scsi	Fix misc new gcc warnings	2021-05-11 14:47:36 +02:00
soc	net: dsa: felix: implement port flushing on .phylink_mac_link_down	2021-02-17 11:02:27 +01:00
sound	ALSA: hda: intel-nhlt: verify config type	2021-03-09 11:11:14 +01:00
target	scsi: target: core: Add cmd length set before cmd complete	2021-03-17 17:06:25 +01:00
trace	SUNRPC: Remove trace_xprt_transmit_queued	2021-05-19 10:13:03 +02:00
uapi	icmp: don't send out ICMP messages with a source address of 0.0.0.0	2021-06-23 14:42:47 +02:00
vdso
video	gpu: ipu-v3: remove unused functions	2020-10-26 10:42:38 +01:00
xen	Xen/gntdev: correct error checking in gntdev_map_grant_pages()	2021-02-23 15:53:24 +01:00