LLVM 24.0.0git
MemorySanitizer.cpp
Go to the documentation of this file.
1//===- MemorySanitizer.cpp - detector of uninitialized reads --------------===//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8//
9/// \file
10/// This file is a part of MemorySanitizer, a detector of uninitialized
11/// reads.
12///
13/// The algorithm of the tool is similar to Memcheck
14/// (https://static.usenix.org/event/usenix05/tech/general/full_papers/seward/seward_html/usenix2005.html)
15/// We associate a few shadow bits with every byte of the application memory,
16/// poison the shadow of the malloc-ed or alloca-ed memory, load the shadow,
17/// bits on every memory read, propagate the shadow bits through some of the
18/// arithmetic instruction (including MOV), store the shadow bits on every
19/// memory write, report a bug on some other instructions (e.g. JMP) if the
20/// associated shadow is poisoned.
21///
22/// But there are differences too. The first and the major one:
23/// compiler instrumentation instead of binary instrumentation. This
24/// gives us much better register allocation, possible compiler
25/// optimizations and a fast start-up. But this brings the major issue
26/// as well: msan needs to see all program events, including system
27/// calls and reads/writes in system libraries, so we either need to
28/// compile *everything* with msan or use a binary translation
29/// component (e.g. DynamoRIO) to instrument pre-built libraries.
30/// Another difference from Memcheck is that we use 8 shadow bits per
31/// byte of application memory and use a direct shadow mapping. This
32/// greatly simplifies the instrumentation code and avoids races on
33/// shadow updates (Memcheck is single-threaded so races are not a
34/// concern there. Memcheck uses 2 shadow bits per byte with a slow
35/// path storage that uses 8 bits per byte).
36///
37/// The default value of shadow is 0, which means "clean" (not poisoned).
38///
39/// Every module initializer should call __msan_init to ensure that the
40/// shadow memory is ready. On error, __msan_warning is called. Since
41/// parameters and return values may be passed via registers, we have a
42/// specialized thread-local shadow for return values
43/// (__msan_retval_tls) and parameters (__msan_param_tls).
44///
45/// Origin tracking.
46///
47/// MemorySanitizer can track origins (allocation points) of all uninitialized
48/// values. This behavior is controlled with a flag (msan-track-origins) and is
49/// disabled by default.
50///
51/// Origins are 4-byte values created and interpreted by the runtime library.
52/// They are stored in a second shadow mapping, one 4-byte value for 4 bytes
53/// of application memory. Propagation of origins is basically a bunch of
54/// "select" instructions that pick the origin of a dirty argument, if an
55/// instruction has one.
56///
57/// Every 4 aligned, consecutive bytes of application memory have one origin
58/// value associated with them. If these bytes contain uninitialized data
59/// coming from 2 different allocations, the last store wins. Because of this,
60/// MemorySanitizer reports can show unrelated origins, but this is unlikely in
61/// practice.
62///
63/// Origins are meaningless for fully initialized values, so MemorySanitizer
64/// avoids storing origin to memory when a fully initialized value is stored.
65/// This way it avoids needless overwriting origin of the 4-byte region on
66/// a short (i.e. 1 byte) clean store, and it is also good for performance.
67///
68/// Atomic handling.
69///
70/// Ideally, every atomic store of application value should update the
71/// corresponding shadow location in an atomic way. Unfortunately, atomic store
72/// of two disjoint locations can not be done without severe slowdown.
73///
74/// Therefore, we implement an approximation that may err on the safe side.
75/// In this implementation, every atomically accessed location in the program
76/// may only change from (partially) uninitialized to fully initialized, but
77/// not the other way around. We load the shadow _after_ the application load,
78/// and we store the shadow _before_ the app store. Also, we always store clean
79/// shadow (if the application store is atomic). This way, if the store-load
80/// pair constitutes a happens-before arc, shadow store and load are correctly
81/// ordered such that the load will get either the value that was stored, or
82/// some later value (which is always clean).
83///
84/// This does not work very well with Compare-And-Swap (CAS) and
85/// Read-Modify-Write (RMW) operations. To follow the above logic, CAS and RMW
86/// must store the new shadow before the app operation, and load the shadow
87/// after the app operation. Computers don't work this way. Current
88/// implementation ignores the load aspect of CAS/RMW, always returning a clean
89/// value. It implements the store part as a simple atomic store by storing a
90/// clean shadow.
91///
92/// Instrumenting inline assembly.
93///
94/// For inline assembly code LLVM has little idea about which memory locations
95/// become initialized depending on the arguments. It can be possible to figure
96/// out which arguments are meant to point to inputs and outputs, but the
97/// actual semantics can be only visible at runtime. In the Linux kernel it's
98/// also possible that the arguments only indicate the offset for a base taken
99/// from a segment register, so it's dangerous to treat any asm() arguments as
100/// pointers. We take a conservative approach generating calls to
101/// __msan_instrument_asm_store(ptr, size)
102/// , which defer the memory unpoisoning to the runtime library.
103/// The latter can perform more complex address checks to figure out whether
104/// it's safe to touch the shadow memory.
105/// Like with atomic operations, we call __msan_instrument_asm_store() before
106/// the assembly call, so that changes to the shadow memory will be seen by
107/// other threads together with main memory initialization.
108///
109/// KernelMemorySanitizer (KMSAN) implementation.
110///
111/// The major differences between KMSAN and MSan instrumentation are:
112/// - KMSAN always tracks the origins and implies msan-keep-going=true;
113/// - KMSAN allocates shadow and origin memory for each page separately, so
114/// there are no explicit accesses to shadow and origin in the
115/// instrumentation.
116/// Shadow and origin values for a particular X-byte memory location
117/// (X=1,2,4,8) are accessed through pointers obtained via the
118/// __msan_metadata_ptr_for_load_X(ptr)
119/// __msan_metadata_ptr_for_store_X(ptr)
120/// functions. The corresponding functions check that the X-byte accesses
121/// are possible and returns the pointers to shadow and origin memory.
122/// Arbitrary sized accesses are handled with:
123/// __msan_metadata_ptr_for_load_n(ptr, size)
124/// __msan_metadata_ptr_for_store_n(ptr, size);
125/// Note that the sanitizer code has to deal with how shadow/origin pairs
126/// returned by the these functions are represented in different ABIs. In
127/// the X86_64 ABI they are returned in RDX:RAX, in PowerPC64 they are
128/// returned in r3 and r4, and in the SystemZ ABI they are written to memory
129/// pointed to by a hidden parameter.
130/// - TLS variables are stored in a single per-task struct. A call to a
131/// function __msan_get_context_state() returning a pointer to that struct
132/// is inserted into every instrumented function before the entry block;
133/// - __msan_warning() takes a 32-bit origin parameter;
134/// - local variables are poisoned with __msan_poison_alloca() upon function
135/// entry and unpoisoned with __msan_unpoison_alloca() before leaving the
136/// function;
137/// - the pass doesn't declare any global variables or add global constructors
138/// to the translation unit.
139///
140/// Also, KMSAN currently ignores uninitialized memory passed into inline asm
141/// calls, making sure we're on the safe side wrt. possible false positives.
142///
143/// KernelMemorySanitizer only supports X86_64, SystemZ and PowerPC64 at the
144/// moment.
145///
146//
147// FIXME: This sanitizer does not yet handle scalable vectors
148//
149//===----------------------------------------------------------------------===//
150
153#include "llvm/ADT/APInt.h"
154#include "llvm/ADT/ArrayRef.h"
155#include "llvm/ADT/DenseMap.h"
157#include "llvm/ADT/SetVector.h"
158#include "llvm/ADT/SmallPtrSet.h"
159#include "llvm/ADT/SmallVector.h"
161#include "llvm/ADT/StringRef.h"
165#include "llvm/IR/Argument.h"
167#include "llvm/IR/Attributes.h"
168#include "llvm/IR/BasicBlock.h"
169#include "llvm/IR/CallingConv.h"
170#include "llvm/IR/Constant.h"
171#include "llvm/IR/Constants.h"
172#include "llvm/IR/DataLayout.h"
173#include "llvm/IR/DerivedTypes.h"
174#include "llvm/IR/Function.h"
175#include "llvm/IR/GlobalValue.h"
177#include "llvm/IR/IRBuilder.h"
178#include "llvm/IR/InlineAsm.h"
179#include "llvm/IR/InstVisitor.h"
180#include "llvm/IR/InstrTypes.h"
181#include "llvm/IR/Instruction.h"
182#include "llvm/IR/Instructions.h"
184#include "llvm/IR/Intrinsics.h"
185#include "llvm/IR/IntrinsicsAArch64.h"
186#include "llvm/IR/IntrinsicsX86.h"
187#include "llvm/IR/MDBuilder.h"
188#include "llvm/IR/Module.h"
189#include "llvm/IR/Type.h"
190#include "llvm/IR/Value.h"
191#include "llvm/IR/ValueMap.h"
194#include "llvm/Support/Casting.h"
195#include "llvm/Support/Debug.h"
205#include <algorithm>
206#include <cassert>
207#include <cstddef>
208#include <cstdint>
209#include <memory>
210#include <numeric>
211#include <string>
212#include <tuple>
213
214using namespace llvm;
215
216#define DEBUG_TYPE "msan"
217
218DEBUG_COUNTER(DebugInsertCheck, "msan-insert-check",
219 "Controls which checks to insert");
220
221DEBUG_COUNTER(DebugInstrumentInstruction, "msan-instrument-instruction",
222 "Controls which instruction to instrument");
223
224static const unsigned kOriginSize = 4;
227
228// These constants must be kept in sync with the ones in msan.h.
229// TODO: increase size to match SVE/SVE2/SME/SME2 limits
230static const unsigned kParamTLSSize = 800;
231static const unsigned kRetvalTLSSize = 800;
232
233// Accesses sizes are powers of two: 1, 2, 4, 8.
234static const size_t kNumberOfAccessSizes = 4;
235
236const char kMsanModuleCtorName[] = "msan.module_ctor";
237const char kMsanInitName[] = "__msan_init";
238
239namespace {
240
241// Memory map parameters used in application-to-shadow address calculation.
242// Offset = (Addr & ~AndMask) ^ XorMask
243// Shadow = ShadowBase + Offset
244// Origin = OriginBase + Offset
245struct MemoryMapParams {
246 uint64_t AndMask;
247 uint64_t XorMask;
248 uint64_t ShadowBase;
249 uint64_t OriginBase;
250};
251
252struct PlatformMemoryMapParams {
253 const MemoryMapParams *bits32;
254 const MemoryMapParams *bits64;
255};
256
257} // end anonymous namespace
258
259// i386 Linux
260static const MemoryMapParams Linux_I386_MemoryMapParams = {
261 0x000080000000, // AndMask
262 0, // XorMask (not used)
263 0, // ShadowBase (not used)
264 0x000040000000, // OriginBase
265};
266
267// x86_64 Linux
268static const MemoryMapParams Linux_X86_64_MemoryMapParams = {
269 0, // AndMask (not used)
270 0x500000000000, // XorMask
271 0, // ShadowBase (not used)
272 0x100000000000, // OriginBase
273};
274
275// mips32 Linux
276// FIXME: Remove -msan-origin-base -msan-and-mask added by PR #109284 to tests
277// after picking good constants
278
279// mips64 Linux
280static const MemoryMapParams Linux_MIPS64_MemoryMapParams = {
281 0, // AndMask (not used)
282 0x008000000000, // XorMask
283 0, // ShadowBase (not used)
284 0x002000000000, // OriginBase
285};
286
287// ppc32 Linux
288// FIXME: Remove -msan-origin-base -msan-and-mask added by PR #109284 to tests
289// after picking good constants
290
291// ppc64 Linux
292static const MemoryMapParams Linux_PowerPC64_MemoryMapParams = {
293 0xE00000000000, // AndMask
294 0x100000000000, // XorMask
295 0x080000000000, // ShadowBase
296 0x1C0000000000, // OriginBase
297};
298
299// s390x Linux
300static const MemoryMapParams Linux_S390X_MemoryMapParams = {
301 0xC00000000000, // AndMask
302 0, // XorMask (not used)
303 0x080000000000, // ShadowBase
304 0x1C0000000000, // OriginBase
305};
306
307// arm32 Linux
308// FIXME: Remove -msan-origin-base -msan-and-mask added by PR #109284 to tests
309// after picking good constants
310
311// aarch64 Linux
312static const MemoryMapParams Linux_AArch64_MemoryMapParams = {
313 0, // AndMask (not used)
314 0x0B00000000000, // XorMask
315 0, // ShadowBase (not used)
316 0x0200000000000, // OriginBase
317};
318
319// loongarch64 Linux
320static const MemoryMapParams Linux_LoongArch64_MemoryMapParams = {
321 0, // AndMask (not used)
322 0x500000000000, // XorMask
323 0, // ShadowBase (not used)
324 0x100000000000, // OriginBase
325};
326
327// hexagon Linux
328static const MemoryMapParams Linux_Hexagon_MemoryMapParams = {
329 0, // AndMask (not used)
330 0x20000000, // XorMask
331 0, // ShadowBase (not used)
332 0x50000000, // OriginBase
333};
334
335// riscv32 Linux
336// FIXME: Remove -msan-origin-base -msan-and-mask added by PR #109284 to tests
337// after picking good constants
338
339// aarch64 FreeBSD
340static const MemoryMapParams FreeBSD_AArch64_MemoryMapParams = {
341 0x1800000000000, // AndMask
342 0x0400000000000, // XorMask
343 0x0200000000000, // ShadowBase
344 0x0700000000000, // OriginBase
345};
346
347// i386 FreeBSD
348static const MemoryMapParams FreeBSD_I386_MemoryMapParams = {
349 0x000180000000, // AndMask
350 0x000040000000, // XorMask
351 0x000020000000, // ShadowBase
352 0x000700000000, // OriginBase
353};
354
355// x86_64 FreeBSD
356static const MemoryMapParams FreeBSD_X86_64_MemoryMapParams = {
357 0xc00000000000, // AndMask
358 0x200000000000, // XorMask
359 0x100000000000, // ShadowBase
360 0x380000000000, // OriginBase
361};
362
363// x86_64 NetBSD
364static const MemoryMapParams NetBSD_X86_64_MemoryMapParams = {
365 0, // AndMask
366 0x500000000000, // XorMask
367 0, // ShadowBase
368 0x100000000000, // OriginBase
369};
370
371static const PlatformMemoryMapParams Linux_X86_MemoryMapParams = {
374};
375
376static const PlatformMemoryMapParams Linux_MIPS_MemoryMapParams = {
377 nullptr,
379};
380
381static const PlatformMemoryMapParams Linux_PowerPC_MemoryMapParams = {
382 nullptr,
384};
385
386static const PlatformMemoryMapParams Linux_S390_MemoryMapParams = {
387 nullptr,
389};
390
391static const PlatformMemoryMapParams Linux_ARM_MemoryMapParams = {
392 nullptr,
394};
395
396static const PlatformMemoryMapParams Linux_LoongArch_MemoryMapParams = {
397 nullptr,
399};
400
401static const PlatformMemoryMapParams Linux_Hexagon_MemoryMapParams_P = {
403 nullptr,
404};
405
406static const PlatformMemoryMapParams FreeBSD_ARM_MemoryMapParams = {
407 nullptr,
409};
410
411static const PlatformMemoryMapParams FreeBSD_X86_MemoryMapParams = {
414};
415
416static const PlatformMemoryMapParams NetBSD_X86_MemoryMapParams = {
417 nullptr,
419};
420
422
423namespace {
424
425/// Instrument functions of a module to detect uninitialized reads.
426///
427/// Instantiating MemorySanitizer inserts the msan runtime library API function
428/// declarations into the module if they don't exist already. Instantiating
429/// ensures the __msan_init function is in the list of global constructors for
430/// the module.
431class MemorySanitizer {
432public:
433 MemorySanitizer(const InstrumentationOptions &Opts, Module &M,
434 MemorySanitizerOptions Options)
435 : Opts(Opts), CompileKernel(Options.Kernel),
436 TrackOrigins(Options.TrackOrigins), Recover(Options.Recover),
437 EagerChecks(Options.EagerChecks) {
438 initializeModule(M);
439 }
440
441 // MSan cannot be moved or copied because of MapParams.
442 MemorySanitizer(MemorySanitizer &&) = delete;
443 MemorySanitizer &operator=(MemorySanitizer &&) = delete;
444 MemorySanitizer(const MemorySanitizer &) = delete;
445 MemorySanitizer &operator=(const MemorySanitizer &) = delete;
446
447 bool sanitizeFunction(Function &F, TargetLibraryInfo &TLI);
448
449private:
450 friend struct MemorySanitizerVisitor;
451 friend struct VarArgHelperBase;
452 friend struct VarArgAMD64Helper;
453 friend struct VarArgAArch64Helper;
454 friend struct VarArgPowerPC64Helper;
455 friend struct VarArgPowerPC32Helper;
456 friend struct VarArgSystemZHelper;
457 friend struct VarArgI386Helper;
458 friend struct VarArgGenericHelper;
459
460 void initializeModule(Module &M);
461 void initializeCallbacks(Module &M, const TargetLibraryInfo &TLI);
462 void createKernelApi(Module &M, const TargetLibraryInfo &TLI);
463 void createUserspaceApi(Module &M, const TargetLibraryInfo &TLI);
464
465 template <typename... ArgsTy>
466 FunctionCallee getOrInsertMsanMetadataFunction(Module &M, StringRef Name,
467 ArgsTy... Args);
468
469 const InstrumentationOptions &Opts;
470 /// True if we're compiling the Linux kernel.
471 bool CompileKernel;
472 /// Track origins (allocation points) of uninitialized values.
473 int TrackOrigins;
474 bool Recover;
475 bool EagerChecks;
476
477 Triple TargetTriple;
478 LLVMContext *C;
479 Type *IntptrTy; ///< Integer type with the size of a ptr in default AS.
480 Type *OriginTy;
481 PointerType *PtrTy; ///< Integer type with the size of a ptr in default AS.
482
483 // XxxTLS variables represent the per-thread state in MSan and per-task state
484 // in KMSAN.
485 // For the userspace these point to thread-local globals. In the kernel land
486 // they point to the members of a per-task struct obtained via a call to
487 // __msan_get_context_state().
488
489 /// Thread-local shadow storage for function parameters.
490 Value *ParamTLS;
491
492 /// Thread-local origin storage for function parameters.
493 Value *ParamOriginTLS;
494
495 /// Thread-local shadow storage for function return value.
496 Value *RetvalTLS;
497
498 /// Thread-local origin storage for function return value.
499 Value *RetvalOriginTLS;
500
501 /// Thread-local shadow storage for in-register va_arg function.
502 Value *VAArgTLS;
503
504 /// Thread-local shadow storage for in-register va_arg function.
505 Value *VAArgOriginTLS;
506
507 /// Thread-local shadow storage for va_arg overflow area.
508 Value *VAArgOverflowSizeTLS;
509
510 /// Are the instrumentation callbacks set up?
511 bool CallbacksInitialized = false;
512
513 /// The run-time callback to print a warning.
514 FunctionCallee WarningFn;
515
516 // These arrays are indexed by log2(AccessSize).
517 FunctionCallee MaybeWarningFn[kNumberOfAccessSizes];
518 FunctionCallee MaybeWarningVarSizeFn;
519 FunctionCallee MaybeStoreOriginFn[kNumberOfAccessSizes];
520
521 /// Run-time helper that generates a new origin value for a stack
522 /// allocation.
523 FunctionCallee MsanSetAllocaOriginWithDescriptionFn;
524 // No description version
525 FunctionCallee MsanSetAllocaOriginNoDescriptionFn;
526
527 /// Run-time helper that poisons stack on function entry.
528 FunctionCallee MsanPoisonStackFn;
529
530 /// Run-time helper that records a store (or any event) of an
531 /// uninitialized value and returns an updated origin id encoding this info.
532 FunctionCallee MsanChainOriginFn;
533
534 /// Run-time helper that paints an origin over a region.
535 FunctionCallee MsanSetOriginFn;
536
537 /// MSan runtime replacements for memmove, memcpy and memset.
538 FunctionCallee MemmoveFn, MemcpyFn, MemsetFn;
539
540 /// KMSAN callback for task-local function argument shadow.
541 StructType *MsanContextStateTy;
542 FunctionCallee MsanGetContextStateFn;
543
544 /// Functions for poisoning/unpoisoning local variables
545 FunctionCallee MsanPoisonAllocaFn, MsanUnpoisonAllocaFn;
546
547 /// Pair of shadow/origin pointers.
548 Type *MsanMetadata;
549
550 /// Each of the MsanMetadataPtrXxx functions returns a MsanMetadata.
551 FunctionCallee MsanMetadataPtrForLoadN, MsanMetadataPtrForStoreN;
552 FunctionCallee MsanMetadataPtrForLoad_1_8[4];
553 FunctionCallee MsanMetadataPtrForStore_1_8[4];
554 FunctionCallee MsanInstrumentAsmStoreFn;
555
556 /// Storage for return values of the MsanMetadataPtrXxx functions.
557 Value *MsanMetadataAlloca;
558
559 /// Helper to choose between different MsanMetadataPtrXxx().
560 FunctionCallee getKmsanShadowOriginAccessFn(bool isStore, int size);
561
562 /// Memory map parameters used in application-to-shadow calculation.
563 const MemoryMapParams *MapParams;
564
565 /// Custom memory map parameters used when -msan-shadow-base or
566 // -msan-origin-base is provided.
567 MemoryMapParams CustomMapParams;
568
569 MDNode *ColdCallWeights;
570
571 /// Branch weights for origin store.
572 MDNode *OriginStoreWeights;
573};
574
575void insertModuleCtor(const InstrumentationOptions &Opts, Module &M) {
578 /*InitArgTypes=*/{},
579 /*InitArgs=*/{},
580 // This callback is invoked when the functions are created the first
581 // time. Hook them into the global ctors list in that case:
582 [&](Function *Ctor, FunctionCallee) {
583 if (!Opts.msan_with_comdat) {
584 appendToGlobalCtors(M, Ctor, 0);
585 return;
586 }
587 Comdat *MsanCtorComdat = M.getOrInsertComdat(kMsanModuleCtorName);
588 Ctor->setComdat(MsanCtorComdat);
589 appendToGlobalCtors(M, Ctor, 0, Ctor);
590 });
591}
592
593} // end anonymous namespace
594
596 bool EC) {
597 const InstrumentationOptions &Opts = InstrumentationOptions::Global;
598 Kernel = valueOr(Opts.msan_kernel, K);
599 TrackOrigins = Opts.msan_track_origins.value_or(Kernel ? 2 : TO);
600 Recover = valueOr(Opts.msan_keep_going, Kernel || R);
601 EagerChecks = valueOr(Opts.msan_eager_checks, EC);
602}
603
606 // Return early if nosanitize_memory module flag is present for the module.
607 if (checkIfAlreadyInstrumented(M, "nosanitize_memory"))
608 return PreservedAnalyses::all();
609 const InstrumentationOptions &Opts = InstrumentationOptions::Global;
610 bool Modified = false;
611 if (!Options.Kernel) {
612 insertModuleCtor(Opts, M);
613 Modified = true;
614 }
615
616 auto &FAM = AM.getResult<FunctionAnalysisManagerModuleProxy>(M).getManager();
617 for (Function &F : M) {
618 if (F.empty())
619 continue;
620 MemorySanitizer Msan(Opts, *F.getParent(), Options);
621 Modified |=
622 Msan.sanitizeFunction(F, FAM.getResult<TargetLibraryAnalysis>(F));
623 }
624
625 if (!Modified)
626 return PreservedAnalyses::all();
627
629 // GlobalsAA is considered stateless and does not get invalidated unless
630 // explicitly invalidated; PreservedAnalyses::none() is not enough. Sanitizers
631 // make changes that require GlobalsAA to be invalidated.
632 PA.abandon<GlobalsAA>();
633 return PA;
634}
635
637 raw_ostream &OS, function_ref<StringRef(StringRef)> MapClassName2PassName) {
638 static_cast<PassInfoMixin<MemorySanitizerPass> *>(this)->printPipeline(
639 OS, MapClassName2PassName);
640 OS << '<';
641 if (Options.Recover)
642 OS << "recover;";
643 if (Options.Kernel)
644 OS << "kernel;";
645 if (Options.EagerChecks)
646 OS << "eager-checks;";
647 OS << "track-origins=" << Options.TrackOrigins;
648 OS << '>';
649}
650
651/// Create a non-const global initialized with the given string.
652///
653/// Creates a writable global for Str so that we can pass it to the
654/// run-time lib. Runtime uses first 4 bytes of the string to store the
655/// frame ID, so the string needs to be mutable.
657 StringRef Str) {
658 Constant *StrConst = ConstantDataArray::getString(M.getContext(), Str);
659 return new GlobalVariable(M, StrConst->getType(), /*isConstant=*/true,
660 GlobalValue::PrivateLinkage, StrConst, "");
661}
662
663template <typename... ArgsTy>
665MemorySanitizer::getOrInsertMsanMetadataFunction(Module &M, StringRef Name,
666 ArgsTy... Args) {
667 if (TargetTriple.getArch() == Triple::systemz) {
668 // SystemZ ABI: shadow/origin pair is returned via a hidden parameter.
669 return M.getOrInsertFunction(Name, Type::getVoidTy(*C), PtrTy,
670 std::forward<ArgsTy>(Args)...);
671 }
672
673 return M.getOrInsertFunction(Name, MsanMetadata,
674 std::forward<ArgsTy>(Args)...);
675}
676
677/// Create KMSAN API callbacks.
678void MemorySanitizer::createKernelApi(Module &M, const TargetLibraryInfo &TLI) {
679 IRBuilder<> IRB(M);
680
681 // These will be initialized in insertKmsanPrologue().
682 RetvalTLS = nullptr;
683 RetvalOriginTLS = nullptr;
684 ParamTLS = nullptr;
685 ParamOriginTLS = nullptr;
686 VAArgTLS = nullptr;
687 VAArgOriginTLS = nullptr;
688 VAArgOverflowSizeTLS = nullptr;
689
690 WarningFn = M.getOrInsertFunction("__msan_warning",
691 TLI.getAttrList(C, {0}, /*Signed=*/false),
692 IRB.getVoidTy(), IRB.getInt32Ty());
693
694 // Requests the per-task context state (kmsan_context_state*) from the
695 // runtime library.
696 MsanContextStateTy = StructType::get(
697 ArrayType::get(IRB.getInt64Ty(), kParamTLSSize / 8),
698 ArrayType::get(IRB.getInt64Ty(), kRetvalTLSSize / 8),
699 ArrayType::get(IRB.getInt64Ty(), kParamTLSSize / 8),
700 ArrayType::get(IRB.getInt64Ty(), kParamTLSSize / 8), /* va_arg_origin */
701 IRB.getInt64Ty(), ArrayType::get(OriginTy, kParamTLSSize / 4), OriginTy,
702 OriginTy);
703 MsanGetContextStateFn =
704 M.getOrInsertFunction("__msan_get_context_state", PtrTy);
705
706 MsanMetadata = StructType::get(PtrTy, PtrTy);
707
708 for (int ind = 0, size = 1; ind < 4; ind++, size <<= 1) {
709 std::string name_load =
710 "__msan_metadata_ptr_for_load_" + std::to_string(size);
711 std::string name_store =
712 "__msan_metadata_ptr_for_store_" + std::to_string(size);
713 MsanMetadataPtrForLoad_1_8[ind] =
714 getOrInsertMsanMetadataFunction(M, name_load, PtrTy);
715 MsanMetadataPtrForStore_1_8[ind] =
716 getOrInsertMsanMetadataFunction(M, name_store, PtrTy);
717 }
718
719 MsanMetadataPtrForLoadN = getOrInsertMsanMetadataFunction(
720 M, "__msan_metadata_ptr_for_load_n", PtrTy, IntptrTy);
721 MsanMetadataPtrForStoreN = getOrInsertMsanMetadataFunction(
722 M, "__msan_metadata_ptr_for_store_n", PtrTy, IntptrTy);
723
724 // Functions for poisoning and unpoisoning memory.
725 MsanPoisonAllocaFn = M.getOrInsertFunction(
726 "__msan_poison_alloca", IRB.getVoidTy(), PtrTy, IntptrTy, PtrTy);
727 MsanUnpoisonAllocaFn = M.getOrInsertFunction(
728 "__msan_unpoison_alloca", IRB.getVoidTy(), PtrTy, IntptrTy);
729}
730
732 return M.getOrInsertGlobal(Name, Ty, [&] {
733 return new GlobalVariable(M, Ty, false, GlobalVariable::ExternalLinkage,
734 nullptr, Name, nullptr,
736 });
737}
738
739/// Insert declarations for userspace-specific functions and globals.
740void MemorySanitizer::createUserspaceApi(Module &M,
741 const TargetLibraryInfo &TLI) {
742 IRBuilder<> IRB(M);
743
744 // Create the callback.
745 // FIXME: this function should have "Cold" calling conv,
746 // which is not yet implemented.
747 if (TrackOrigins) {
748 StringRef WarningFnName = Recover ? "__msan_warning_with_origin"
749 : "__msan_warning_with_origin_noreturn";
750 WarningFn = M.getOrInsertFunction(WarningFnName,
751 TLI.getAttrList(C, {0}, /*Signed=*/false),
752 IRB.getVoidTy(), IRB.getInt32Ty());
753 } else {
754 StringRef WarningFnName =
755 Recover ? "__msan_warning" : "__msan_warning_noreturn";
756 WarningFn = M.getOrInsertFunction(WarningFnName, IRB.getVoidTy());
757 }
758
759 // Create the global TLS variables.
760 RetvalTLS =
761 getOrInsertGlobal(M, "__msan_retval_tls",
762 ArrayType::get(IRB.getInt64Ty(), kRetvalTLSSize / 8));
763
764 RetvalOriginTLS = getOrInsertGlobal(M, "__msan_retval_origin_tls", OriginTy);
765
766 ParamTLS =
767 getOrInsertGlobal(M, "__msan_param_tls",
768 ArrayType::get(IRB.getInt64Ty(), kParamTLSSize / 8));
769
770 ParamOriginTLS =
771 getOrInsertGlobal(M, "__msan_param_origin_tls",
772 ArrayType::get(OriginTy, kParamTLSSize / 4));
773
774 VAArgTLS =
775 getOrInsertGlobal(M, "__msan_va_arg_tls",
776 ArrayType::get(IRB.getInt64Ty(), kParamTLSSize / 8));
777
778 VAArgOriginTLS =
779 getOrInsertGlobal(M, "__msan_va_arg_origin_tls",
780 ArrayType::get(OriginTy, kParamTLSSize / 4));
781
782 VAArgOverflowSizeTLS = getOrInsertGlobal(M, "__msan_va_arg_overflow_size_tls",
783 IRB.getIntPtrTy(M.getDataLayout()));
784
785 for (size_t AccessSizeIndex = 0; AccessSizeIndex < kNumberOfAccessSizes;
786 AccessSizeIndex++) {
787 unsigned AccessSize = 1 << AccessSizeIndex;
788 std::string FunctionName = "__msan_maybe_warning_" + itostr(AccessSize);
789 MaybeWarningFn[AccessSizeIndex] = M.getOrInsertFunction(
790 FunctionName, TLI.getAttrList(C, {0, 1}, /*Signed=*/false),
791 IRB.getVoidTy(), IRB.getIntNTy(AccessSize * 8), IRB.getInt32Ty());
792 MaybeWarningVarSizeFn = M.getOrInsertFunction(
793 "__msan_maybe_warning_N", TLI.getAttrList(C, {}, /*Signed=*/false),
794 IRB.getVoidTy(), PtrTy, IRB.getInt64Ty(), IRB.getInt32Ty());
795 FunctionName = "__msan_maybe_store_origin_" + itostr(AccessSize);
796 MaybeStoreOriginFn[AccessSizeIndex] = M.getOrInsertFunction(
797 FunctionName, TLI.getAttrList(C, {0, 2}, /*Signed=*/false),
798 IRB.getVoidTy(), IRB.getIntNTy(AccessSize * 8), PtrTy,
799 IRB.getInt32Ty());
800 }
801
802 MsanSetAllocaOriginWithDescriptionFn =
803 M.getOrInsertFunction("__msan_set_alloca_origin_with_descr",
804 IRB.getVoidTy(), PtrTy, IntptrTy, PtrTy, PtrTy);
805 MsanSetAllocaOriginNoDescriptionFn =
806 M.getOrInsertFunction("__msan_set_alloca_origin_no_descr",
807 IRB.getVoidTy(), PtrTy, IntptrTy, PtrTy);
808 MsanPoisonStackFn = M.getOrInsertFunction("__msan_poison_stack",
809 IRB.getVoidTy(), PtrTy, IntptrTy);
810}
811
812/// Insert extern declaration of runtime-provided functions and globals.
813void MemorySanitizer::initializeCallbacks(Module &M,
814 const TargetLibraryInfo &TLI) {
815 // Only do this once.
816 if (CallbacksInitialized)
817 return;
818
819 IRBuilder<> IRB(M);
820 // Initialize callbacks that are common for kernel and userspace
821 // instrumentation.
822 MsanChainOriginFn = M.getOrInsertFunction(
823 "__msan_chain_origin",
824 TLI.getAttrList(C, {0}, /*Signed=*/false, /*Ret=*/true), IRB.getInt32Ty(),
825 IRB.getInt32Ty());
826 MsanSetOriginFn = M.getOrInsertFunction(
827 "__msan_set_origin", TLI.getAttrList(C, {2}, /*Signed=*/false),
828 IRB.getVoidTy(), PtrTy, IntptrTy, IRB.getInt32Ty());
829 MemmoveFn =
830 M.getOrInsertFunction("__msan_memmove", PtrTy, PtrTy, PtrTy, IntptrTy);
831 MemcpyFn =
832 M.getOrInsertFunction("__msan_memcpy", PtrTy, PtrTy, PtrTy, IntptrTy);
833 MemsetFn = M.getOrInsertFunction("__msan_memset",
834 TLI.getAttrList(C, {1}, /*Signed=*/true),
835 PtrTy, PtrTy, IRB.getInt32Ty(), IntptrTy);
836
837 MsanInstrumentAsmStoreFn = M.getOrInsertFunction(
838 "__msan_instrument_asm_store", IRB.getVoidTy(), PtrTy, IntptrTy);
839
840 if (CompileKernel) {
841 createKernelApi(M, TLI);
842 } else {
843 createUserspaceApi(M, TLI);
844 }
845 CallbacksInitialized = true;
846}
847
848FunctionCallee MemorySanitizer::getKmsanShadowOriginAccessFn(bool isStore,
849 int size) {
850 FunctionCallee *Fns =
851 isStore ? MsanMetadataPtrForStore_1_8 : MsanMetadataPtrForLoad_1_8;
852 switch (size) {
853 case 1:
854 return Fns[0];
855 case 2:
856 return Fns[1];
857 case 4:
858 return Fns[2];
859 case 8:
860 return Fns[3];
861 default:
862 return nullptr;
863 }
864}
865
866/// Module-level initialization.
867///
868/// inserts a call to __msan_init to the module's constructor list.
869void MemorySanitizer::initializeModule(Module &M) {
870 auto &DL = M.getDataLayout();
871
872 TargetTriple = M.getTargetTriple();
873
874 bool ShadowPassed = Opts.msan_shadow_base.has_value();
875 bool OriginPassed = Opts.msan_origin_base.has_value();
876 // Check the overrides first
877 if (ShadowPassed || OriginPassed) {
878 CustomMapParams.AndMask = Opts.msan_and_mask;
879 CustomMapParams.XorMask = Opts.msan_xor_mask;
880 CustomMapParams.ShadowBase = Opts.msan_shadow_base.value_or(0);
881 CustomMapParams.OriginBase = Opts.msan_origin_base.value_or(0);
882 MapParams = &CustomMapParams;
883 } else {
884 switch (TargetTriple.getOS()) {
885 case Triple::FreeBSD:
886 switch (TargetTriple.getArch()) {
887 case Triple::aarch64:
888 MapParams = FreeBSD_ARM_MemoryMapParams.bits64;
889 break;
890 case Triple::x86_64:
891 MapParams = FreeBSD_X86_MemoryMapParams.bits64;
892 break;
893 case Triple::x86:
894 MapParams = FreeBSD_X86_MemoryMapParams.bits32;
895 break;
896 default:
897 report_fatal_error("unsupported architecture");
898 }
899 break;
900 case Triple::NetBSD:
901 switch (TargetTriple.getArch()) {
902 case Triple::x86_64:
903 MapParams = NetBSD_X86_MemoryMapParams.bits64;
904 break;
905 default:
906 report_fatal_error("unsupported architecture");
907 }
908 break;
909 case Triple::Linux:
910 switch (TargetTriple.getArch()) {
911 case Triple::x86_64:
912 MapParams = Linux_X86_MemoryMapParams.bits64;
913 break;
914 case Triple::x86:
915 MapParams = Linux_X86_MemoryMapParams.bits32;
916 break;
917 case Triple::mips64:
918 case Triple::mips64el:
919 MapParams = Linux_MIPS_MemoryMapParams.bits64;
920 break;
921 case Triple::ppc64:
922 case Triple::ppc64le:
923 MapParams = Linux_PowerPC_MemoryMapParams.bits64;
924 break;
925 case Triple::systemz:
926 MapParams = Linux_S390_MemoryMapParams.bits64;
927 break;
928 case Triple::aarch64:
930 MapParams = Linux_ARM_MemoryMapParams.bits64;
931 break;
933 MapParams = Linux_LoongArch_MemoryMapParams.bits64;
934 break;
935 case Triple::hexagon:
936 MapParams = Linux_Hexagon_MemoryMapParams_P.bits32;
937 break;
938 default:
939 report_fatal_error("unsupported architecture");
940 }
941 break;
942 default:
943 report_fatal_error("unsupported operating system");
944 }
945 }
946
947 C = &(M.getContext());
948 IRBuilder<> IRB(M);
949 IntptrTy = IRB.getIntPtrTy(DL);
950 OriginTy = IRB.getInt32Ty();
951 PtrTy = IRB.getPtrTy();
952
953 ColdCallWeights = MDBuilder(*C).createUnlikelyBranchWeights();
954 OriginStoreWeights = MDBuilder(*C).createUnlikelyBranchWeights();
955
956 if (!CompileKernel) {
957 if (TrackOrigins)
958 M.getOrInsertGlobal("__msan_track_origins", IRB.getInt32Ty(), [&] {
959 return new GlobalVariable(
960 M, IRB.getInt32Ty(), true, GlobalValue::WeakODRLinkage,
961 IRB.getInt32(TrackOrigins), "__msan_track_origins");
962 });
963
964 if (Recover)
965 M.getOrInsertGlobal("__msan_keep_going", IRB.getInt32Ty(), [&] {
966 return new GlobalVariable(M, IRB.getInt32Ty(), true,
967 GlobalValue::WeakODRLinkage,
968 IRB.getInt32(Recover), "__msan_keep_going");
969 });
970 }
971}
972
973namespace {
974
975/// A helper class that handles instrumentation of VarArg
976/// functions on a particular platform.
977///
978/// Implementations are expected to insert the instrumentation
979/// necessary to propagate argument shadow through VarArg function
980/// calls. Visit* methods are called during an InstVisitor pass over
981/// the function, and should avoid creating new basic blocks. A new
982/// instance of this class is created for each instrumented function.
983struct VarArgHelper {
984 virtual ~VarArgHelper() = default;
985
986 /// Visit a CallBase.
987 virtual void visitCallBase(CallBase &CB, IRBuilder<> &IRB) = 0;
988
989 /// Visit a va_start call.
990 virtual void visitVAStartInst(VAStartInst &I) = 0;
991
992 /// Visit a va_copy call.
993 virtual void visitVACopyInst(VACopyInst &I) = 0;
994
995 /// Finalize function instrumentation.
996 ///
997 /// This method is called after visiting all interesting (see above)
998 /// instructions in a function.
999 virtual void finalizeInstrumentation() = 0;
1000};
1001
1002struct MemorySanitizerVisitor;
1003
1004} // end anonymous namespace
1005
1006static VarArgHelper *CreateVarArgHelper(Function &Func, MemorySanitizer &Msan,
1007 MemorySanitizerVisitor &Visitor);
1008
1009static unsigned TypeSizeToSizeIndex(TypeSize TS) {
1010 if (TS.isScalable())
1011 // Scalable types unconditionally take slowpaths.
1012 return kNumberOfAccessSizes;
1013 unsigned TypeSizeFixed = TS.getFixedValue();
1014 if (TypeSizeFixed <= 8)
1015 return 0;
1016 return Log2_32_Ceil((TypeSizeFixed + 7) / 8);
1017}
1018
1019namespace {
1020
1021/// Helper class to attach debug information of the given instruction onto new
1022/// instructions inserted after.
1023class NextNodeIRBuilder : public IRBuilder<> {
1024public:
1025 explicit NextNodeIRBuilder(Instruction *IP) : IRBuilder<>(IP->getNextNode()) {
1026 SetCurrentDebugLocation(IP->getDebugLoc());
1027 }
1028};
1029
1030/// This class does all the work for a given function. Store and Load
1031/// instructions store and load corresponding shadow and origin
1032/// values. Most instructions propagate shadow from arguments to their
1033/// return values. Certain instructions (most importantly, BranchInst)
1034/// test their argument shadow and print reports (with a runtime call) if it's
1035/// non-zero.
1036struct MemorySanitizerVisitor : public InstVisitor<MemorySanitizerVisitor> {
1037 const InstrumentationOptions &Opts;
1038 Function &F;
1039 MemorySanitizer &MS;
1040 SmallVector<PHINode *, 16> ShadowPHINodes, OriginPHINodes;
1041 ValueMap<Value *, Value *> ShadowMap, OriginMap;
1042 std::unique_ptr<VarArgHelper> VAHelper;
1043 const TargetLibraryInfo *TLI;
1044 Instruction *FnPrologueEnd;
1045 SmallVector<Instruction *, 16> Instructions;
1046
1047 // The following flags disable parts of MSan instrumentation based on
1048 // exclusion list contents and command-line options.
1049 bool InsertChecks;
1050 bool PropagateShadow;
1051 bool PoisonStack;
1052 bool PoisonUndef;
1053 bool PoisonUndefVectors;
1054
1055 struct ShadowOriginAndInsertPoint {
1056 Value *Shadow;
1057 Value *Origin;
1058 Instruction *OrigIns;
1059
1060 ShadowOriginAndInsertPoint(Value *S, Value *O, Instruction *I)
1061 : Shadow(S), Origin(O), OrigIns(I) {}
1062 };
1064 DenseMap<const DILocation *, int> LazyWarningDebugLocationCount;
1065 SmallSetVector<AllocaInst *, 16> AllocaSet;
1068 int64_t SplittableBlocksCount = 0;
1069
1070 MemorySanitizerVisitor(Function &F, MemorySanitizer &MS,
1071 const TargetLibraryInfo &TLI)
1072 : Opts(MS.Opts), F(F), MS(MS), VAHelper(CreateVarArgHelper(F, MS, *this)),
1073 TLI(&TLI) {
1074 bool SanitizeFunction = F.hasFnAttribute(Attribute::SanitizeMemory) &&
1075 !Opts.msan_disable_checks;
1076 InsertChecks = SanitizeFunction;
1077 PropagateShadow = SanitizeFunction;
1078 PoisonStack = SanitizeFunction && Opts.msan_poison_stack;
1079 PoisonUndef = SanitizeFunction && Opts.msan_poison_undef;
1080 PoisonUndefVectors = SanitizeFunction && Opts.msan_poison_undef_vectors;
1081
1082 // In the presence of unreachable blocks, we may see Phi nodes with
1083 // incoming nodes from such blocks. Since InstVisitor skips unreachable
1084 // blocks, such nodes will not have any shadow value associated with them.
1085 // It's easier to remove unreachable blocks than deal with missing shadow.
1087
1088 MS.initializeCallbacks(*F.getParent(), TLI);
1089 FnPrologueEnd =
1090 IRBuilder<>(F.getEntryBlock().getFirstNonPHIIt())
1091 .CreateIntrinsicWithoutFolding(Intrinsic::donothing, {});
1092
1093 if (MS.CompileKernel) {
1094 IRBuilder<> IRB(FnPrologueEnd);
1095 insertKmsanPrologue(IRB);
1096 }
1097
1098 LLVM_DEBUG(if (!InsertChecks) dbgs()
1099 << "MemorySanitizer is not inserting checks into '"
1100 << F.getName() << "'\n");
1101 }
1102
1103 bool instrumentWithCalls(Value *V) {
1104 // Constants likely will be eliminated by follow-up passes.
1105 if (isa<Constant>(V))
1106 return false;
1107 ++SplittableBlocksCount;
1108 return Opts.msan_instrumentation_with_call_threshold >= 0 &&
1109 SplittableBlocksCount >
1110 Opts.msan_instrumentation_with_call_threshold;
1111 }
1112
1113 bool isInPrologue(Instruction &I) {
1114 return I.getParent() == FnPrologueEnd->getParent() &&
1115 (&I == FnPrologueEnd || I.comesBefore(FnPrologueEnd));
1116 }
1117
1118 // Creates a new origin and records the stack trace. In general we can call
1119 // this function for any origin manipulation we like. However it will cost
1120 // runtime resources. So use this wisely only if it can provide additional
1121 // information helpful to a user.
1122 Value *updateOrigin(Value *V, IRBuilder<> &IRB) {
1123 if (MS.TrackOrigins <= 1)
1124 return V;
1125 return IRB.CreateCall(MS.MsanChainOriginFn, V);
1126 }
1127
1128 Value *originToIntptr(IRBuilder<> &IRB, Value *Origin) {
1129 const DataLayout &DL = F.getDataLayout();
1130 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
1131 if (IntptrSize == kOriginSize)
1132 return Origin;
1133 assert(IntptrSize == kOriginSize * 2);
1134 Origin = IRB.CreateIntCast(Origin, MS.IntptrTy, /* isSigned */ false);
1135 return IRB.CreateOr(Origin, IRB.CreateShl(Origin, kOriginSize * 8));
1136 }
1137
1138 /// Fill memory range with the given origin value.
1139 void paintOrigin(IRBuilder<> &IRB, Value *Origin, Value *OriginPtr,
1140 TypeSize TS, Align Alignment) {
1141 const DataLayout &DL = F.getDataLayout();
1142 const Align IntptrAlignment = DL.getABITypeAlign(MS.IntptrTy);
1143 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
1144 assert(IntptrAlignment >= kMinOriginAlignment);
1145 assert(IntptrSize >= kOriginSize);
1146
1147 // Note: The loop based formation works for fixed length vectors too,
1148 // however we prefer to unroll and specialize alignment below.
1149 if (TS.isScalable()) {
1150 Value *Size = IRB.CreateTypeSize(MS.IntptrTy, TS);
1151 Value *RoundUp =
1152 IRB.CreateAdd(Size, ConstantInt::get(MS.IntptrTy, kOriginSize - 1));
1153 Value *End =
1154 IRB.CreateUDiv(RoundUp, ConstantInt::get(MS.IntptrTy, kOriginSize));
1155 auto [InsertPt, Index] =
1157 IRB.SetInsertPoint(InsertPt);
1158
1159 Value *GEP = IRB.CreateGEP(MS.OriginTy, OriginPtr, Index);
1161 return;
1162 }
1163
1164 unsigned Size = TS.getFixedValue();
1165
1166 unsigned Ofs = 0;
1167 Align CurrentAlignment = Alignment;
1168 if (Alignment >= IntptrAlignment && IntptrSize > kOriginSize) {
1169 Value *IntptrOrigin = originToIntptr(IRB, Origin);
1170 Value *IntptrOriginPtr = IRB.CreatePointerCast(OriginPtr, MS.PtrTy);
1171 for (unsigned i = 0; i < Size / IntptrSize; ++i) {
1172 Value *Ptr = i ? IRB.CreateConstGEP1_32(MS.IntptrTy, IntptrOriginPtr, i)
1173 : IntptrOriginPtr;
1174 IRB.CreateAlignedStore(IntptrOrigin, Ptr, CurrentAlignment);
1175 Ofs += IntptrSize / kOriginSize;
1176 CurrentAlignment = IntptrAlignment;
1177 }
1178 }
1179
1180 for (unsigned i = Ofs; i < (Size + kOriginSize - 1) / kOriginSize; ++i) {
1181 Value *GEP =
1182 i ? IRB.CreateConstGEP1_32(MS.OriginTy, OriginPtr, i) : OriginPtr;
1183 IRB.CreateAlignedStore(Origin, GEP, CurrentAlignment);
1184 CurrentAlignment = kMinOriginAlignment;
1185 }
1186 }
1187
1188 void storeOrigin(IRBuilder<> &IRB, Value *Addr, Value *Shadow, Value *Origin,
1189 Value *OriginPtr, Align Alignment) {
1190 const DataLayout &DL = F.getDataLayout();
1191 const Align OriginAlignment = std::max(kMinOriginAlignment, Alignment);
1192 TypeSize StoreSize = DL.getTypeStoreSize(Shadow->getType());
1193 // ZExt cannot convert between vector and scalar
1194 Value *ConvertedShadow = convertShadowToScalar(Shadow, IRB);
1195 if (auto *ConstantShadow = dyn_cast<Constant>(ConvertedShadow)) {
1196 if (!Opts.msan_check_constant_shadow || ConstantShadow->isNullValue()) {
1197 // Origin is not needed: value is initialized or const shadow is
1198 // ignored.
1199 return;
1200 }
1201 if (llvm::isKnownNonZero(ConvertedShadow, DL)) {
1202 // Copy origin as the value is definitely uninitialized.
1203 paintOrigin(IRB, updateOrigin(Origin, IRB), OriginPtr, StoreSize,
1204 OriginAlignment);
1205 return;
1206 }
1207 // Fallback to runtime check, which still can be optimized out later.
1208 }
1209
1210 TypeSize TypeSizeInBits = DL.getTypeSizeInBits(ConvertedShadow->getType());
1211 unsigned SizeIndex = TypeSizeToSizeIndex(TypeSizeInBits);
1212 if (instrumentWithCalls(ConvertedShadow) &&
1213 SizeIndex < kNumberOfAccessSizes && !MS.CompileKernel) {
1214 FunctionCallee Fn = MS.MaybeStoreOriginFn[SizeIndex];
1215 Value *ConvertedShadow2 =
1216 IRB.CreateZExt(ConvertedShadow, IRB.getIntNTy(8 * (1 << SizeIndex)));
1217 CallBase *CB = IRB.CreateCall(Fn, {ConvertedShadow2, Addr, Origin});
1218 CB->addParamAttr(0, Attribute::ZExt);
1219 CB->addParamAttr(2, Attribute::ZExt);
1220 } else {
1221 Value *Cmp = convertToBool(ConvertedShadow, IRB, "_mscmp");
1223 Cmp, &*IRB.GetInsertPoint(), false, MS.OriginStoreWeights);
1224 IRBuilder<> IRBNew(CheckTerm);
1225 paintOrigin(IRBNew, updateOrigin(Origin, IRBNew), OriginPtr, StoreSize,
1226 OriginAlignment);
1227 }
1228 }
1229
1230 void materializeStores() {
1231 for (StoreInst *SI : StoreList) {
1232 IRBuilder<> IRB(SI);
1233 Value *Val = SI->getValueOperand();
1234 Value *Addr = SI->getPointerOperand();
1235 Value *Shadow = SI->isAtomic() ? getCleanShadow(Val) : getShadow(Val);
1236 Value *ShadowPtr, *OriginPtr;
1237 Type *ShadowTy = Shadow->getType();
1238 const Align Alignment = SI->getAlign();
1239 const Align OriginAlignment = std::max(kMinOriginAlignment, Alignment);
1240 std::tie(ShadowPtr, OriginPtr) =
1241 getShadowOriginPtr(Addr, IRB, ShadowTy, Alignment, /*isStore*/ true);
1242
1243 [[maybe_unused]] StoreInst *NewSI =
1244 IRB.CreateAlignedStore(Shadow, ShadowPtr, Alignment);
1245 LLVM_DEBUG(dbgs() << " STORE: " << *NewSI << "\n");
1246
1247 if (SI->isAtomic())
1248 SI->setOrdering(addReleaseOrdering(SI->getOrdering()));
1249
1250 if (MS.TrackOrigins && !SI->isAtomic())
1251 storeOrigin(IRB, Addr, Shadow, getOrigin(Val), OriginPtr,
1252 OriginAlignment);
1253 }
1254 }
1255
1256 // Returns true if Debug Location corresponds to multiple warnings.
1257 bool shouldDisambiguateWarningLocation(const DebugLoc &DebugLoc) {
1258 if (MS.TrackOrigins < 2)
1259 return false;
1260
1261 if (LazyWarningDebugLocationCount.empty())
1262 for (const auto &I : InstrumentationList)
1263 ++LazyWarningDebugLocationCount[I.OrigIns->getDebugLoc()];
1264
1265 return LazyWarningDebugLocationCount[DebugLoc] >=
1266 Opts.msan_disambiguate_warning_threshold;
1267 }
1268
1269 /// Helper function to insert a warning at IRB's current insert point.
1270 void insertWarningFn(IRBuilder<> &IRB, Value *Origin) {
1271 if (!Origin)
1272 Origin = (Value *)IRB.getInt32(0);
1273 assert(Origin->getType()->isIntegerTy());
1274
1275 if (shouldDisambiguateWarningLocation(IRB.getCurrentDebugLocation())) {
1276 // Try to create additional origin with debug info of the last origin
1277 // instruction. It may provide additional information to the user.
1278 if (Instruction *OI = dyn_cast_or_null<Instruction>(Origin)) {
1279 assert(MS.TrackOrigins);
1280 auto NewDebugLoc = OI->getDebugLoc();
1281 // Origin update with missing or the same debug location provides no
1282 // additional value.
1283 if (NewDebugLoc && NewDebugLoc != IRB.getCurrentDebugLocation()) {
1284 // Insert update just before the check, so we call runtime only just
1285 // before the report.
1286 IRBuilder<> IRBOrigin(&*IRB.GetInsertPoint());
1287 IRBOrigin.SetCurrentDebugLocation(NewDebugLoc);
1288 Origin = updateOrigin(Origin, IRBOrigin);
1289 }
1290 }
1291 }
1292
1293 if (MS.CompileKernel || MS.TrackOrigins)
1294 IRB.CreateCall(MS.WarningFn, Origin)->setCannotMerge();
1295 else
1296 IRB.CreateCall(MS.WarningFn)->setCannotMerge();
1297 // FIXME: Insert UnreachableInst if !MS.Recover?
1298 // This may invalidate some of the following checks and needs to be done
1299 // at the very end.
1300 }
1301
1302 void materializeOneCheck(IRBuilder<> &IRB, Value *ConvertedShadow,
1303 Value *Origin) {
1304 const DataLayout &DL = F.getDataLayout();
1305 TypeSize TypeSizeInBits = DL.getTypeSizeInBits(ConvertedShadow->getType());
1306 unsigned SizeIndex = TypeSizeToSizeIndex(TypeSizeInBits);
1307 if (instrumentWithCalls(ConvertedShadow) && !MS.CompileKernel) {
1308 // ZExt cannot convert between vector and scalar
1309 ConvertedShadow = convertShadowToScalar(ConvertedShadow, IRB);
1310 Value *ConvertedShadow2 =
1311 IRB.CreateZExt(ConvertedShadow, IRB.getIntNTy(8 * (1 << SizeIndex)));
1312
1313 if (SizeIndex < kNumberOfAccessSizes) {
1314 FunctionCallee Fn = MS.MaybeWarningFn[SizeIndex];
1315 CallBase *CB = IRB.CreateCall(
1316 Fn,
1317 {ConvertedShadow2,
1318 MS.TrackOrigins && Origin ? Origin : (Value *)IRB.getInt32(0)});
1319 CB->addParamAttr(0, Attribute::ZExt);
1320 CB->addParamAttr(1, Attribute::ZExt);
1321 } else {
1322 FunctionCallee Fn = MS.MaybeWarningVarSizeFn;
1323 Value *ShadowAlloca = IRB.CreateAlloca(ConvertedShadow2->getType(), 0u);
1324 IRB.CreateStore(ConvertedShadow2, ShadowAlloca);
1325 unsigned ShadowSize = DL.getTypeAllocSize(ConvertedShadow2->getType());
1326 CallBase *CB = IRB.CreateCall(
1327 Fn,
1328 {ShadowAlloca, ConstantInt::get(IRB.getInt64Ty(), ShadowSize),
1329 MS.TrackOrigins && Origin ? Origin : (Value *)IRB.getInt32(0)});
1330 CB->addParamAttr(1, Attribute::ZExt);
1331 CB->addParamAttr(2, Attribute::ZExt);
1332 }
1333 } else {
1334 Value *Cmp = convertToBool(ConvertedShadow, IRB, "_mscmp");
1336 Cmp, &*IRB.GetInsertPoint(),
1337 /* Unreachable */ !MS.Recover, MS.ColdCallWeights);
1338
1339 IRB.SetInsertPoint(CheckTerm);
1340 insertWarningFn(IRB, Origin);
1341 LLVM_DEBUG(dbgs() << " CHECK: " << *Cmp << "\n");
1342 }
1343 }
1344
1345 void materializeInstructionChecks(
1346 ArrayRef<ShadowOriginAndInsertPoint> InstructionChecks) {
1347 const DataLayout &DL = F.getDataLayout();
1348 // Disable combining in some cases. TrackOrigins checks each shadow to pick
1349 // correct origin.
1350 bool Combine = !MS.TrackOrigins;
1351 Instruction *Instruction = InstructionChecks.front().OrigIns;
1352 Value *Shadow = nullptr;
1353 for (const auto &ShadowData : InstructionChecks) {
1354 assert(ShadowData.OrigIns == Instruction);
1355 IRBuilder<> IRB(Instruction);
1356
1357 Value *ConvertedShadow = ShadowData.Shadow;
1358
1359 if (auto *ConstantShadow = dyn_cast<Constant>(ConvertedShadow)) {
1360 if (!Opts.msan_check_constant_shadow || ConstantShadow->isNullValue()) {
1361 // Skip, value is initialized or const shadow is ignored.
1362 continue;
1363 }
1364 if (llvm::isKnownNonZero(ConvertedShadow, DL)) {
1365 // Report as the value is definitely uninitialized.
1366 insertWarningFn(IRB, ShadowData.Origin);
1367 if (!MS.Recover)
1368 return; // Always fail and stop here, not need to check the rest.
1369 // Skip entire instruction,
1370 continue;
1371 }
1372 // Fallback to runtime check, which still can be optimized out later.
1373 }
1374
1375 if (!Combine) {
1376 materializeOneCheck(IRB, ConvertedShadow, ShadowData.Origin);
1377 continue;
1378 }
1379
1380 if (!Shadow) {
1381 Shadow = ConvertedShadow;
1382 continue;
1383 }
1384
1385 Shadow = convertToBool(Shadow, IRB, "_mscmp");
1386 ConvertedShadow = convertToBool(ConvertedShadow, IRB, "_mscmp");
1387 Shadow = IRB.CreateOr(Shadow, ConvertedShadow, "_msor");
1388 }
1389
1390 if (Shadow) {
1391 assert(Combine);
1392 IRBuilder<> IRB(Instruction);
1393 materializeOneCheck(IRB, Shadow, nullptr);
1394 }
1395 }
1396
1397 static bool isAArch64SVCount(Type *Ty) {
1398 if (TargetExtType *TTy = dyn_cast<TargetExtType>(Ty))
1399 return TTy->getName() == "aarch64.svcount";
1400 return false;
1401 }
1402
1403 // This is intended to match the "AArch64 Predicate-as-Counter Type" (aka
1404 // 'target("aarch64.svcount")', but not e.g., <vscale x 4 x i32>.
1405 static bool isScalableNonVectorType(Type *Ty) {
1406 if (!isAArch64SVCount(Ty))
1407 LLVM_DEBUG(dbgs() << "isScalableNonVectorType: Unexpected type " << *Ty
1408 << "\n");
1409
1410 return Ty->isScalableTy() && !isa<VectorType>(Ty);
1411 }
1412
1413 void materializeChecks() {
1414#ifndef NDEBUG
1415 // For assert below.
1416 SmallPtrSet<Instruction *, 16> Done;
1417#endif
1418
1419 for (auto I = InstrumentationList.begin();
1420 I != InstrumentationList.end();) {
1421 auto OrigIns = I->OrigIns;
1422 // Checks are grouped by the original instruction. We call all
1423 // `insertShadowCheck` for an instruction at once.
1424 assert(Done.insert(OrigIns).second);
1425 auto J = std::find_if(I + 1, InstrumentationList.end(),
1426 [OrigIns](const ShadowOriginAndInsertPoint &R) {
1427 return OrigIns != R.OrigIns;
1428 });
1429 // Process all checks of instruction at once.
1430 materializeInstructionChecks(ArrayRef<ShadowOriginAndInsertPoint>(I, J));
1431 I = J;
1432 }
1433
1434 LLVM_DEBUG(dbgs() << "DONE:\n" << F);
1435 }
1436
1437 // Returns the last instruction in the new prologue
1438 void insertKmsanPrologue(IRBuilder<> &IRB) {
1439 Value *ContextState = IRB.CreateCall(MS.MsanGetContextStateFn, {});
1440 Constant *Zero = IRB.getInt32(0);
1441 MS.ParamTLS = IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1442 {Zero, IRB.getInt32(0)}, "param_shadow");
1443 MS.RetvalTLS = IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1444 {Zero, IRB.getInt32(1)}, "retval_shadow");
1445 MS.VAArgTLS = IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1446 {Zero, IRB.getInt32(2)}, "va_arg_shadow");
1447 MS.VAArgOriginTLS = IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1448 {Zero, IRB.getInt32(3)}, "va_arg_origin");
1449 MS.VAArgOverflowSizeTLS =
1450 IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1451 {Zero, IRB.getInt32(4)}, "va_arg_overflow_size");
1452 MS.ParamOriginTLS = IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1453 {Zero, IRB.getInt32(5)}, "param_origin");
1454 MS.RetvalOriginTLS =
1455 IRB.CreateGEP(MS.MsanContextStateTy, ContextState,
1456 {Zero, IRB.getInt32(6)}, "retval_origin");
1457 if (MS.TargetTriple.getArch() == Triple::systemz)
1458 MS.MsanMetadataAlloca = IRB.CreateAlloca(MS.MsanMetadata, 0u);
1459 }
1460
1461 /// Add MemorySanitizer instrumentation to a function.
1462 bool runOnFunction() {
1463 // Iterate all BBs in depth-first order and create shadow instructions
1464 // for all instructions (where applicable).
1465 // For PHI nodes we create dummy shadow PHIs which will be finalized later.
1466 for (BasicBlock *BB : depth_first(FnPrologueEnd->getParent()))
1467 visit(*BB);
1468
1469 // `visit` above only collects instructions. Process them after iterating
1470 // CFG to avoid requirement on CFG transformations.
1471 for (Instruction *I : Instructions)
1473
1474 // Finalize PHI nodes.
1475 for (PHINode *PN : ShadowPHINodes) {
1476 PHINode *PNS = cast<PHINode>(getShadow(PN));
1477 PHINode *PNO = MS.TrackOrigins ? cast<PHINode>(getOrigin(PN)) : nullptr;
1478 size_t NumValues = PN->getNumIncomingValues();
1479 for (size_t v = 0; v < NumValues; v++) {
1480 PNS->addIncoming(getShadow(PN, v), PN->getIncomingBlock(v));
1481 if (PNO)
1482 PNO->addIncoming(getOrigin(PN, v), PN->getIncomingBlock(v));
1483 }
1484 }
1485
1486 VAHelper->finalizeInstrumentation();
1487
1488 // Poison llvm.lifetime.start intrinsics, if we haven't fallen back to
1489 // instrumenting only allocas.
1490 if (Opts.msan_handle_lifetime_intrinsics) {
1491 for (auto Item : LifetimeStartList) {
1492 instrumentAlloca(*Item.second, Item.first);
1493 AllocaSet.remove(Item.second);
1494 }
1495 }
1496 // Poison the allocas for which we didn't instrument the corresponding
1497 // lifetime intrinsics.
1498 for (AllocaInst *AI : AllocaSet)
1499 instrumentAlloca(*AI);
1500
1501 // Insert shadow value checks.
1502 materializeChecks();
1503
1504 // Delayed instrumentation of StoreInst.
1505 // This may not add new address checks.
1506 materializeStores();
1507
1508 return true;
1509 }
1510
1511 /// Compute the shadow type that corresponds to a given Value.
1512 Type *getShadowTy(Value *V) { return getShadowTy(V->getType()); }
1513
1514 /// Compute the shadow type that corresponds to a given Type.
1515 Type *getShadowTy(Type *OrigTy) {
1516 if (!OrigTy->isSized()) {
1517 return nullptr;
1518 }
1519 // For integer type, shadow is the same as the original type.
1520 // This may return weird-sized types like i1.
1521 if (IntegerType *IT = dyn_cast<IntegerType>(OrigTy))
1522 return IT;
1523 const DataLayout &DL = F.getDataLayout();
1524 if (VectorType *VT = dyn_cast<VectorType>(OrigTy)) {
1525 uint32_t EltSize = DL.getTypeSizeInBits(VT->getElementType());
1526 return VectorType::get(IntegerType::get(*MS.C, EltSize),
1527 VT->getElementCount());
1528 }
1529 if (ArrayType *AT = dyn_cast<ArrayType>(OrigTy)) {
1530 return ArrayType::get(getShadowTy(AT->getElementType()),
1531 AT->getNumElements());
1532 }
1533 if (StructType *ST = dyn_cast<StructType>(OrigTy)) {
1535 for (unsigned i = 0, n = ST->getNumElements(); i < n; i++)
1536 Elements.push_back(getShadowTy(ST->getElementType(i)));
1537 StructType *Res = StructType::get(*MS.C, Elements, ST->isPacked());
1538 LLVM_DEBUG(dbgs() << "getShadowTy: " << *ST << " ===> " << *Res << "\n");
1539 return Res;
1540 }
1541 if (isScalableNonVectorType(OrigTy)) {
1542 LLVM_DEBUG(dbgs() << "getShadowTy: Scalable non-vector type: " << *OrigTy
1543 << "\n");
1544 return OrigTy;
1545 }
1546
1547 uint32_t TypeSize = DL.getTypeSizeInBits(OrigTy);
1548 return IntegerType::get(*MS.C, TypeSize);
1549 }
1550
1551 /// Extract combined shadow of struct elements as a bool
1552 Value *collapseStructShadow(StructType *Struct, Value *Shadow,
1553 IRBuilder<> &IRB) {
1554 Value *FalseVal = IRB.getIntN(/* width */ 1, /* value */ 0);
1555 Value *Aggregator = FalseVal;
1556
1557 for (unsigned Idx = 0; Idx < Struct->getNumElements(); Idx++) {
1558 // Combine by ORing together each element's bool shadow
1559 Value *ShadowItem = IRB.CreateExtractValue(Shadow, Idx);
1560 Value *ShadowBool = convertToBool(ShadowItem, IRB);
1561
1562 if (Aggregator != FalseVal)
1563 Aggregator = IRB.CreateOr(Aggregator, ShadowBool);
1564 else
1565 Aggregator = ShadowBool;
1566 }
1567
1568 return Aggregator;
1569 }
1570
1571 // Extract combined shadow of array elements
1572 Value *collapseArrayShadow(ArrayType *Array, Value *Shadow,
1573 IRBuilder<> &IRB) {
1574 if (!Array->getNumElements())
1575 return IRB.getIntN(/* width */ 1, /* value */ 0);
1576
1577 Value *FirstItem = IRB.CreateExtractValue(Shadow, 0);
1578 Value *Aggregator = convertShadowToScalar(FirstItem, IRB);
1579
1580 for (unsigned Idx = 1; Idx < Array->getNumElements(); Idx++) {
1581 Value *ShadowItem = IRB.CreateExtractValue(Shadow, Idx);
1582 Value *ShadowInner = convertShadowToScalar(ShadowItem, IRB);
1583 Aggregator = IRB.CreateOr(Aggregator, ShadowInner);
1584 }
1585 return Aggregator;
1586 }
1587
1588 /// Convert a shadow value to it's flattened variant. The resulting
1589 /// shadow may not necessarily have the same bit width as the input
1590 /// value, but it will always be comparable to zero.
1591 Value *convertShadowToScalar(Value *V, IRBuilder<> &IRB) {
1592 if (StructType *Struct = dyn_cast<StructType>(V->getType()))
1593 return collapseStructShadow(Struct, V, IRB);
1594 if (ArrayType *Array = dyn_cast<ArrayType>(V->getType()))
1595 return collapseArrayShadow(Array, V, IRB);
1596 if (isa<VectorType>(V->getType())) {
1597 if (isa<ScalableVectorType>(V->getType()))
1598 return convertShadowToScalar(IRB.CreateOrReduce(V), IRB);
1599 unsigned BitWidth =
1600 V->getType()->getPrimitiveSizeInBits().getFixedValue();
1601 return IRB.CreateBitCast(V, IntegerType::get(*MS.C, BitWidth));
1602 }
1603 return V;
1604 }
1605
1606 // Convert a scalar value to an i1 by comparing with 0
1607 Value *convertToBool(Value *V, IRBuilder<> &IRB, const Twine &name = "") {
1608 Type *VTy = V->getType();
1609 if (!VTy->isIntegerTy())
1610 return convertToBool(convertShadowToScalar(V, IRB), IRB, name);
1611 if (VTy->getIntegerBitWidth() == 1)
1612 // Just converting a bool to a bool, so do nothing.
1613 return V;
1614 return IRB.CreateICmpNE(V, ConstantInt::get(VTy, 0), name);
1615 }
1616
1617 Type *ptrToIntPtrType(Type *PtrTy) const {
1618 if (VectorType *VectTy = dyn_cast<VectorType>(PtrTy)) {
1619 return VectorType::get(ptrToIntPtrType(VectTy->getElementType()),
1620 VectTy->getElementCount());
1621 }
1622 assert(PtrTy->isIntOrPtrTy());
1623 return MS.IntptrTy;
1624 }
1625
1626 Type *getPtrToShadowPtrType(Type *IntPtrTy, Type *ShadowTy) const {
1627 if (VectorType *VectTy = dyn_cast<VectorType>(IntPtrTy)) {
1628 return VectorType::get(
1629 getPtrToShadowPtrType(VectTy->getElementType(), ShadowTy),
1630 VectTy->getElementCount());
1631 }
1632 assert(IntPtrTy == MS.IntptrTy);
1633 return MS.PtrTy;
1634 }
1635
1636 Constant *constToIntPtr(Type *IntPtrTy, uint64_t C) const {
1637 if (VectorType *VectTy = dyn_cast<VectorType>(IntPtrTy)) {
1639 VectTy->getElementCount(),
1640 constToIntPtr(VectTy->getElementType(), C));
1641 }
1642 assert(IntPtrTy == MS.IntptrTy);
1643 // TODO: Avoid implicit trunc?
1644 // See https://github.com/llvm/llvm-project/issues/112510.
1645 return ConstantInt::get(MS.IntptrTy, C, /*IsSigned=*/false,
1646 /*ImplicitTrunc=*/true);
1647 }
1648
1649 /// Returns the integer shadow offset that corresponds to a given
1650 /// application address, whereby:
1651 ///
1652 /// Offset = (Addr & ~AndMask) ^ XorMask
1653 /// Shadow = ShadowBase + Offset
1654 /// Origin = (OriginBase + Offset) & ~Alignment
1655 ///
1656 /// Note: for efficiency, many shadow mappings only require use the XorMask
1657 /// and OriginBase; the AndMask and ShadowBase are often zero.
1658 Value *getShadowPtrOffset(Value *Addr, IRBuilder<> &IRB) {
1659 Type *IntptrTy = ptrToIntPtrType(Addr->getType());
1660 Value *OffsetLong = IRB.CreatePointerCast(Addr, IntptrTy);
1661
1662 if (uint64_t AndMask = MS.MapParams->AndMask)
1663 OffsetLong = IRB.CreateAnd(OffsetLong, constToIntPtr(IntptrTy, ~AndMask));
1664
1665 if (uint64_t XorMask = MS.MapParams->XorMask)
1666 OffsetLong = IRB.CreateXor(OffsetLong, constToIntPtr(IntptrTy, XorMask));
1667 return OffsetLong;
1668 }
1669
1670 /// Compute the shadow and origin addresses corresponding to a given
1671 /// application address.
1672 ///
1673 /// Shadow = ShadowBase + Offset
1674 /// Origin = (OriginBase + Offset) & ~3ULL
1675 /// Addr can be a ptr or <N x ptr>. In both cases ShadowTy the shadow type of
1676 /// a single pointee.
1677 /// Returns <shadow_ptr, origin_ptr> or <<N x shadow_ptr>, <N x origin_ptr>>.
1678 std::pair<Value *, Value *>
1679 getShadowOriginPtrUserspace(Value *Addr, IRBuilder<> &IRB, Type *ShadowTy,
1680 MaybeAlign Alignment) {
1681 VectorType *VectTy = dyn_cast<VectorType>(Addr->getType());
1682 if (!VectTy) {
1683 assert(Addr->getType()->isPointerTy());
1684 } else {
1685 assert(VectTy->getElementType()->isPointerTy());
1686 }
1687 Type *IntptrTy = ptrToIntPtrType(Addr->getType());
1688 Value *ShadowOffset = getShadowPtrOffset(Addr, IRB);
1689 Value *ShadowLong = ShadowOffset;
1690 if (uint64_t ShadowBase = MS.MapParams->ShadowBase) {
1691 ShadowLong =
1692 IRB.CreateAdd(ShadowLong, constToIntPtr(IntptrTy, ShadowBase));
1693 }
1694 Value *ShadowPtr = IRB.CreateIntToPtr(
1695 ShadowLong, getPtrToShadowPtrType(IntptrTy, ShadowTy));
1696
1697 Value *OriginPtr = nullptr;
1698 if (MS.TrackOrigins) {
1699 Value *OriginLong = ShadowOffset;
1700 uint64_t OriginBase = MS.MapParams->OriginBase;
1701 if (OriginBase != 0)
1702 OriginLong =
1703 IRB.CreateAdd(OriginLong, constToIntPtr(IntptrTy, OriginBase));
1704 if (!Alignment || *Alignment < kMinOriginAlignment) {
1706 OriginLong = IRB.CreateAnd(OriginLong, constToIntPtr(IntptrTy, ~Mask));
1707 }
1708 OriginPtr = IRB.CreateIntToPtr(
1709 OriginLong, getPtrToShadowPtrType(IntptrTy, MS.OriginTy));
1710 }
1711 return std::make_pair(ShadowPtr, OriginPtr);
1712 }
1713
1714 template <typename... ArgsTy>
1715 Value *createMetadataCall(IRBuilder<> &IRB, FunctionCallee Callee,
1716 ArgsTy... Args) {
1717 if (MS.TargetTriple.getArch() == Triple::systemz) {
1718 IRB.CreateCall(Callee,
1719 {MS.MsanMetadataAlloca, std::forward<ArgsTy>(Args)...});
1720 return IRB.CreateLoad(MS.MsanMetadata, MS.MsanMetadataAlloca);
1721 }
1722
1723 return IRB.CreateCall(Callee, {std::forward<ArgsTy>(Args)...});
1724 }
1725
1726 std::pair<Value *, Value *> getShadowOriginPtrKernelNoVec(Value *Addr,
1727 IRBuilder<> &IRB,
1728 Type *ShadowTy,
1729 bool isStore) {
1730 Value *ShadowOriginPtrs;
1731 const DataLayout &DL = F.getDataLayout();
1732 TypeSize Size = DL.getTypeStoreSize(ShadowTy);
1733
1734 FunctionCallee Getter = MS.getKmsanShadowOriginAccessFn(isStore, Size);
1735 Value *AddrCast = IRB.CreatePointerCast(Addr, MS.PtrTy);
1736 if (Getter) {
1737 ShadowOriginPtrs = createMetadataCall(IRB, Getter, AddrCast);
1738 } else {
1739 Value *SizeVal = ConstantInt::get(MS.IntptrTy, Size);
1740 ShadowOriginPtrs = createMetadataCall(
1741 IRB,
1742 isStore ? MS.MsanMetadataPtrForStoreN : MS.MsanMetadataPtrForLoadN,
1743 AddrCast, SizeVal);
1744 }
1745 Value *ShadowPtr = IRB.CreateExtractValue(ShadowOriginPtrs, 0);
1746 ShadowPtr = IRB.CreatePointerCast(ShadowPtr, MS.PtrTy);
1747 Value *OriginPtr = IRB.CreateExtractValue(ShadowOriginPtrs, 1);
1748
1749 return std::make_pair(ShadowPtr, OriginPtr);
1750 }
1751
1752 /// Addr can be a ptr or <N x ptr>. In both cases ShadowTy the shadow type of
1753 /// a single pointee.
1754 /// Returns <shadow_ptr, origin_ptr> or <<N x shadow_ptr>, <N x origin_ptr>>.
1755 std::pair<Value *, Value *> getShadowOriginPtrKernel(Value *Addr,
1756 IRBuilder<> &IRB,
1757 Type *ShadowTy,
1758 bool isStore) {
1759 VectorType *VectTy = dyn_cast<VectorType>(Addr->getType());
1760 if (!VectTy) {
1761 assert(Addr->getType()->isPointerTy());
1762 return getShadowOriginPtrKernelNoVec(Addr, IRB, ShadowTy, isStore);
1763 }
1764
1765 // TODO: Support callbacs with vectors of addresses.
1766 unsigned NumElements = cast<FixedVectorType>(VectTy)->getNumElements();
1767 Value *ShadowPtrs = ConstantInt::getNullValue(
1768 FixedVectorType::get(IRB.getPtrTy(), NumElements));
1769 Value *OriginPtrs = nullptr;
1770 if (MS.TrackOrigins)
1771 OriginPtrs = ConstantInt::getNullValue(
1772 FixedVectorType::get(IRB.getPtrTy(), NumElements));
1773 for (unsigned i = 0; i < NumElements; ++i) {
1774 Value *OneAddr =
1775 IRB.CreateExtractElement(Addr, ConstantInt::get(IRB.getInt32Ty(), i));
1776 auto [ShadowPtr, OriginPtr] =
1777 getShadowOriginPtrKernelNoVec(OneAddr, IRB, ShadowTy, isStore);
1778
1779 ShadowPtrs = IRB.CreateInsertElement(
1780 ShadowPtrs, ShadowPtr, ConstantInt::get(IRB.getInt32Ty(), i));
1781 if (MS.TrackOrigins)
1782 OriginPtrs = IRB.CreateInsertElement(
1783 OriginPtrs, OriginPtr, ConstantInt::get(IRB.getInt32Ty(), i));
1784 }
1785 return {ShadowPtrs, OriginPtrs};
1786 }
1787
1788 std::pair<Value *, Value *> getShadowOriginPtr(Value *Addr, IRBuilder<> &IRB,
1789 Type *ShadowTy,
1790 MaybeAlign Alignment,
1791 bool isStore) {
1792 if (MS.CompileKernel)
1793 return getShadowOriginPtrKernel(Addr, IRB, ShadowTy, isStore);
1794 return getShadowOriginPtrUserspace(Addr, IRB, ShadowTy, Alignment);
1795 }
1796
1797 /// Compute the shadow address for a given function argument.
1798 ///
1799 /// Shadow = ParamTLS+ArgOffset.
1800 Value *getShadowPtrForArgument(IRBuilder<> &IRB, int ArgOffset) {
1801 return IRB.CreatePtrAdd(MS.ParamTLS,
1802 ConstantInt::get(MS.IntptrTy, ArgOffset), "_msarg");
1803 }
1804
1805 /// Compute the origin address for a given function argument.
1806 Value *getOriginPtrForArgument(IRBuilder<> &IRB, int ArgOffset) {
1807 if (!MS.TrackOrigins)
1808 return nullptr;
1809 return IRB.CreatePtrAdd(MS.ParamOriginTLS,
1810 ConstantInt::get(MS.IntptrTy, ArgOffset),
1811 "_msarg_o");
1812 }
1813
1814 /// Compute the shadow address for a retval.
1815 Value *getShadowPtrForRetval(IRBuilder<> &IRB) {
1816 return IRB.CreatePointerCast(MS.RetvalTLS, IRB.getPtrTy(0), "_msret");
1817 }
1818
1819 /// Compute the origin address for a retval.
1820 Value *getOriginPtrForRetval() {
1821 // We keep a single origin for the entire retval. Might be too optimistic.
1822 return MS.RetvalOriginTLS;
1823 }
1824
1825 /// Set SV to be the shadow value for V.
1826 void setShadow(Value *V, Value *SV) {
1827 assert(!ShadowMap.count(V) && "Values may only have one shadow");
1828 ShadowMap[V] = PropagateShadow ? SV : getCleanShadow(V);
1829 }
1830
1831 /// Set Origin to be the origin value for V.
1832 void setOrigin(Value *V, Value *Origin) {
1833 if (!MS.TrackOrigins)
1834 return;
1835 assert(!OriginMap.count(V) && "Values may only have one origin");
1836 LLVM_DEBUG(dbgs() << "ORIGIN: " << *V << " ==> " << *Origin << "\n");
1837 OriginMap[V] = Origin;
1838 }
1839
1840 Constant *getCleanShadow(Type *OrigTy) {
1841 Type *ShadowTy = getShadowTy(OrigTy);
1842 if (!ShadowTy)
1843 return nullptr;
1844 return Constant::getNullValue(ShadowTy);
1845 }
1846
1847 /// Create a clean shadow value for a given value.
1848 ///
1849 /// Clean shadow (all zeroes) means all bits of the value are defined
1850 /// (initialized).
1851 Constant *getCleanShadow(Value *V) { return getCleanShadow(V->getType()); }
1852
1853 /// Create a dirty shadow of a given shadow type.
1854 Constant *getPoisonedShadow(Type *ShadowTy) {
1855 assert(ShadowTy);
1856 if (isa<IntegerType>(ShadowTy) || isa<VectorType>(ShadowTy))
1857 return Constant::getAllOnesValue(ShadowTy);
1858 if (ArrayType *AT = dyn_cast<ArrayType>(ShadowTy)) {
1859 SmallVector<Constant *, 4> Vals(AT->getNumElements(),
1860 getPoisonedShadow(AT->getElementType()));
1861 return ConstantArray::get(AT, Vals);
1862 }
1863 if (StructType *ST = dyn_cast<StructType>(ShadowTy)) {
1864 SmallVector<Constant *, 4> Vals;
1865 for (unsigned i = 0, n = ST->getNumElements(); i < n; i++)
1866 Vals.push_back(getPoisonedShadow(ST->getElementType(i)));
1867 return ConstantStruct::get(ST, Vals);
1868 }
1869 llvm_unreachable("Unexpected shadow type");
1870 }
1871
1872 /// Create a dirty shadow for a given value.
1873 Constant *getPoisonedShadow(Value *V) {
1874 Type *ShadowTy = getShadowTy(V);
1875 if (!ShadowTy)
1876 return nullptr;
1877 return getPoisonedShadow(ShadowTy);
1878 }
1879
1880 /// Create a clean (zero) origin.
1881 Value *getCleanOrigin() { return Constant::getNullValue(MS.OriginTy); }
1882
1883 /// Get the shadow value for a given Value.
1884 ///
1885 /// This function either returns the value set earlier with setShadow,
1886 /// or extracts if from ParamTLS (for function arguments).
1887 Value *getShadow(Value *V) {
1888 if (Instruction *I = dyn_cast<Instruction>(V)) {
1889 if (!PropagateShadow || I->getMetadata(LLVMContext::MD_nosanitize))
1890 return getCleanShadow(V);
1891 // For instructions the shadow is already stored in the map.
1892 Value *Shadow = ShadowMap[V];
1893 if (!Shadow) {
1894 LLVM_DEBUG(dbgs() << "No shadow: " << *V << "\n" << *(I->getParent()));
1895 assert(Shadow && "No shadow for a value");
1896 }
1897 return Shadow;
1898 }
1899 // Handle fully undefined values
1900 // (partially undefined constant vectors are handled later)
1901 if ([[maybe_unused]] UndefValue *U = dyn_cast<UndefValue>(V)) {
1902 Value *AllOnes = (PropagateShadow && PoisonUndef) ? getPoisonedShadow(V)
1903 : getCleanShadow(V);
1904 LLVM_DEBUG(dbgs() << "Undef: " << *U << " ==> " << *AllOnes << "\n");
1905 return AllOnes;
1906 }
1907 if (Argument *A = dyn_cast<Argument>(V)) {
1908 // For arguments we compute the shadow on demand and store it in the map.
1909 Value *&ShadowPtr = ShadowMap[V];
1910 if (ShadowPtr)
1911 return ShadowPtr;
1912 Function *F = A->getParent();
1913 IRBuilder<> EntryIRB(FnPrologueEnd);
1914 unsigned ArgOffset = 0;
1915 const DataLayout &DL = F->getDataLayout();
1916 for (auto &FArg : F->args()) {
1917 if (!FArg.getType()->isSized() || FArg.getType()->isScalableTy()) {
1918 LLVM_DEBUG(dbgs() << (FArg.getType()->isScalableTy()
1919 ? "vscale not fully supported\n"
1920 : "Arg is not sized\n"));
1921 if (A == &FArg) {
1922 ShadowPtr = getCleanShadow(V);
1923 setOrigin(A, getCleanOrigin());
1924 break;
1925 }
1926 continue;
1927 }
1928
1929 unsigned Size = FArg.hasByValAttr()
1930 ? DL.getTypeAllocSize(FArg.getParamByValType())
1931 : DL.getTypeAllocSize(FArg.getType());
1932
1933 if (A == &FArg) {
1934 bool Overflow = ArgOffset + Size > kParamTLSSize;
1935 if (FArg.hasByValAttr()) {
1936 // ByVal pointer itself has clean shadow. We copy the actual
1937 // argument shadow to the underlying memory.
1938 // Figure out maximal valid memcpy alignment.
1939 const Align ArgAlign = DL.getValueOrABITypeAlignment(
1940 FArg.getParamAlign(), FArg.getParamByValType());
1941 Value *CpShadowPtr, *CpOriginPtr;
1942 std::tie(CpShadowPtr, CpOriginPtr) =
1943 getShadowOriginPtr(V, EntryIRB, EntryIRB.getInt8Ty(), ArgAlign,
1944 /*isStore*/ true);
1945 if (!PropagateShadow || Overflow) {
1946 // ParamTLS overflow.
1947 EntryIRB.CreateMemSet(
1948 CpShadowPtr, Constant::getNullValue(EntryIRB.getInt8Ty()),
1949 Size, ArgAlign);
1950 } else {
1951 Value *Base = getShadowPtrForArgument(EntryIRB, ArgOffset);
1952 const Align CopyAlign = std::min(ArgAlign, kShadowTLSAlignment);
1953 [[maybe_unused]] Value *Cpy = EntryIRB.CreateMemCpy(
1954 CpShadowPtr, CopyAlign, Base, CopyAlign, Size);
1955 LLVM_DEBUG(dbgs() << " ByValCpy: " << *Cpy << "\n");
1956
1957 if (MS.TrackOrigins) {
1958 Value *OriginPtr = getOriginPtrForArgument(EntryIRB, ArgOffset);
1959 // FIXME: OriginSize should be:
1960 // alignTo(V % kMinOriginAlignment + Size, kMinOriginAlignment)
1961 unsigned OriginSize = alignTo(Size, kMinOriginAlignment);
1962 EntryIRB.CreateMemCpy(
1963 CpOriginPtr,
1964 /* by getShadowOriginPtr */ kMinOriginAlignment, OriginPtr,
1965 /* by origin_tls[ArgOffset] */ kMinOriginAlignment,
1966 OriginSize);
1967 }
1968 }
1969 }
1970
1971 if (!PropagateShadow || Overflow || FArg.hasByValAttr() ||
1972 (MS.EagerChecks && FArg.hasAttribute(Attribute::NoUndef))) {
1973 ShadowPtr = getCleanShadow(V);
1974 setOrigin(A, getCleanOrigin());
1975 } else {
1976 // Shadow over TLS
1977 Value *Base = getShadowPtrForArgument(EntryIRB, ArgOffset);
1978 ShadowPtr = EntryIRB.CreateAlignedLoad(getShadowTy(&FArg), Base,
1980 if (MS.TrackOrigins) {
1981 Value *OriginPtr = getOriginPtrForArgument(EntryIRB, ArgOffset);
1982 setOrigin(A, EntryIRB.CreateLoad(MS.OriginTy, OriginPtr));
1983 }
1984 }
1986 << " ARG: " << FArg << " ==> " << *ShadowPtr << "\n");
1987 break;
1988 }
1989
1990 ArgOffset += alignTo(Size, kShadowTLSAlignment);
1991 }
1992 assert(ShadowPtr && "Could not find shadow for an argument");
1993 return ShadowPtr;
1994 }
1995
1996 // Check for partially-undefined constant vectors
1997 // TODO: scalable vectors (this is hard because we do not have IRBuilder)
1998 if (isa<FixedVectorType>(V->getType()) && isa<Constant>(V) &&
1999 cast<Constant>(V)->containsUndefOrPoisonElement() && PropagateShadow &&
2000 PoisonUndefVectors) {
2001 unsigned NumElems = cast<FixedVectorType>(V->getType())->getNumElements();
2002 SmallVector<Constant *, 32> ShadowVector(NumElems);
2003 for (unsigned i = 0; i != NumElems; ++i) {
2004 Constant *Elem = cast<Constant>(V)->getAggregateElement(i);
2005 ShadowVector[i] = isa<UndefValue>(Elem) ? getPoisonedShadow(Elem)
2006 : getCleanShadow(Elem);
2007 }
2008
2009 Value *ShadowConstant = ConstantVector::get(ShadowVector);
2010 LLVM_DEBUG(dbgs() << "Partial undef constant vector: " << *V << " ==> "
2011 << *ShadowConstant << "\n");
2012
2013 return ShadowConstant;
2014 }
2015
2016 // TODO: partially-undefined constant arrays, structures, and nested types
2017
2018 // For everything else the shadow is zero.
2019 return getCleanShadow(V);
2020 }
2021
2022 /// Get the shadow for i-th argument of the instruction I.
2023 Value *getShadow(Instruction *I, int i) {
2024 return getShadow(I->getOperand(i));
2025 }
2026
2027 /// Get the origin for a value.
2028 Value *getOrigin(Value *V) {
2029 if (!MS.TrackOrigins)
2030 return nullptr;
2031 if (!PropagateShadow || isa<Constant>(V) || isa<InlineAsm>(V))
2032 return getCleanOrigin();
2034 "Unexpected value type in getOrigin()");
2035 if (Instruction *I = dyn_cast<Instruction>(V)) {
2036 if (I->getMetadata(LLVMContext::MD_nosanitize))
2037 return getCleanOrigin();
2038 }
2039 Value *Origin = OriginMap[V];
2040 assert(Origin && "Missing origin");
2041 return Origin;
2042 }
2043
2044 /// Get the origin for i-th argument of the instruction I.
2045 Value *getOrigin(Instruction *I, int i) {
2046 return getOrigin(I->getOperand(i));
2047 }
2048
2049 /// Remember the place where a shadow check should be inserted.
2050 ///
2051 /// This location will be later instrumented with a check that will print a
2052 /// UMR warning in runtime if the shadow value is not 0.
2053 void insertCheckShadow(Value *Shadow, Value *Origin, Instruction *OrigIns) {
2054 assert(Shadow);
2055 if (!InsertChecks)
2056 return;
2057
2058 if (!DebugCounter::shouldExecute(DebugInsertCheck)) {
2059 LLVM_DEBUG(dbgs() << "Skipping check of " << *Shadow << " before "
2060 << *OrigIns << "\n");
2061 return;
2062 }
2063
2064 Type *ShadowTy = Shadow->getType();
2065 if (isScalableNonVectorType(ShadowTy)) {
2066 LLVM_DEBUG(dbgs() << "Skipping check of scalable non-vector " << *Shadow
2067 << " before " << *OrigIns << "\n");
2068 return;
2069 }
2070#ifndef NDEBUG
2071 assert((isa<IntegerType>(ShadowTy) || isa<VectorType>(ShadowTy) ||
2072 isa<StructType>(ShadowTy) || isa<ArrayType>(ShadowTy)) &&
2073 "Can only insert checks for integer, vector, and aggregate shadow "
2074 "types");
2075#endif
2076 InstrumentationList.push_back(
2077 ShadowOriginAndInsertPoint(Shadow, Origin, OrigIns));
2078 }
2079
2080 /// Get shadow for value, and remember the place where a shadow check should
2081 /// be inserted.
2082 ///
2083 /// This location will be later instrumented with a check that will print a
2084 /// UMR warning in runtime if the value is not fully defined.
2085 void insertCheckShadowOf(Value *Val, Instruction *OrigIns) {
2086 assert(Val);
2087 Value *Shadow, *Origin;
2088 if (Opts.msan_check_constant_shadow) {
2089 Shadow = getShadow(Val);
2090 if (!Shadow)
2091 return;
2092 Origin = getOrigin(Val);
2093 } else {
2094 Shadow = dyn_cast_or_null<Instruction>(getShadow(Val));
2095 if (!Shadow)
2096 return;
2097 Origin = dyn_cast_or_null<Instruction>(getOrigin(Val));
2098 }
2099 insertCheckShadow(Shadow, Origin, OrigIns);
2100 }
2101
2103 switch (a) {
2104 case AtomicOrdering::NotAtomic:
2105 return AtomicOrdering::NotAtomic;
2106 case AtomicOrdering::Unordered:
2107 case AtomicOrdering::Monotonic:
2108 case AtomicOrdering::Release:
2109 return AtomicOrdering::Release;
2110 case AtomicOrdering::Acquire:
2111 case AtomicOrdering::AcquireRelease:
2112 return AtomicOrdering::AcquireRelease;
2113 case AtomicOrdering::SequentiallyConsistent:
2114 return AtomicOrdering::SequentiallyConsistent;
2115 }
2116 llvm_unreachable("Unknown ordering");
2117 }
2118
2119 Value *makeAddReleaseOrderingTable(IRBuilder<> &IRB) {
2120 constexpr int NumOrderings = (int)AtomicOrderingCABI::seq_cst + 1;
2121 uint32_t OrderingTable[NumOrderings] = {};
2122
2123 OrderingTable[(int)AtomicOrderingCABI::relaxed] =
2124 OrderingTable[(int)AtomicOrderingCABI::release] =
2125 (int)AtomicOrderingCABI::release;
2126 OrderingTable[(int)AtomicOrderingCABI::consume] =
2127 OrderingTable[(int)AtomicOrderingCABI::acquire] =
2128 OrderingTable[(int)AtomicOrderingCABI::acq_rel] =
2129 (int)AtomicOrderingCABI::acq_rel;
2130 OrderingTable[(int)AtomicOrderingCABI::seq_cst] =
2131 (int)AtomicOrderingCABI::seq_cst;
2132
2133 return ConstantDataVector::get(IRB.getContext(), OrderingTable);
2134 }
2135
2137 switch (a) {
2138 case AtomicOrdering::NotAtomic:
2139 return AtomicOrdering::NotAtomic;
2140 case AtomicOrdering::Unordered:
2141 case AtomicOrdering::Monotonic:
2142 case AtomicOrdering::Acquire:
2143 return AtomicOrdering::Acquire;
2144 case AtomicOrdering::Release:
2145 case AtomicOrdering::AcquireRelease:
2146 return AtomicOrdering::AcquireRelease;
2147 case AtomicOrdering::SequentiallyConsistent:
2148 return AtomicOrdering::SequentiallyConsistent;
2149 }
2150 llvm_unreachable("Unknown ordering");
2151 }
2152
2153 Value *makeAddAcquireOrderingTable(IRBuilder<> &IRB) {
2154 constexpr int NumOrderings = (int)AtomicOrderingCABI::seq_cst + 1;
2155 uint32_t OrderingTable[NumOrderings] = {};
2156
2157 OrderingTable[(int)AtomicOrderingCABI::relaxed] =
2158 OrderingTable[(int)AtomicOrderingCABI::acquire] =
2159 OrderingTable[(int)AtomicOrderingCABI::consume] =
2160 (int)AtomicOrderingCABI::acquire;
2161 OrderingTable[(int)AtomicOrderingCABI::release] =
2162 OrderingTable[(int)AtomicOrderingCABI::acq_rel] =
2163 (int)AtomicOrderingCABI::acq_rel;
2164 OrderingTable[(int)AtomicOrderingCABI::seq_cst] =
2165 (int)AtomicOrderingCABI::seq_cst;
2166
2167 return ConstantDataVector::get(IRB.getContext(), OrderingTable);
2168 }
2169
2170 // ------------------- Visitors.
2171 using InstVisitor<MemorySanitizerVisitor>::visit;
2172 void visit(Instruction &I) {
2173 if (I.getMetadata(LLVMContext::MD_nosanitize))
2174 return;
2175 // Don't want to visit if we're in the prologue
2176 if (isInPrologue(I))
2177 return;
2178 if (!DebugCounter::shouldExecute(DebugInstrumentInstruction)) {
2179 LLVM_DEBUG(dbgs() << "Skipping instruction: " << I << "\n");
2180 // We still need to set the shadow and origin to clean values.
2181 setShadow(&I, getCleanShadow(&I));
2182 setOrigin(&I, getCleanOrigin());
2183 return;
2184 }
2185
2186 Instructions.push_back(&I);
2187 }
2188
2189 /// Instrument LoadInst
2190 ///
2191 /// Loads the corresponding shadow and (optionally) origin.
2192 /// Optionally, checks that the load address is fully defined.
2193 void visitLoadInst(LoadInst &I) {
2194 assert(I.getType()->isSized() && "Load type must have size");
2195 assert(!I.getMetadata(LLVMContext::MD_nosanitize));
2196 NextNodeIRBuilder IRB(&I);
2197 Type *ShadowTy = getShadowTy(&I);
2198 Value *Addr = I.getPointerOperand();
2199 Value *ShadowPtr = nullptr, *OriginPtr = nullptr;
2200 const Align Alignment = I.getAlign();
2201 if (PropagateShadow) {
2202 std::tie(ShadowPtr, OriginPtr) =
2203 getShadowOriginPtr(Addr, IRB, ShadowTy, Alignment, /*isStore*/ false);
2204 setShadow(&I,
2205 IRB.CreateAlignedLoad(ShadowTy, ShadowPtr, Alignment, "_msld"));
2206 } else {
2207 setShadow(&I, getCleanShadow(&I));
2208 }
2209
2210 if (Opts.msan_check_access_address)
2211 insertCheckShadowOf(I.getPointerOperand(), &I);
2212
2213 if (I.isAtomic())
2214 I.setOrdering(addAcquireOrdering(I.getOrdering()));
2215
2216 if (MS.TrackOrigins) {
2217 if (PropagateShadow) {
2218 const Align OriginAlignment = std::max(kMinOriginAlignment, Alignment);
2219 setOrigin(
2220 &I, IRB.CreateAlignedLoad(MS.OriginTy, OriginPtr, OriginAlignment));
2221 } else {
2222 setOrigin(&I, getCleanOrigin());
2223 }
2224 }
2225 }
2226
2227 /// Instrument StoreInst
2228 ///
2229 /// Stores the corresponding shadow and (optionally) origin.
2230 /// Optionally, checks that the store address is fully defined.
2231 void visitStoreInst(StoreInst &I) {
2232 StoreList.push_back(&I);
2233 if (Opts.msan_check_access_address)
2234 insertCheckShadowOf(I.getPointerOperand(), &I);
2235 }
2236
2237 void handleCASOrRMW(Instruction &I) {
2239
2240 IRBuilder<> IRB(&I);
2241 Value *Addr = I.getOperand(0);
2242 Value *Val = I.getOperand(1);
2243 Value *ShadowPtr = getShadowOriginPtr(Addr, IRB, getShadowTy(Val), Align(1),
2244 /*isStore*/ true)
2245 .first;
2246
2247 if (Opts.msan_check_access_address)
2248 insertCheckShadowOf(Addr, &I);
2249
2250 // Only test the conditional argument of cmpxchg instruction.
2251 // The other argument can potentially be uninitialized, but we can not
2252 // detect this situation reliably without possible false positives.
2254 insertCheckShadowOf(Val, &I);
2255
2256 IRB.CreateStore(getCleanShadow(Val), ShadowPtr);
2257
2258 setShadow(&I, getCleanShadow(&I));
2259 setOrigin(&I, getCleanOrigin());
2260 }
2261
2262 void visitAtomicRMWInst(AtomicRMWInst &I) {
2263 handleCASOrRMW(I);
2264 I.setOrdering(addReleaseOrdering(I.getOrdering()));
2265 }
2266
2267 void visitAtomicCmpXchgInst(AtomicCmpXchgInst &I) {
2268 handleCASOrRMW(I);
2269 I.setSuccessOrdering(addReleaseOrdering(I.getSuccessOrdering()));
2270 }
2271
2272 /// Generic handler to compute shadow for == and != comparisons.
2273 ///
2274 /// This function is used by handleEqualityComparison and visitSwitchInst.
2275 ///
2276 /// Sometimes the comparison result is known even if some of the bits of the
2277 /// arguments are not.
2278 Value *propagateEqualityComparison(IRBuilder<> &IRB, Value *A, Value *B,
2279 Value *Sa, Value *Sb) {
2280 assert(getShadowTy(A) == Sa->getType());
2281 assert(getShadowTy(B) == Sb->getType());
2282
2283 // Get rid of pointers and vectors of pointers.
2284 // For ints (and vectors of ints), types of A and Sa match,
2285 // and this is a no-op.
2286 A = IRB.CreatePointerCast(A, Sa->getType());
2287 B = IRB.CreatePointerCast(B, Sb->getType());
2288
2289 // A == B <==> (C = A^B) == 0
2290 // A != B <==> (C = A^B) != 0
2291 // Sc = Sa | Sb
2292 Value *C = IRB.CreateXor(A, B);
2293 Value *Sc = IRB.CreateOr(Sa, Sb);
2294 // Now dealing with i = (C == 0) comparison (or C != 0, does not matter now)
2295 // Result is defined if one of the following is true
2296 // * there is a defined 1 bit in C
2297 // * C is fully defined
2298 // Si = !(C & ~Sc) && Sc
2300 Value *MinusOne = Constant::getAllOnesValue(Sc->getType());
2301 Value *LHS = IRB.CreateICmpNE(Sc, Zero);
2302 Value *RHS =
2303 IRB.CreateICmpEQ(IRB.CreateAnd(IRB.CreateXor(Sc, MinusOne), C), Zero);
2304 Value *Si = IRB.CreateAnd(LHS, RHS);
2305 Si->setName("_msprop_icmp");
2306
2307 return Si;
2308 }
2309
2310 // Instrument:
2311 // switch i32 %Val, label %else [ i32 0, label %A
2312 // i32 1, label %B
2313 // i32 2, label %C ]
2314 //
2315 // Typically, the switch input value (%Val) is fully initialized.
2316 //
2317 // Sometimes the compiler may convert (icmp + br) into a switch statement.
2318 // MSan allows icmp eq/ne with partly initialized inputs to still result in a
2319 // fully initialized output, if there exists a bit that is initialized in
2320 // both inputs with a differing value. For compatibility, we support this in
2321 // the switch instrumentation as well. Note that this edge case only applies
2322 // if the switch input value does not match *any* of the cases (matching any
2323 // of the cases requires an exact, fully initialized match).
2324 //
2325 // ShadowCases = 0
2326 // | propagateEqualityComparison(Val, 0)
2327 // | propagateEqualityComparison(Val, 1)
2328 // | propagateEqualityComparison(Val, 2))
2329 void visitSwitchInst(SwitchInst &SI) {
2330 IRBuilder<> IRB(&SI);
2331
2332 Value *Val = SI.getCondition();
2333 Value *ShadowVal = getShadow(Val);
2334 // TODO: add fast path - if the condition is fully initialized, we know
2335 // there is no UUM, without needing to consider the case values below.
2336
2337 // Some code (e.g., AMDGPUGenMCCodeEmitter.inc) has tens of thousands of
2338 // cases. This results in an extremely long chained expression for MSan's
2339 // switch instrumentation, which can cause the JumpThreadingPass to have a
2340 // stack overflow or excessive runtime. We limit the number of cases
2341 // considered, with the tradeoff of niche false negatives.
2342 // TODO: figure out a better solution.
2343 int casesToConsider = Opts.msan_switch_precision;
2344
2345 Value *ShadowCases = nullptr;
2346 for (auto Case : SI.cases()) {
2347 if (casesToConsider <= 0)
2348 break;
2349
2350 Value *Comparator = Case.getCaseValue();
2351 // TODO: some simplification is possible when comparing multiple cases
2352 // simultaneously.
2353 Value *ComparisonShadow = propagateEqualityComparison(
2354 IRB, Val, Comparator, ShadowVal, getShadow(Comparator));
2355
2356 if (ShadowCases)
2357 ShadowCases = IRB.CreateOr(ShadowCases, ComparisonShadow);
2358 else
2359 ShadowCases = ComparisonShadow;
2360
2361 casesToConsider--;
2362 }
2363
2364 if (ShadowCases)
2365 insertCheckShadow(ShadowCases, getOrigin(Val), &SI);
2366 }
2367
2368 // Vector manipulation.
2369 void visitExtractElementInst(ExtractElementInst &I) {
2370 insertCheckShadowOf(I.getOperand(1), &I);
2371 IRBuilder<> IRB(&I);
2372 setShadow(&I, IRB.CreateExtractElement(getShadow(&I, 0), I.getOperand(1),
2373 "_msprop"));
2374 setOrigin(&I, getOrigin(&I, 0));
2375 }
2376
2377 void visitInsertElementInst(InsertElementInst &I) {
2378 insertCheckShadowOf(I.getOperand(2), &I);
2379 IRBuilder<> IRB(&I);
2380 auto *Shadow0 = getShadow(&I, 0);
2381 auto *Shadow1 = getShadow(&I, 1);
2382 setShadow(&I, IRB.CreateInsertElement(Shadow0, Shadow1, I.getOperand(2),
2383 "_msprop"));
2384 setOriginForNaryOp(I);
2385 }
2386
2387 void visitShuffleVectorInst(ShuffleVectorInst &I) {
2388 IRBuilder<> IRB(&I);
2389 auto *Shadow0 = getShadow(&I, 0);
2390 auto *Shadow1 = getShadow(&I, 1);
2391 setShadow(&I, IRB.CreateShuffleVector(Shadow0, Shadow1, I.getShuffleMask(),
2392 "_msprop"));
2393 setOriginForNaryOp(I);
2394 }
2395
2396 // Casts.
2397 void visitSExtInst(SExtInst &I) {
2398 IRBuilder<> IRB(&I);
2399 setShadow(&I, IRB.CreateSExt(getShadow(&I, 0), I.getType(), "_msprop"));
2400 setOrigin(&I, getOrigin(&I, 0));
2401 }
2402
2403 void visitZExtInst(ZExtInst &I) {
2404 IRBuilder<> IRB(&I);
2405 setShadow(&I, IRB.CreateZExt(getShadow(&I, 0), I.getType(), "_msprop"));
2406 setOrigin(&I, getOrigin(&I, 0));
2407 }
2408
2409 void visitTruncInst(TruncInst &I) {
2410 IRBuilder<> IRB(&I);
2411 setShadow(&I, IRB.CreateTrunc(getShadow(&I, 0), I.getType(), "_msprop"));
2412 setOrigin(&I, getOrigin(&I, 0));
2413 }
2414
2415 void visitBitCastInst(BitCastInst &I) {
2416 // Special case: if this is the bitcast (there is exactly 1 allowed) between
2417 // a musttail call and a ret, don't instrument. New instructions are not
2418 // allowed after a musttail call.
2419 if (auto *CI = dyn_cast<CallInst>(I.getOperand(0)))
2420 if (CI->isMustTailCall())
2421 return;
2422 IRBuilder<> IRB(&I);
2423 setShadow(&I, IRB.CreateBitCast(getShadow(&I, 0), getShadowTy(&I)));
2424 setOrigin(&I, getOrigin(&I, 0));
2425 }
2426
2427 void visitAddrSpaceCastInst(AddrSpaceCastInst &I) {
2428 IRBuilder<> IRB(&I);
2429 setShadow(&I, IRB.CreateIntCast(getShadow(&I, 0), getShadowTy(&I), false,
2430 "_msprop_addrspacecast"));
2431 setOrigin(&I, getOrigin(&I, 0));
2432 }
2433
2434 void visitPtrToIntInst(PtrToIntInst &I) {
2435 IRBuilder<> IRB(&I);
2436 setShadow(&I, IRB.CreateIntCast(getShadow(&I, 0), getShadowTy(&I), false,
2437 "_msprop_ptrtoint"));
2438 setOrigin(&I, getOrigin(&I, 0));
2439 }
2440
2441 void visitPtrToAddrInst(PtrToAddrInst &I) {
2442 IRBuilder<> IRB(&I);
2443 setShadow(&I, IRB.CreateIntCast(getShadow(&I, 0), getShadowTy(&I), false,
2444 "_msprop_ptrtoaddr"));
2445 setOrigin(&I, getOrigin(&I, 0));
2446 }
2447
2448 void visitIntToPtrInst(IntToPtrInst &I) {
2449 IRBuilder<> IRB(&I);
2450 setShadow(&I, IRB.CreateIntCast(getShadow(&I, 0), getShadowTy(&I), false,
2451 "_msprop_inttoptr"));
2452 setOrigin(&I, getOrigin(&I, 0));
2453 }
2454
2455 /// Handle LLVM and NEON vector convert intrinsics.
2456 ///
2457 /// e.g., <4 x i32> @llvm.aarch64.neon.fcvtpu.v4i32.v4f32(<4 x float>)
2458 /// i32 @llvm.aarch64.neon.fcvtms.i32.f64 (double)
2459 /// <2 x i32> @fptoui (<2 x float>)
2460 /// i64 @llvm.fptosi.sat.i64.f64(double)
2461 ///
2462 /// Note that the size of input/output elements can differ e.g.,
2463 /// double @sitofp(i32)
2464 /// but the number of elements must be the same.
2465 ///
2466 /// For conversions to or from fixed-point, there is a trailing argument to
2467 /// indicate the fixed-point precision:
2468 /// - <4 x float> llvm.aarch64.neon.vcvtfxs2fp.v4f32.v4i32(<4 x i32>, i32)
2469 /// - <4 x i32> llvm.aarch64.neon.vcvtfp2fxu.v4i32.v4f32(<4 x float>, i32)
2470 ///
2471 /// For x86 SSE vector convert intrinsics, see
2472 /// handleSSEVectorConvertIntrinsic().
2473 void handleGenericVectorConvertIntrinsic(Instruction &I, bool FixedPoint) {
2474 [[maybe_unused]] unsigned NumArgs = I.getNumOperands();
2475 if (auto *CI = dyn_cast<CallInst>(&I))
2476 NumArgs = CI->arg_size();
2477
2478 if (FixedPoint) {
2479 assert(NumArgs == 2);
2480 Value *Precision = I.getOperand(1);
2481 insertCheckShadowOf(Precision, &I);
2482 } else {
2483 assert(NumArgs == 1);
2484 }
2485
2486 IRBuilder<> IRB(&I);
2487 Value *S0 = getShadow(&I, 0);
2488
2489 /// For scalars:
2490 /// Since they are converting from floating-point to integer, or between
2491 /// different width floating-point values, the output is:
2492 /// - fully uninitialized if *any* bit of the input is uninitialized
2493 /// - fully ininitialized if all bits of the input are ininitialized
2494 /// We apply the same principle on a per-field basis for vectors.
2495 Value *OutShadow = IRB.CreateSExt(IRB.CreateICmpNE(S0, getCleanShadow(S0)),
2496 getShadowTy(&I));
2497 setShadow(&I, OutShadow);
2498 setOriginForNaryOp(I);
2499 }
2500
2501 void visitFPToSIInst(CastInst &I) {
2502 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2503 }
2504 void visitFPToUIInst(CastInst &I) {
2505 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2506 }
2507 void visitSIToFPInst(CastInst &I) {
2508 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2509 }
2510 void visitUIToFPInst(CastInst &I) {
2511 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2512 }
2513
2514 void visitFPExtInst(CastInst &I) {
2515 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2516 }
2517 void visitFPTruncInst(CastInst &I) {
2518 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
2519 }
2520
2521 /// Generic handler to compute shadow for bitwise AND.
2522 ///
2523 /// This is used by 'visitAnd' but also as a primitive for other handlers.
2524 ///
2525 /// This code is precise: it implements the rule that "And" of an initialized
2526 /// zero bit always results in an initialized value:
2527 // 1&1 => 1; 0&1 => 0; p&1 => p;
2528 // 1&0 => 0; 0&0 => 0; p&0 => 0;
2529 // 1&p => p; 0&p => 0; p&p => p;
2530 //
2531 // S = (S1 & S2) | (V1 & S2) | (S1 & V2)
2532 Value *handleBitwiseAnd(IRBuilder<> &IRB, Value *V1, Value *V2, Value *S1,
2533 Value *S2) {
2534 // "The two arguments to the ‘and’ instruction must be integer or vector
2535 // of integer values. Both arguments must have identical types."
2536 //
2537 // We enforce this condition for all callers to handleBitwiseAnd(); callers
2538 // with non-integer types should call CreateAppToShadowCast() themselves.
2539 assert(V1->getType()->isIntOrIntVectorTy());
2540 assert(V1->getType() == V2->getType());
2541
2542 // Conveniently, getShadowTy() of Int/IntVector returns the original type.
2543 assert(V1->getType() == S1->getType());
2544 assert(V2->getType() == S2->getType());
2545
2546 Value *S1S2 = IRB.CreateAnd(S1, S2);
2547 Value *V1S2 = IRB.CreateAnd(V1, S2);
2548 Value *S1V2 = IRB.CreateAnd(S1, V2);
2549
2550 return IRB.CreateOr({S1S2, V1S2, S1V2});
2551 }
2552
2553 /// Handler for bitwise AND operator.
2554 void visitAnd(BinaryOperator &I) {
2555 IRBuilder<> IRB(&I);
2556 Value *V1 = I.getOperand(0);
2557 Value *V2 = I.getOperand(1);
2558 Value *S1 = getShadow(&I, 0);
2559 Value *S2 = getShadow(&I, 1);
2560
2561 Value *OutShadow = handleBitwiseAnd(IRB, V1, V2, S1, S2);
2562
2563 setShadow(&I, OutShadow);
2564 setOriginForNaryOp(I);
2565 }
2566
2567 void visitOr(BinaryOperator &I) {
2568 IRBuilder<> IRB(&I);
2569 // "Or" of 1 and a poisoned value results in unpoisoned value:
2570 // 1|1 => 1; 0|1 => 1; p|1 => 1;
2571 // 1|0 => 1; 0|0 => 0; p|0 => p;
2572 // 1|p => 1; 0|p => p; p|p => p;
2573 //
2574 // S = (S1 & S2) | (~V1 & S2) | (S1 & ~V2)
2575 //
2576 // If the "disjoint OR" property is violated, the result is poison, and
2577 // hence the entire shadow is uninitialized:
2578 // S = S | SignExt(V1 & V2 != 0)
2579 Value *S1 = getShadow(&I, 0);
2580 Value *S2 = getShadow(&I, 1);
2581 Value *V1 = I.getOperand(0);
2582 Value *V2 = I.getOperand(1);
2583
2584 // "The two arguments to the ‘or’ instruction must be integer or vector
2585 // of integer values. Both arguments must have identical types."
2586 assert(V1->getType()->isIntOrIntVectorTy());
2587 assert(V1->getType() == V2->getType());
2588
2589 // Conveniently, getShadowTy() of Int/IntVector returns the original type.
2590 assert(V1->getType() == S1->getType());
2591 assert(V2->getType() == S2->getType());
2592
2593 Value *NotV1 = IRB.CreateNot(V1);
2594 Value *NotV2 = IRB.CreateNot(V2);
2595
2596 Value *S1S2 = IRB.CreateAnd(S1, S2);
2597 Value *S2NotV1 = IRB.CreateAnd(NotV1, S2);
2598 Value *S1NotV2 = IRB.CreateAnd(S1, NotV2);
2599
2600 Value *S = IRB.CreateOr({S1S2, S2NotV1, S1NotV2});
2601
2602 if (Opts.msan_precise_disjoint_or &&
2603 cast<PossiblyDisjointInst>(&I)->isDisjoint()) {
2604 Value *V1V2 = IRB.CreateAnd(V1, V2);
2605 Value *DisjointOrShadow = IRB.CreateSExt(
2606 IRB.CreateICmpNE(V1V2, getCleanShadow(V1V2)), V1V2->getType());
2607 S = IRB.CreateOr(S, DisjointOrShadow, "_ms_disjoint");
2608 }
2609
2610 setShadow(&I, S);
2611 setOriginForNaryOp(I);
2612 }
2613
2614 /// Default propagation of shadow and/or origin.
2615 ///
2616 /// This class implements the general case of shadow propagation, used in all
2617 /// cases where we don't know and/or don't care about what the operation
2618 /// actually does. It converts all input shadow values to a common type
2619 /// (extending or truncating as necessary), and bitwise OR's them.
2620 ///
2621 /// This is much cheaper than inserting checks (i.e. requiring inputs to be
2622 /// fully initialized), and less prone to false positives.
2623 ///
2624 /// This class also implements the general case of origin propagation. For a
2625 /// Nary operation, result origin is set to the origin of an argument that is
2626 /// not entirely initialized. If there is more than one such arguments, the
2627 /// rightmost of them is picked. It does not matter which one is picked if all
2628 /// arguments are initialized.
2629 template <bool CombineShadow> class Combiner {
2630 Value *Shadow = nullptr;
2631 Value *Origin = nullptr;
2632 IRBuilder<> &IRB;
2633 MemorySanitizerVisitor *MSV;
2634
2635 public:
2636 Combiner(MemorySanitizerVisitor *MSV, IRBuilder<> &IRB)
2637 : IRB(IRB), MSV(MSV) {}
2638
2639 /// Add a pair of shadow and origin values to the mix.
2640 Combiner &Add(Value *OpShadow, Value *OpOrigin) {
2641 if (CombineShadow) {
2642 assert(OpShadow);
2643 if (!Shadow)
2644 Shadow = OpShadow;
2645 else {
2646 OpShadow = MSV->CreateShadowCast(IRB, OpShadow, Shadow->getType());
2647 Shadow = IRB.CreateOr(Shadow, OpShadow, "_msprop");
2648 }
2649 }
2650
2651 if (MSV->MS.TrackOrigins) {
2652 assert(OpOrigin);
2653 if (!Origin) {
2654 Origin = OpOrigin;
2655 } else {
2656 Constant *ConstOrigin = dyn_cast<Constant>(OpOrigin);
2657 // No point in adding something that might result in 0 origin value.
2658 if (!ConstOrigin || !ConstOrigin->isNullValue()) {
2659 Value *Cond = MSV->convertToBool(OpShadow, IRB);
2660 Origin = IRB.CreateSelect(Cond, OpOrigin, Origin);
2661 }
2662 }
2663 }
2664 return *this;
2665 }
2666
2667 /// Add an application value to the mix.
2668 Combiner &Add(Value *V) {
2669 Value *OpShadow = MSV->getShadow(V);
2670 Value *OpOrigin = MSV->MS.TrackOrigins ? MSV->getOrigin(V) : nullptr;
2671 return Add(OpShadow, OpOrigin);
2672 }
2673
2674 /// Set the current combined values as the given instruction's shadow
2675 /// and origin.
2676 void Done(Instruction *I) {
2677 if (CombineShadow) {
2678 assert(Shadow);
2679 Shadow = MSV->CreateShadowCast(IRB, Shadow, MSV->getShadowTy(I));
2680 MSV->setShadow(I, Shadow);
2681 }
2682 if (MSV->MS.TrackOrigins) {
2683 assert(Origin);
2684 MSV->setOrigin(I, Origin);
2685 }
2686 }
2687
2688 /// Store the current combined value at the specified origin
2689 /// location.
2690 void DoneAndStoreOrigin(TypeSize TS, Value *OriginPtr) {
2691 if (MSV->MS.TrackOrigins) {
2692 assert(Origin);
2693 MSV->paintOrigin(IRB, Origin, OriginPtr, TS, kMinOriginAlignment);
2694 }
2695 }
2696 };
2697
2698 using ShadowAndOriginCombiner = Combiner<true>;
2699 using OriginCombiner = Combiner<false>;
2700
2701 /// Propagate origin for arbitrary operation.
2702 void setOriginForNaryOp(Instruction &I) {
2703 if (!MS.TrackOrigins)
2704 return;
2705 IRBuilder<> IRB(&I);
2706 OriginCombiner OC(this, IRB);
2707 for (Use &Op : I.operands())
2708 OC.Add(Op.get());
2709 OC.Done(&I);
2710 }
2711
2712 size_t VectorOrPrimitiveTypeSizeInBits(Type *Ty) {
2713 assert(!(Ty->isVectorTy() && Ty->getScalarType()->isPointerTy()) &&
2714 "Vector of pointers is not a valid shadow type");
2715 return Ty->isVectorTy() ? cast<FixedVectorType>(Ty)->getNumElements() *
2717 : Ty->getPrimitiveSizeInBits();
2718 }
2719
2720 /// Cast between two shadow types, extending or truncating as
2721 /// necessary.
2722 Value *CreateShadowCast(IRBuilder<> &IRB, Value *V, Type *dstTy,
2723 bool Signed = false) {
2724 Type *srcTy = V->getType();
2725 if (srcTy == dstTy)
2726 return V;
2727 size_t srcSizeInBits = VectorOrPrimitiveTypeSizeInBits(srcTy);
2728 size_t dstSizeInBits = VectorOrPrimitiveTypeSizeInBits(dstTy);
2729 if (srcSizeInBits > 1 && dstSizeInBits == 1)
2730 return IRB.CreateICmpNE(V, getCleanShadow(V));
2731
2732 if (dstTy->isIntegerTy() && srcTy->isIntegerTy())
2733 return IRB.CreateIntCast(V, dstTy, Signed);
2734 if (dstTy->isVectorTy() && srcTy->isVectorTy() &&
2735 cast<VectorType>(dstTy)->getElementCount() ==
2736 cast<VectorType>(srcTy)->getElementCount())
2737 return IRB.CreateIntCast(V, dstTy, Signed);
2738 Value *V1 = IRB.CreateBitCast(V, Type::getIntNTy(*MS.C, srcSizeInBits));
2739 Value *V2 =
2740 IRB.CreateIntCast(V1, Type::getIntNTy(*MS.C, dstSizeInBits), Signed);
2741 return IRB.CreateBitCast(V2, dstTy);
2742 // TODO: handle struct types.
2743 }
2744
2745 /// Cast an application value to the type of its own shadow.
2746 Value *CreateAppToShadowCast(IRBuilder<> &IRB, Value *V) {
2747 Type *ShadowTy = getShadowTy(V);
2748 if (V->getType() == ShadowTy)
2749 return V;
2750 if (V->getType()->isPtrOrPtrVectorTy())
2751 return IRB.CreatePtrToInt(V, ShadowTy);
2752 else
2753 return IRB.CreateBitCast(V, ShadowTy);
2754 }
2755
2756 /// Propagate shadow for arbitrary operation.
2757 void handleShadowOr(Instruction &I) {
2758 IRBuilder<> IRB(&I);
2759 ShadowAndOriginCombiner SC(this, IRB);
2760 for (Use &Op : I.operands())
2761 SC.Add(Op.get());
2762 SC.Done(&I);
2763 }
2764
2765 // Perform a bitwise OR on the horizontal pairs (or other specified grouping)
2766 // of elements.
2767 //
2768 // For example, suppose we have:
2769 // VectorA: <a0, a1, a2, a3, a4, a5>
2770 // VectorB: <b0, b1, b2, b3, b4, b5>
2771 // ReductionFactor: 3
2772 // Shards: 1
2773 // The output would be:
2774 // <a0|a1|a2, a3|a4|a5, b0|b1|b2, b3|b4|b5>
2775 //
2776 // If we have:
2777 // VectorA: <a0, a1, a2, a3, a4, a5, a6, a7>
2778 // VectorB: <b0, b1, b2, b3, b4, b5, b6, b7>
2779 // ReductionFactor: 2
2780 // Shards: 2
2781 // then a and be each have 2 "shards", resulting in the output being
2782 // interleaved:
2783 // <a0|a1, a2|a3, b0|b1, b2|b3, a4|a5, a6|a7, b4|b5, b6|b7>
2784 //
2785 // This is convenient for instrumenting horizontal add/sub.
2786 // For bitwise OR on "vertical" pairs, see maybeHandleSimpleNomemIntrinsic().
2787 Value *horizontalReduce(IntrinsicInst &I, unsigned ReductionFactor,
2788 unsigned Shards, Value *VectorA, Value *VectorB) {
2789 assert(isa<FixedVectorType>(VectorA->getType()));
2790 unsigned NumElems =
2791 cast<FixedVectorType>(VectorA->getType())->getNumElements();
2792
2793 [[maybe_unused]] unsigned TotalNumElems = NumElems;
2794 if (VectorB) {
2795 assert(VectorA->getType() == VectorB->getType());
2796 TotalNumElems *= 2;
2797 }
2798
2799 assert(NumElems % (ReductionFactor * Shards) == 0);
2800
2801 Value *Or = nullptr;
2802
2803 IRBuilder<> IRB(&I);
2804 for (unsigned i = 0; i < ReductionFactor; i++) {
2805 SmallVector<int, 16> Mask;
2806
2807 for (unsigned j = 0; j < Shards; j++) {
2808 unsigned Offset = NumElems / Shards * j;
2809
2810 for (unsigned X = 0; X < NumElems / Shards; X += ReductionFactor)
2811 Mask.push_back(Offset + X + i);
2812
2813 if (VectorB) {
2814 for (unsigned X = 0; X < NumElems / Shards; X += ReductionFactor)
2815 Mask.push_back(NumElems + Offset + X + i);
2816 }
2817 }
2818
2819 Value *Masked;
2820 if (VectorB)
2821 Masked = IRB.CreateShuffleVector(VectorA, VectorB, Mask);
2822 else
2823 Masked = IRB.CreateShuffleVector(VectorA, Mask);
2824
2825 if (Or)
2826 Or = IRB.CreateOr(Or, Masked);
2827 else
2828 Or = Masked;
2829 }
2830
2831 return Or;
2832 }
2833
2834 /// Propagate shadow for 1- or 2-vector intrinsics that combine adjacent
2835 /// fields.
2836 ///
2837 /// e.g., <2 x i32> @llvm.aarch64.neon.saddlp.v2i32.v4i16(<4 x i16>)
2838 /// <16 x i8> @llvm.aarch64.neon.addp.v16i8(<16 x i8>, <16 x i8>)
2839 void handlePairwiseShadowOrIntrinsic(IntrinsicInst &I, unsigned Shards) {
2840 assert(I.arg_size() == 1 || I.arg_size() == 2);
2841
2842 assert(I.getType()->isVectorTy());
2843 assert(I.getArgOperand(0)->getType()->isVectorTy());
2844
2845 [[maybe_unused]] FixedVectorType *ParamType =
2846 cast<FixedVectorType>(I.getArgOperand(0)->getType());
2847 assert((I.arg_size() != 2) ||
2848 (ParamType == cast<FixedVectorType>(I.getArgOperand(1)->getType())));
2849 [[maybe_unused]] FixedVectorType *ReturnType =
2850 cast<FixedVectorType>(I.getType());
2851 assert(ParamType->getNumElements() * I.arg_size() ==
2852 2 * ReturnType->getNumElements());
2853
2854 IRBuilder<> IRB(&I);
2855
2856 // Horizontal OR of shadow
2857 Value *FirstArgShadow = getShadow(&I, 0);
2858 Value *SecondArgShadow = nullptr;
2859 if (I.arg_size() == 2)
2860 SecondArgShadow = getShadow(&I, 1);
2861
2862 Value *OrShadow = horizontalReduce(I, /*ReductionFactor=*/2, Shards,
2863 FirstArgShadow, SecondArgShadow);
2864
2865 OrShadow = CreateShadowCast(IRB, OrShadow, getShadowTy(&I));
2866
2867 setShadow(&I, OrShadow);
2868 setOriginForNaryOp(I);
2869 }
2870
2871 /// Propagate shadow for 1- or 2-vector intrinsics that combine adjacent
2872 /// fields, with the parameters reinterpreted to have elements of a specified
2873 /// width. For example:
2874 /// @llvm.x86.ssse3.phadd.w(<1 x i64> [[VAR1]], <1 x i64> [[VAR2]])
2875 /// conceptually operates on
2876 /// (<4 x i16> [[VAR1]], <4 x i16> [[VAR2]])
2877 /// and can be handled with ReinterpretElemWidth == 16.
2878 void handlePairwiseShadowOrIntrinsic(IntrinsicInst &I, unsigned Shards,
2879 int ReinterpretElemWidth) {
2880 assert(I.arg_size() == 1 || I.arg_size() == 2);
2881
2882 assert(I.getType()->isVectorTy());
2883 assert(I.getArgOperand(0)->getType()->isVectorTy());
2884
2885 FixedVectorType *ParamType =
2886 cast<FixedVectorType>(I.getArgOperand(0)->getType());
2887 assert((I.arg_size() != 2) ||
2888 (ParamType == cast<FixedVectorType>(I.getArgOperand(1)->getType())));
2889
2890 [[maybe_unused]] FixedVectorType *ReturnType =
2891 cast<FixedVectorType>(I.getType());
2892 assert(ParamType->getNumElements() * I.arg_size() ==
2893 2 * ReturnType->getNumElements());
2894
2895 IRBuilder<> IRB(&I);
2896
2897 FixedVectorType *ReinterpretShadowTy = nullptr;
2898 assert(isAligned(Align(ReinterpretElemWidth),
2899 ParamType->getPrimitiveSizeInBits()));
2900 ReinterpretShadowTy = FixedVectorType::get(
2901 IRB.getIntNTy(ReinterpretElemWidth),
2902 ParamType->getPrimitiveSizeInBits() / ReinterpretElemWidth);
2903
2904 // Horizontal OR of shadow
2905 Value *FirstArgShadow = getShadow(&I, 0);
2906 FirstArgShadow = IRB.CreateBitCast(FirstArgShadow, ReinterpretShadowTy);
2907
2908 // If we had two parameters each with an odd number of elements, the total
2909 // number of elements is even, but we have never seen this in extant
2910 // instruction sets, so we enforce that each parameter must have an even
2911 // number of elements.
2913 Align(2),
2914 cast<FixedVectorType>(FirstArgShadow->getType())->getNumElements()));
2915
2916 Value *SecondArgShadow = nullptr;
2917 if (I.arg_size() == 2) {
2918 SecondArgShadow = getShadow(&I, 1);
2919 SecondArgShadow = IRB.CreateBitCast(SecondArgShadow, ReinterpretShadowTy);
2920 }
2921
2922 Value *OrShadow = horizontalReduce(I, /*ReductionFactor=*/2, Shards,
2923 FirstArgShadow, SecondArgShadow);
2924
2925 OrShadow = CreateShadowCast(IRB, OrShadow, getShadowTy(&I));
2926
2927 setShadow(&I, OrShadow);
2928 setOriginForNaryOp(I);
2929 }
2930
2931 void visitFNeg(UnaryOperator &I) { handleShadowOr(I); }
2932
2933 // Handle multiplication by constant.
2934 //
2935 // Handle a special case of multiplication by constant that may have one or
2936 // more zeros in the lower bits. This makes corresponding number of lower bits
2937 // of the result zero as well. We model it by shifting the other operand
2938 // shadow left by the required number of bits. Effectively, we transform
2939 // (X * (A * 2**B)) to ((X << B) * A) and instrument (X << B) as (Sx << B).
2940 // We use multiplication by 2**N instead of shift to cover the case of
2941 // multiplication by 0, which may occur in some elements of a vector operand.
2942 void handleMulByConstant(BinaryOperator &I, Constant *ConstArg,
2943 Value *OtherArg) {
2944 Constant *ShadowMul;
2945 Type *Ty = ConstArg->getType();
2946 if (auto *VTy = dyn_cast<VectorType>(Ty)) {
2947 unsigned NumElements = cast<FixedVectorType>(VTy)->getNumElements();
2948 Type *EltTy = VTy->getElementType();
2950 for (unsigned Idx = 0; Idx < NumElements; ++Idx) {
2951 if (ConstantInt *Elt =
2953 const APInt &V = Elt->getValue();
2954 APInt V2 = APInt(V.getBitWidth(), 1) << V.countr_zero();
2955 Elements.push_back(ConstantInt::get(EltTy, V2));
2956 } else {
2957 Elements.push_back(ConstantInt::get(EltTy, 1));
2958 }
2959 }
2960 ShadowMul = ConstantVector::get(Elements);
2961 } else {
2962 if (ConstantInt *Elt = dyn_cast<ConstantInt>(ConstArg)) {
2963 const APInt &V = Elt->getValue();
2964 APInt V2 = APInt(V.getBitWidth(), 1) << V.countr_zero();
2965 ShadowMul = ConstantInt::get(Ty, V2);
2966 } else {
2967 ShadowMul = ConstantInt::get(Ty, 1);
2968 }
2969 }
2970
2971 IRBuilder<> IRB(&I);
2972 setShadow(&I,
2973 IRB.CreateMul(getShadow(OtherArg), ShadowMul, "msprop_mul_cst"));
2974 setOrigin(&I, getOrigin(OtherArg));
2975 }
2976
2977 void visitMul(BinaryOperator &I) {
2978 Constant *constOp0 = dyn_cast<Constant>(I.getOperand(0));
2979 Constant *constOp1 = dyn_cast<Constant>(I.getOperand(1));
2980 if (constOp0 && !constOp1)
2981 handleMulByConstant(I, constOp0, I.getOperand(1));
2982 else if (constOp1 && !constOp0)
2983 handleMulByConstant(I, constOp1, I.getOperand(0));
2984 else
2985 handleShadowOr(I);
2986 }
2987
2988 void visitFAdd(BinaryOperator &I) { handleShadowOr(I); }
2989 void visitFSub(BinaryOperator &I) { handleShadowOr(I); }
2990 void visitFMul(BinaryOperator &I) { handleShadowOr(I); }
2991 void visitAdd(BinaryOperator &I) { handleShadowOr(I); }
2992 void visitSub(BinaryOperator &I) { handleShadowOr(I); }
2993 void visitXor(BinaryOperator &I) { handleShadowOr(I); }
2994
2995 void handleIntegerDiv(Instruction &I) {
2996 IRBuilder<> IRB(&I);
2997 // Strict on the second argument.
2998 insertCheckShadowOf(I.getOperand(1), &I);
2999 setShadow(&I, getShadow(&I, 0));
3000 setOrigin(&I, getOrigin(&I, 0));
3001 }
3002
3003 void visitUDiv(BinaryOperator &I) { handleIntegerDiv(I); }
3004 void visitSDiv(BinaryOperator &I) { handleIntegerDiv(I); }
3005 void visitURem(BinaryOperator &I) { handleIntegerDiv(I); }
3006 void visitSRem(BinaryOperator &I) { handleIntegerDiv(I); }
3007
3008 // Floating point division is side-effect free. We can not require that the
3009 // divisor is fully initialized and must propagate shadow. See PR37523.
3010 void visitFDiv(BinaryOperator &I) { handleShadowOr(I); }
3011 void visitFRem(BinaryOperator &I) { handleShadowOr(I); }
3012
3013 /// Instrument == and != comparisons.
3014 ///
3015 /// Sometimes the comparison result is known even if some of the bits of the
3016 /// arguments are not.
3017 void handleEqualityComparison(ICmpInst &I) {
3018 IRBuilder<> IRB(&I);
3019 Value *A = I.getOperand(0);
3020 Value *B = I.getOperand(1);
3021 Value *Sa = getShadow(A);
3022 Value *Sb = getShadow(B);
3023
3024 Value *Si = propagateEqualityComparison(IRB, A, B, Sa, Sb);
3025
3026 setShadow(&I, Si);
3027 setOriginForNaryOp(I);
3028 }
3029
3030 /// Instrument relational comparisons.
3031 ///
3032 /// This function does exact shadow propagation for all relational
3033 /// comparisons of integers, pointers and vectors of those.
3034 /// FIXME: output seems suboptimal when one of the operands is a constant
3035 void handleRelationalComparisonExact(ICmpInst &I) {
3036 IRBuilder<> IRB(&I);
3037 Value *A = I.getOperand(0);
3038 Value *B = I.getOperand(1);
3039 Value *Sa = getShadow(A);
3040 Value *Sb = getShadow(B);
3041
3042 // Get rid of pointers and vectors of pointers.
3043 // For ints (and vectors of ints), types of A and Sa match,
3044 // and this is a no-op.
3045 A = IRB.CreatePointerCast(A, Sa->getType());
3046 B = IRB.CreatePointerCast(B, Sb->getType());
3047
3048 // Let [a0, a1] be the interval of possible values of A, taking into account
3049 // its undefined bits. Let [b0, b1] be the interval of possible values of B.
3050 // Then (A cmp B) is defined iff (a0 cmp b1) == (a1 cmp b0).
3051 bool IsSigned = I.isSigned();
3052
3053 auto GetMinMaxUnsigned = [&](Value *V, Value *S) {
3054 if (IsSigned) {
3055 // Sign-flip to map from signed range to unsigned range. Relation A vs B
3056 // should be preserved, if checked with `getUnsignedPredicate()`.
3057 // Relationship between Amin, Amax, Bmin, Bmax also will not be
3058 // affected, as they are created by effectively adding/substructing from
3059 // A (or B) a value, derived from shadow, with no overflow, either
3060 // before or after sign flip.
3061 APInt MinVal =
3062 APInt::getSignedMinValue(V->getType()->getScalarSizeInBits());
3063 V = IRB.CreateXor(V, ConstantInt::get(V->getType(), MinVal));
3064 }
3065 // Minimize undefined bits.
3066 Value *Min = IRB.CreateAnd(V, IRB.CreateNot(S));
3067 Value *Max = IRB.CreateOr(V, S);
3068 return std::make_pair(Min, Max);
3069 };
3070
3071 auto [Amin, Amax] = GetMinMaxUnsigned(A, Sa);
3072 auto [Bmin, Bmax] = GetMinMaxUnsigned(B, Sb);
3073 Value *S1 = IRB.CreateICmp(I.getUnsignedPredicate(), Amin, Bmax);
3074 Value *S2 = IRB.CreateICmp(I.getUnsignedPredicate(), Amax, Bmin);
3075
3076 Value *Si = IRB.CreateXor(S1, S2);
3077 setShadow(&I, Si);
3078 setOriginForNaryOp(I);
3079 }
3080
3081 /// Instrument signed relational comparisons.
3082 ///
3083 /// Handle sign bit tests: x<0, x>=0, x<=-1, x>-1 by propagating the highest
3084 /// bit of the shadow. Everything else is delegated to handleShadowOr().
3085 void handleSignedRelationalComparison(ICmpInst &I) {
3086 Constant *constOp;
3087 Value *op = nullptr;
3089 if ((constOp = dyn_cast<Constant>(I.getOperand(1)))) {
3090 op = I.getOperand(0);
3091 pre = I.getPredicate();
3092 } else if ((constOp = dyn_cast<Constant>(I.getOperand(0)))) {
3093 op = I.getOperand(1);
3094 pre = I.getSwappedPredicate();
3095 } else {
3096 handleShadowOr(I);
3097 return;
3098 }
3099
3100 if ((constOp->isNullValue() &&
3101 (pre == CmpInst::ICMP_SLT || pre == CmpInst::ICMP_SGE)) ||
3102 (constOp->isAllOnesValue() &&
3103 (pre == CmpInst::ICMP_SGT || pre == CmpInst::ICMP_SLE))) {
3104 IRBuilder<> IRB(&I);
3105 Value *Shadow = IRB.CreateICmpSLT(getShadow(op), getCleanShadow(op),
3106 "_msprop_icmp_s");
3107 setShadow(&I, Shadow);
3108 setOrigin(&I, getOrigin(op));
3109 } else {
3110 handleShadowOr(I);
3111 }
3112 }
3113
3114 void visitICmpInst(ICmpInst &I) {
3115 if (!Opts.msan_handle_icmp) {
3116 handleShadowOr(I);
3117 return;
3118 }
3119 if (I.isEquality()) {
3120 handleEqualityComparison(I);
3121 return;
3122 }
3123
3124 assert(I.isRelational());
3125 if (Opts.msan_handle_icmp_exact) {
3126 handleRelationalComparisonExact(I);
3127 return;
3128 }
3129 if (I.isSigned()) {
3130 handleSignedRelationalComparison(I);
3131 return;
3132 }
3133
3134 assert(I.isUnsigned());
3135 if ((isa<Constant>(I.getOperand(0)) || isa<Constant>(I.getOperand(1)))) {
3136 handleRelationalComparisonExact(I);
3137 return;
3138 }
3139
3140 handleShadowOr(I);
3141 }
3142
3143 void visitFCmpInst(FCmpInst &I) { handleShadowOr(I); }
3144
3145 void handleShift(BinaryOperator &I) {
3146 IRBuilder<> IRB(&I);
3147 // If any of the S2 bits are poisoned, the whole thing is poisoned.
3148 // Otherwise perform the same shift on S1.
3149 Value *S1 = getShadow(&I, 0);
3150 Value *S2 = getShadow(&I, 1);
3151 Value *S2Conv =
3152 IRB.CreateSExt(IRB.CreateICmpNE(S2, getCleanShadow(S2)), S2->getType());
3153 Value *V2 = I.getOperand(1);
3154 Value *Shift = IRB.CreateBinOp(I.getOpcode(), S1, V2);
3155 setShadow(&I, IRB.CreateOr(Shift, S2Conv));
3156 setOriginForNaryOp(I);
3157 }
3158
3159 void visitShl(BinaryOperator &I) { handleShift(I); }
3160 void visitAShr(BinaryOperator &I) { handleShift(I); }
3161 void visitLShr(BinaryOperator &I) { handleShift(I); }
3162
3163 void handleFunnelShift(IntrinsicInst &I) {
3164 IRBuilder<> IRB(&I);
3165 // If any of the S2 bits are poisoned, the whole thing is poisoned.
3166 // Otherwise perform the same shift on S0 and S1.
3167 Value *S0 = getShadow(&I, 0);
3168 Value *S1 = getShadow(&I, 1);
3169 Value *S2 = getShadow(&I, 2);
3170 Value *S2Conv =
3171 IRB.CreateSExt(IRB.CreateICmpNE(S2, getCleanShadow(S2)), S2->getType());
3172 Value *V2 = I.getOperand(2);
3173 Value *Shift = IRB.CreateIntrinsic(I.getIntrinsicID(), S2Conv->getType(),
3174 {S0, S1, V2});
3175 setShadow(&I, IRB.CreateOr(Shift, S2Conv));
3176 setOriginForNaryOp(I);
3177 }
3178
3179 // Instrument bit manipulation intrinsics.
3180 // All of these intrinsics are Z = I(SRC, MASK)
3181 // where the types of all operands and the result match.
3182 // The following instrumentation happens to work for all of them:
3183 // Sz = I(Ssrc, MASK) | (sext (Smask != 0))
3184 void handleGenericBitManipulation(IntrinsicInst &I) {
3185 IRBuilder<> IRB(&I);
3186 Type *ShadowTy = getShadowTy(&I);
3187
3188 // If any bit of the mask operand is poisoned, then the whole thing is.
3189 Value *SMask = getShadow(&I, 1);
3190 SMask = IRB.CreateSExt(IRB.CreateICmpNE(SMask, getCleanShadow(ShadowTy)),
3191 ShadowTy);
3192 // Apply the same intrinsic to the shadow of the first operand.
3193 Value *S;
3194 if (Function *Func = I.getCalledFunction())
3195 S = IRB.CreateCall(Func, {getShadow(&I, 0), I.getOperand(1)});
3196 else
3197 S = IRB.CreateIntrinsic(I.getIntrinsicID(), ShadowTy,
3198 {getShadow(&I, 0), I.getOperand(1)});
3199
3200 setShadow(&I, IRB.CreateOr(SMask, S));
3201 setOriginForNaryOp(I);
3202 }
3203
3204 /// Instrument llvm.memmove
3205 ///
3206 /// At this point we don't know if llvm.memmove will be inlined or not.
3207 /// If we don't instrument it and it gets inlined,
3208 /// our interceptor will not kick in and we will lose the memmove.
3209 /// If we instrument the call here, but it does not get inlined,
3210 /// we will memmove the shadow twice: which is bad in case
3211 /// of overlapping regions. So, we simply lower the intrinsic to a call.
3212 ///
3213 /// Similar situation exists for memcpy and memset.
3214 void visitMemMoveInst(MemMoveInst &I) {
3215 getShadow(I.getArgOperand(1)); // Ensure shadow initialized
3216 IRBuilder<> IRB(&I);
3217 IRB.CreateCall(MS.MemmoveFn,
3218 {I.getArgOperand(0), I.getArgOperand(1),
3219 IRB.CreateIntCast(I.getArgOperand(2), MS.IntptrTy, false)});
3221 }
3222
3223 /// Instrument memcpy
3224 ///
3225 /// Similar to memmove: avoid copying shadow twice. This is somewhat
3226 /// unfortunate as it may slowdown small constant memcpys.
3227 /// FIXME: consider doing manual inline for small constant sizes and proper
3228 /// alignment.
3229 ///
3230 /// Note: This also handles memcpy.inline, which promises no calls to external
3231 /// functions as an optimization. However, with instrumentation enabled this
3232 /// is difficult to promise; additionally, we know that the MSan runtime
3233 /// exists and provides __msan_memcpy(). Therefore, we assume that with
3234 /// instrumentation it's safe to turn memcpy.inline into a call to
3235 /// __msan_memcpy(). Should this be wrong, such as when implementing memcpy()
3236 /// itself, instrumentation should be disabled with the no_sanitize attribute.
3237 void visitMemCpyInst(MemCpyInst &I) {
3238 getShadow(I.getArgOperand(1)); // Ensure shadow initialized
3239 IRBuilder<> IRB(&I);
3240 IRB.CreateCall(MS.MemcpyFn,
3241 {I.getArgOperand(0), I.getArgOperand(1),
3242 IRB.CreateIntCast(I.getArgOperand(2), MS.IntptrTy, false)});
3244 }
3245
3246 // Same as memcpy.
3247 void visitMemSetInst(MemSetInst &I) {
3248 IRBuilder<> IRB(&I);
3249 IRB.CreateCall(
3250 MS.MemsetFn,
3251 {I.getArgOperand(0),
3252 IRB.CreateIntCast(I.getArgOperand(1), IRB.getInt32Ty(), false),
3253 IRB.CreateIntCast(I.getArgOperand(2), MS.IntptrTy, false)});
3255 }
3256
3257 void visitVAStartInst(VAStartInst &I) { VAHelper->visitVAStartInst(I); }
3258
3259 void visitVACopyInst(VACopyInst &I) { VAHelper->visitVACopyInst(I); }
3260
3261 /// Handle vector store-like intrinsics.
3262 ///
3263 /// Instrument intrinsics that look like a simple SIMD store: writes memory,
3264 /// has 1 pointer argument and 1 vector argument, returns void.
3265 bool handleVectorStoreIntrinsic(IntrinsicInst &I) {
3266 assert(I.arg_size() == 2);
3267
3268 IRBuilder<> IRB(&I);
3269 Value *Addr = I.getArgOperand(0);
3270 Value *Shadow = getShadow(&I, 1);
3271 Value *ShadowPtr, *OriginPtr;
3272
3273 // We don't know the pointer alignment (could be unaligned SSE store!).
3274 // Have to assume to worst case.
3275 std::tie(ShadowPtr, OriginPtr) = getShadowOriginPtr(
3276 Addr, IRB, Shadow->getType(), Align(1), /*isStore*/ true);
3277 IRB.CreateAlignedStore(Shadow, ShadowPtr, Align(1));
3278
3279 if (Opts.msan_check_access_address)
3280 insertCheckShadowOf(Addr, &I);
3281
3282 // FIXME: factor out common code from materializeStores
3283 if (MS.TrackOrigins)
3284 IRB.CreateStore(getOrigin(&I, 1), OriginPtr);
3285 return true;
3286 }
3287
3288 /// Handle vector load-like intrinsics.
3289 ///
3290 /// Instrument intrinsics that look like a simple SIMD load: reads memory,
3291 /// has 1 pointer argument, returns a vector.
3292 bool handleVectorLoadIntrinsic(IntrinsicInst &I) {
3293 assert(I.arg_size() == 1);
3294
3295 IRBuilder<> IRB(&I);
3296 Value *Addr = I.getArgOperand(0);
3297
3298 Type *ShadowTy = getShadowTy(&I);
3299 Value *ShadowPtr = nullptr, *OriginPtr = nullptr;
3300 if (PropagateShadow) {
3301 // We don't know the pointer alignment (could be unaligned SSE load!).
3302 // Have to assume to worst case.
3303 const Align Alignment = Align(1);
3304 std::tie(ShadowPtr, OriginPtr) =
3305 getShadowOriginPtr(Addr, IRB, ShadowTy, Alignment, /*isStore*/ false);
3306 setShadow(&I,
3307 IRB.CreateAlignedLoad(ShadowTy, ShadowPtr, Alignment, "_msld"));
3308 } else {
3309 setShadow(&I, getCleanShadow(&I));
3310 }
3311
3312 if (Opts.msan_check_access_address)
3313 insertCheckShadowOf(Addr, &I);
3314
3315 if (MS.TrackOrigins) {
3316 if (PropagateShadow)
3317 setOrigin(&I, IRB.CreateLoad(MS.OriginTy, OriginPtr));
3318 else
3319 setOrigin(&I, getCleanOrigin());
3320 }
3321 return true;
3322 }
3323
3324 /// Handle (SIMD arithmetic)-like intrinsics.
3325 ///
3326 /// Instrument intrinsics with any number of arguments of the same type [*],
3327 /// equal to the return type, plus a specified number of trailing flags of
3328 /// any type.
3329 ///
3330 /// [*] The type should be simple (no aggregates or pointers; vectors are
3331 /// fine).
3332 ///
3333 /// Caller guarantees that this intrinsic does not access memory.
3334 ///
3335 /// TODO: "horizontal"/"pairwise" intrinsics are often incorrectly matched by
3336 /// by this handler. See horizontalReduce().
3337 ///
3338 /// TODO: permutation intrinsics are also often incorrectly matched.
3339 [[maybe_unused]] bool
3340 maybeHandleSimpleNomemIntrinsic(IntrinsicInst &I,
3341 unsigned int trailingFlags) {
3342 Type *RetTy = I.getType();
3343 if (!(RetTy->isIntOrIntVectorTy() || RetTy->isFPOrFPVectorTy()))
3344 return false;
3345
3346 unsigned NumArgOperands = I.arg_size();
3347 assert(NumArgOperands >= trailingFlags);
3348 for (unsigned i = 0; i < NumArgOperands - trailingFlags; ++i) {
3349 Type *Ty = I.getArgOperand(i)->getType();
3350 if (Ty != RetTy)
3351 return false;
3352 }
3353
3354 IRBuilder<> IRB(&I);
3355 ShadowAndOriginCombiner SC(this, IRB);
3356 for (unsigned i = 0; i < NumArgOperands; ++i)
3357 SC.Add(I.getArgOperand(i));
3358 SC.Done(&I);
3359
3360 return true;
3361 }
3362
3363 /// Returns whether it was able to heuristically instrument unknown
3364 /// intrinsics.
3365 ///
3366 /// The main purpose of this code is to do something reasonable with all
3367 /// random intrinsics we might encounter, most importantly - SIMD intrinsics.
3368 /// We recognize several classes of intrinsics by their argument types and
3369 /// ModRefBehaviour and apply special instrumentation when we are reasonably
3370 /// sure that we know what the intrinsic does.
3371 ///
3372 /// We special-case intrinsics where this approach fails. See llvm.bswap
3373 /// handling as an example of that.
3374 bool maybeHandleUnknownIntrinsicUnlogged(IntrinsicInst &I) {
3375 unsigned NumArgOperands = I.arg_size();
3376 if (NumArgOperands == 0)
3377 return false;
3378
3379 if (NumArgOperands == 2 && I.getArgOperand(0)->getType()->isPointerTy() &&
3380 I.getArgOperand(1)->getType()->isVectorTy() &&
3381 I.getType()->isVoidTy() && !I.onlyReadsMemory()) {
3382 // This looks like a vector store.
3383 return handleVectorStoreIntrinsic(I);
3384 }
3385
3386 if (NumArgOperands == 1 && I.getArgOperand(0)->getType()->isPointerTy() &&
3387 I.getType()->isVectorTy() && I.onlyReadsMemory()) {
3388 // This looks like a vector load.
3389 return handleVectorLoadIntrinsic(I);
3390 }
3391
3392 if (I.doesNotAccessMemory())
3393 if (maybeHandleSimpleNomemIntrinsic(I, /*trailingFlags=*/0))
3394 return true;
3395
3396 // FIXME: detect and handle SSE maskstore/maskload?
3397 // Some cases are now handled in handleAVXMasked{Load,Store}.
3398 return false;
3399 }
3400
3401 bool maybeHandleUnknownIntrinsic(IntrinsicInst &I) {
3402 if (maybeHandleUnknownIntrinsicUnlogged(I)) {
3403 if (Opts.msan_dump_heuristic_instructions)
3404 dumpInst(I, "Heuristic");
3405
3406 LLVM_DEBUG(dbgs() << "UNKNOWN INSTRUCTION HANDLED HEURISTICALLY: " << I
3407 << "\n");
3408 return true;
3409 } else
3410 return false;
3411 }
3412
3413 void handleInvariantGroup(IntrinsicInst &I) {
3414 setShadow(&I, getShadow(&I, 0));
3415 setOrigin(&I, getOrigin(&I, 0));
3416 }
3417
3418 void handleLifetimeStart(IntrinsicInst &I) {
3419 if (!PoisonStack)
3420 return;
3421 AllocaInst *AI = dyn_cast<AllocaInst>(I.getArgOperand(0));
3422 if (AI)
3423 LifetimeStartList.push_back(std::make_pair(&I, AI));
3424 }
3425
3426 void handleBswap(IntrinsicInst &I) {
3427 IRBuilder<> IRB(&I);
3428 Value *Op = I.getArgOperand(0);
3429 Type *OpType = Op->getType();
3430 setShadow(&I, IRB.CreateIntrinsic(Intrinsic::bswap, ArrayRef(&OpType, 1),
3431 getShadow(Op)));
3432 setOrigin(&I, getOrigin(Op));
3433 }
3434
3435 // Uninitialized bits are ok if they appear after the leading/trailing 0's
3436 // and a 1. If the input is all zero, it is fully initialized iff
3437 // !is_zero_poison.
3438 //
3439 // e.g., for ctlz, with little-endian, if 0/1 are initialized bits with
3440 // concrete value 0/1, and ? is an uninitialized bit:
3441 // - 0001 0??? is fully initialized
3442 // - 000? ???? is fully uninitialized (*)
3443 // - ???? ???? is fully uninitialized
3444 // - 0000 0000 is fully uninitialized if is_zero_poison,
3445 // fully initialized otherwise
3446 //
3447 // (*) TODO: arguably, since the number of zeros is in the range [3, 8], we
3448 // only need to poison 4 bits.
3449 //
3450 // OutputShadow =
3451 // ((ConcreteZerosCount >= ShadowZerosCount) && !AllZeroShadow)
3452 // || (is_zero_poison && AllZeroSrc)
3453 void handleCountLeadingTrailingZeros(IntrinsicInst &I) {
3454 IRBuilder<> IRB(&I);
3455 Value *Src = I.getArgOperand(0);
3456 Value *SrcShadow = getShadow(Src);
3457
3458 Value *False = IRB.getInt1(false);
3459 Value *ConcreteZerosCount = IRB.CreateIntrinsic(
3460 I.getType(), I.getIntrinsicID(), {Src, /*is_zero_poison=*/False});
3461 Value *ShadowZerosCount = IRB.CreateIntrinsic(
3462 I.getType(), I.getIntrinsicID(), {SrcShadow, /*is_zero_poison=*/False});
3463
3464 Value *CompareConcreteZeros = IRB.CreateICmpUGE(
3465 ConcreteZerosCount, ShadowZerosCount, "_mscz_cmp_zeros");
3466
3467 Value *NotAllZeroShadow =
3468 IRB.CreateIsNotNull(SrcShadow, "_mscz_shadow_not_null");
3469 Value *OutputShadow =
3470 IRB.CreateAnd(CompareConcreteZeros, NotAllZeroShadow, "_mscz_main");
3471
3472 // If zero poison is requested, mix in with the shadow
3473 Constant *IsZeroPoison = cast<Constant>(I.getOperand(1));
3474 if (!IsZeroPoison->isNullValue()) {
3475 Value *BoolZeroPoison = IRB.CreateIsNull(Src, "_mscz_bzp");
3476 OutputShadow = IRB.CreateOr(OutputShadow, BoolZeroPoison, "_mscz_bs");
3477 }
3478
3479 OutputShadow = IRB.CreateSExt(OutputShadow, getShadowTy(Src), "_mscz_os");
3480
3481 setShadow(&I, OutputShadow);
3482 setOriginForNaryOp(I);
3483 }
3484
3485 /// Some instructions have additional zero-elements in the return type
3486 /// e.g., <16 x i8> @llvm.x86.avx512.mask.pmov.qb.512(<8 x i64>, ...)
3487 ///
3488 /// This function will return a vector type with the same number of elements
3489 /// as the input, but same per-element width as the return value e.g.,
3490 /// <8 x i8>.
3491 FixedVectorType *maybeShrinkVectorShadowType(Value *Src, IntrinsicInst &I) {
3492 assert(isa<FixedVectorType>(getShadowTy(&I)));
3493 FixedVectorType *ShadowType = cast<FixedVectorType>(getShadowTy(&I));
3494
3495 // TODO: generalize beyond 2x?
3496 if (ShadowType->getElementCount() ==
3497 cast<VectorType>(Src->getType())->getElementCount() * 2)
3498 ShadowType = FixedVectorType::getHalfElementsVectorType(ShadowType);
3499
3500 assert(ShadowType->getElementCount() ==
3501 cast<VectorType>(Src->getType())->getElementCount());
3502
3503 return ShadowType;
3504 }
3505
3506 /// Doubles the length of a vector shadow (extending with zeros) if necessary
3507 /// to match the length of the shadow for the instruction.
3508 /// If scalar types of the vectors are different, it will use the type of the
3509 /// input vector.
3510 /// This is more type-safe than CreateShadowCast().
3511 Value *maybeExtendVectorShadowWithZeros(Value *Shadow, IntrinsicInst &I) {
3512 IRBuilder<> IRB(&I);
3514 assert(isa<FixedVectorType>(I.getType()));
3515
3516 Value *FullShadow = getCleanShadow(&I);
3517 unsigned ShadowNumElems =
3518 cast<FixedVectorType>(Shadow->getType())->getNumElements();
3519 unsigned FullShadowNumElems =
3520 cast<FixedVectorType>(FullShadow->getType())->getNumElements();
3521
3522 assert((ShadowNumElems == FullShadowNumElems) ||
3523 (ShadowNumElems * 2 == FullShadowNumElems));
3524
3525 if (ShadowNumElems == FullShadowNumElems) {
3526 FullShadow = Shadow;
3527 } else {
3528 // TODO: generalize beyond 2x?
3529 SmallVector<int, 32> ShadowMask(FullShadowNumElems);
3530 std::iota(ShadowMask.begin(), ShadowMask.end(), 0);
3531
3532 // Append zeros
3533 FullShadow =
3534 IRB.CreateShuffleVector(Shadow, getCleanShadow(Shadow), ShadowMask);
3535 }
3536
3537 return FullShadow;
3538 }
3539
3540 /// Handle x86 SSE vector conversion.
3541 ///
3542 /// e.g., single-precision to half-precision conversion:
3543 /// <8 x i16> @llvm.x86.vcvtps2ph.256(<8 x float> %a0, i32 0)
3544 /// <8 x i16> @llvm.x86.vcvtps2ph.128(<4 x float> %a0, i32 0)
3545 ///
3546 /// floating-point to integer:
3547 /// <4 x i32> @llvm.x86.sse2.cvtps2dq(<4 x float>)
3548 /// <4 x i32> @llvm.x86.sse2.cvtpd2dq(<2 x double>)
3549 ///
3550 /// Note: if the output has more elements, they are zero-initialized (and
3551 /// therefore the shadow will also be initialized).
3552 ///
3553 /// This differs from handleSSEVectorConvertIntrinsic() because it
3554 /// propagates uninitialized shadow (instead of checking the shadow).
3555 void handleSSEVectorConvertIntrinsicByProp(IntrinsicInst &I,
3556 bool HasRoundingMode) {
3557 if (HasRoundingMode) {
3558 assert(I.arg_size() == 2);
3559 [[maybe_unused]] Value *RoundingMode = I.getArgOperand(1);
3560 assert(RoundingMode->getType()->isIntegerTy());
3561 } else {
3562 assert(I.arg_size() == 1);
3563 }
3564
3565 Value *Src = I.getArgOperand(0);
3566 assert(Src->getType()->isVectorTy());
3567
3568 // The return type might have more elements than the input.
3569 // Temporarily shrink the return type's number of elements.
3570 VectorType *ShadowType = maybeShrinkVectorShadowType(Src, I);
3571
3572 IRBuilder<> IRB(&I);
3573 Value *S0 = getShadow(&I, 0);
3574
3575 /// For scalars:
3576 /// Since they are converting to and/or from floating-point, the output is:
3577 /// - fully uninitialized if *any* bit of the input is uninitialized
3578 /// - fully ininitialized if all bits of the input are ininitialized
3579 /// We apply the same principle on a per-field basis for vectors.
3580 Value *Shadow =
3581 IRB.CreateSExt(IRB.CreateICmpNE(S0, getCleanShadow(S0)), ShadowType);
3582
3583 // The return type might have more elements than the input.
3584 // Extend the return type back to its original width if necessary.
3585 Value *FullShadow = maybeExtendVectorShadowWithZeros(Shadow, I);
3586
3587 setShadow(&I, FullShadow);
3588 setOriginForNaryOp(I);
3589 }
3590
3591 // Instrument x86 SSE vector convert intrinsic.
3592 //
3593 // This function instruments intrinsics like cvtsi2ss:
3594 // %Out = int_xxx_cvtyyy(%ConvertOp)
3595 // or
3596 // %Out = int_xxx_cvtyyy(%CopyOp, %ConvertOp)
3597 // Intrinsic converts \p NumUsedElements elements of \p ConvertOp to the same
3598 // number \p Out elements, and (if has 2 arguments) copies the rest of the
3599 // elements from \p CopyOp.
3600 // In most cases conversion involves floating-point value which may trigger a
3601 // hardware exception when not fully initialized. For this reason we require
3602 // \p ConvertOp[0:NumUsedElements] to be fully initialized and trap otherwise.
3603 // We copy the shadow of \p CopyOp[NumUsedElements:] to \p
3604 // Out[NumUsedElements:]. This means that intrinsics without \p CopyOp always
3605 // return a fully initialized value.
3606 //
3607 // For Arm NEON vector convert intrinsics, see
3608 // handleNEONVectorConvertIntrinsic().
3609 void handleSSEVectorConvertIntrinsic(IntrinsicInst &I, int NumUsedElements,
3610 bool HasRoundingMode = false) {
3611 IRBuilder<> IRB(&I);
3612 Value *CopyOp, *ConvertOp;
3613
3614 assert((!HasRoundingMode ||
3615 isa<ConstantInt>(I.getArgOperand(I.arg_size() - 1))) &&
3616 "Invalid rounding mode");
3617
3618 switch (I.arg_size() - HasRoundingMode) {
3619 case 2:
3620 CopyOp = I.getArgOperand(0);
3621 ConvertOp = I.getArgOperand(1);
3622 break;
3623 case 1:
3624 ConvertOp = I.getArgOperand(0);
3625 CopyOp = nullptr;
3626 break;
3627 default:
3628 llvm_unreachable("Cvt intrinsic with unsupported number of arguments.");
3629 }
3630
3631 // The first *NumUsedElements* elements of ConvertOp are converted to the
3632 // same number of output elements. The rest of the output is copied from
3633 // CopyOp, or (if not available) filled with zeroes.
3634 // Combine shadow for elements of ConvertOp that are used in this operation,
3635 // and insert a check.
3636 // FIXME: consider propagating shadow of ConvertOp, at least in the case of
3637 // int->any conversion.
3638 Value *ConvertShadow = getShadow(ConvertOp);
3639 Value *AggShadow = nullptr;
3640 if (ConvertOp->getType()->isVectorTy()) {
3641 AggShadow = IRB.CreateExtractElement(
3642 ConvertShadow, ConstantInt::get(IRB.getInt32Ty(), 0));
3643 for (int i = 1; i < NumUsedElements; ++i) {
3644 Value *MoreShadow = IRB.CreateExtractElement(
3645 ConvertShadow, ConstantInt::get(IRB.getInt32Ty(), i));
3646 AggShadow = IRB.CreateOr(AggShadow, MoreShadow);
3647 }
3648 } else {
3649 AggShadow = ConvertShadow;
3650 }
3651 assert(AggShadow->getType()->isIntegerTy());
3652 insertCheckShadow(AggShadow, getOrigin(ConvertOp), &I);
3653
3654 // Build result shadow by zero-filling parts of CopyOp shadow that come from
3655 // ConvertOp.
3656 if (CopyOp) {
3657 assert(CopyOp->getType() == I.getType());
3658 assert(CopyOp->getType()->isVectorTy());
3659 Value *ResultShadow = getShadow(CopyOp);
3660 Type *EltTy = cast<VectorType>(ResultShadow->getType())->getElementType();
3661 for (int i = 0; i < NumUsedElements; ++i) {
3662 ResultShadow = IRB.CreateInsertElement(
3663 ResultShadow, ConstantInt::getNullValue(EltTy),
3664 ConstantInt::get(IRB.getInt32Ty(), i));
3665 }
3666 setShadow(&I, ResultShadow);
3667 setOrigin(&I, getOrigin(CopyOp));
3668 } else {
3669 setShadow(&I, getCleanShadow(&I));
3670 setOrigin(&I, getCleanOrigin());
3671 }
3672 }
3673
3674 // Given a scalar or vector, extract lower 64 bits (or less), and return all
3675 // zeroes if it is zero, and all ones otherwise.
3676 Value *Lower64ShadowExtend(IRBuilder<> &IRB, Value *S, Type *T) {
3677 if (S->getType()->isVectorTy())
3678 S = CreateShadowCast(IRB, S, IRB.getInt64Ty(), /* Signed */ true);
3679 assert(S->getType()->getPrimitiveSizeInBits() <= 64);
3680 Value *S2 = IRB.CreateICmpNE(S, getCleanShadow(S));
3681 return CreateShadowCast(IRB, S2, T, /* Signed */ true);
3682 }
3683
3684 // Given a vector, extract its first element, and return all
3685 // zeroes if it is zero, and all ones otherwise.
3686 Value *LowerElementShadowExtend(IRBuilder<> &IRB, Value *S, Type *T) {
3687 Value *S1 = IRB.CreateExtractElement(S, (uint64_t)0);
3688 Value *S2 = IRB.CreateICmpNE(S1, getCleanShadow(S1));
3689 return CreateShadowCast(IRB, S2, T, /* Signed */ true);
3690 }
3691
3692 Value *VariableShadowExtend(IRBuilder<> &IRB, Value *S) {
3693 Type *T = S->getType();
3694 assert(T->isVectorTy());
3695 Value *S2 = IRB.CreateICmpNE(S, getCleanShadow(S));
3696 return IRB.CreateSExt(S2, T);
3697 }
3698
3699 // Instrument vector shift intrinsic.
3700 //
3701 // This function instruments intrinsics like int_x86_avx2_psll_w.
3702 // Intrinsic shifts %In by %ShiftSize bits.
3703 // %ShiftSize may be a vector. In that case the lower 64 bits determine shift
3704 // size, and the rest is ignored. Behavior is defined even if shift size is
3705 // greater than register (or field) width.
3706 void handleVectorShiftIntrinsic(IntrinsicInst &I, bool Variable) {
3707 assert(I.arg_size() == 2);
3708 IRBuilder<> IRB(&I);
3709 // If any of the S2 bits are poisoned, the whole thing is poisoned.
3710 // Otherwise perform the same shift on S1.
3711 Value *S1 = getShadow(&I, 0);
3712 Value *S2 = getShadow(&I, 1);
3713 Value *S2Conv = Variable ? VariableShadowExtend(IRB, S2)
3714 : Lower64ShadowExtend(IRB, S2, getShadowTy(&I));
3715 Value *V1 = I.getOperand(0);
3716 Value *V2 = I.getOperand(1);
3717 Value *Shift = IRB.CreateCall(I.getFunctionType(), I.getCalledOperand(),
3718 {IRB.CreateBitCast(S1, V1->getType()), V2});
3719 Shift = IRB.CreateBitCast(Shift, getShadowTy(&I));
3720 setShadow(&I, IRB.CreateOr(Shift, S2Conv));
3721 setOriginForNaryOp(I);
3722 }
3723
3724 // Get an MMX-sized (64-bit) vector type, or optionally, other sized
3725 // vectors.
3726 Type *getMMXVectorTy(unsigned EltSizeInBits,
3727 unsigned X86_MMXSizeInBits = 64) {
3728 assert(EltSizeInBits != 0 && (X86_MMXSizeInBits % EltSizeInBits) == 0 &&
3729 "Illegal MMX vector element size");
3730 return FixedVectorType::get(IntegerType::get(*MS.C, EltSizeInBits),
3731 X86_MMXSizeInBits / EltSizeInBits);
3732 }
3733
3734 // Returns a signed counterpart for an (un)signed-saturate-and-pack
3735 // intrinsic.
3736 Intrinsic::ID getSignedPackIntrinsic(Intrinsic::ID id) {
3737 switch (id) {
3738 case Intrinsic::x86_sse2_packsswb_128:
3739 case Intrinsic::x86_sse2_packuswb_128:
3740 return Intrinsic::x86_sse2_packsswb_128;
3741
3742 case Intrinsic::x86_sse2_packssdw_128:
3743 case Intrinsic::x86_sse41_packusdw:
3744 return Intrinsic::x86_sse2_packssdw_128;
3745
3746 case Intrinsic::x86_avx2_packsswb:
3747 case Intrinsic::x86_avx2_packuswb:
3748 return Intrinsic::x86_avx2_packsswb;
3749
3750 case Intrinsic::x86_avx2_packssdw:
3751 case Intrinsic::x86_avx2_packusdw:
3752 return Intrinsic::x86_avx2_packssdw;
3753
3754 case Intrinsic::x86_mmx_packsswb:
3755 case Intrinsic::x86_mmx_packuswb:
3756 return Intrinsic::x86_mmx_packsswb;
3757
3758 case Intrinsic::x86_mmx_packssdw:
3759 return Intrinsic::x86_mmx_packssdw;
3760
3761 case Intrinsic::x86_avx512_packssdw_512:
3762 case Intrinsic::x86_avx512_packusdw_512:
3763 return Intrinsic::x86_avx512_packssdw_512;
3764
3765 case Intrinsic::x86_avx512_packsswb_512:
3766 case Intrinsic::x86_avx512_packuswb_512:
3767 return Intrinsic::x86_avx512_packsswb_512;
3768
3769 default:
3770 llvm_unreachable("unexpected intrinsic id");
3771 }
3772 }
3773
3774 // Instrument vector pack intrinsic.
3775 //
3776 // This function instruments intrinsics like x86_mmx_packsswb, that
3777 // packs elements of 2 input vectors into half as many bits with saturation.
3778 // Shadow is propagated with the signed variant of the same intrinsic applied
3779 // to sext(Sa != zeroinitializer), sext(Sb != zeroinitializer).
3780 // MMXEltSizeInBits is used only for x86mmx arguments.
3781 //
3782 // TODO: consider using GetMinMaxUnsigned() to handle saturation precisely
3783 void handleVectorPackIntrinsic(IntrinsicInst &I,
3784 unsigned MMXEltSizeInBits = 0) {
3785 assert(I.arg_size() == 2);
3786 IRBuilder<> IRB(&I);
3787 Value *S1 = getShadow(&I, 0);
3788 Value *S2 = getShadow(&I, 1);
3789 assert(S1->getType()->isVectorTy());
3790
3791 // SExt and ICmpNE below must apply to individual elements of input vectors.
3792 // In case of x86mmx arguments, cast them to appropriate vector types and
3793 // back.
3794 Type *T =
3795 MMXEltSizeInBits ? getMMXVectorTy(MMXEltSizeInBits) : S1->getType();
3796 if (MMXEltSizeInBits) {
3797 S1 = IRB.CreateBitCast(S1, T);
3798 S2 = IRB.CreateBitCast(S2, T);
3799 }
3800 Value *S1_ext =
3802 Value *S2_ext =
3804 if (MMXEltSizeInBits) {
3805 S1_ext = IRB.CreateBitCast(S1_ext, getMMXVectorTy(64));
3806 S2_ext = IRB.CreateBitCast(S2_ext, getMMXVectorTy(64));
3807 }
3808
3809 Value *S = IRB.CreateIntrinsic(getSignedPackIntrinsic(I.getIntrinsicID()),
3810 {S1_ext, S2_ext}, /*FMFSource=*/nullptr,
3811 "_msprop_vector_pack");
3812 if (MMXEltSizeInBits)
3813 S = IRB.CreateBitCast(S, getShadowTy(&I));
3814 setShadow(&I, S);
3815 setOriginForNaryOp(I);
3816 }
3817
3818 // Convert `Mask` into `<n x i1>`.
3819 Constant *createDppMask(unsigned Width, unsigned Mask) {
3820 SmallVector<Constant *, 4> R(Width);
3821 for (auto &M : R) {
3822 M = ConstantInt::getBool(F.getContext(), Mask & 1);
3823 Mask >>= 1;
3824 }
3825 return ConstantVector::get(R);
3826 }
3827
3828 // Calculate output shadow as array of booleans `<n x i1>`, assuming if any
3829 // arg is poisoned, entire dot product is poisoned.
3830 Value *findDppPoisonedOutput(IRBuilder<> &IRB, Value *S, unsigned SrcMask,
3831 unsigned DstMask) {
3832 const unsigned Width =
3833 cast<FixedVectorType>(S->getType())->getNumElements();
3834
3835 S = IRB.CreateSelect(createDppMask(Width, SrcMask), S,
3837 Value *SElem = IRB.CreateOrReduce(S);
3838 Value *IsClean = IRB.CreateIsNull(SElem, "_msdpp");
3839 Value *DstMaskV = createDppMask(Width, DstMask);
3840
3841 return IRB.CreateSelect(
3842 IsClean, Constant::getNullValue(DstMaskV->getType()), DstMaskV);
3843 }
3844
3845 // See `Intel Intrinsics Guide` for `_dp_p*` instructions.
3846 //
3847 // 2 and 4 element versions produce single scalar of dot product, and then
3848 // puts it into elements of output vector, selected by 4 lowest bits of the
3849 // mask. Top 4 bits of the mask control which elements of input to use for dot
3850 // product.
3851 //
3852 // 8 element version mask still has only 4 bit for input, and 4 bit for output
3853 // mask. According to the spec it just operates as 4 element version on first
3854 // 4 elements of inputs and output, and then on last 4 elements of inputs and
3855 // output.
3856 void handleDppIntrinsic(IntrinsicInst &I) {
3857 IRBuilder<> IRB(&I);
3858
3859 Value *S0 = getShadow(&I, 0);
3860 Value *S1 = getShadow(&I, 1);
3861 Value *S = IRB.CreateOr(S0, S1);
3862
3863 const unsigned Width =
3864 cast<FixedVectorType>(S->getType())->getNumElements();
3865 assert(Width == 2 || Width == 4 || Width == 8);
3866
3867 const unsigned Mask = cast<ConstantInt>(I.getArgOperand(2))->getZExtValue();
3868 const unsigned SrcMask = Mask >> 4;
3869 const unsigned DstMask = Mask & 0xf;
3870
3871 // Calculate shadow as `<n x i1>`.
3872 Value *SI1 = findDppPoisonedOutput(IRB, S, SrcMask, DstMask);
3873 if (Width == 8) {
3874 // First 4 elements of shadow are already calculated. `makeDppShadow`
3875 // operats on 32 bit masks, so we can just shift masks, and repeat.
3876 SI1 = IRB.CreateOr(
3877 SI1, findDppPoisonedOutput(IRB, S, SrcMask << 4, DstMask << 4));
3878 }
3879 // Extend to real size of shadow, poisoning either all or none bits of an
3880 // element.
3881 S = IRB.CreateSExt(SI1, S->getType(), "_msdpp");
3882
3883 setShadow(&I, S);
3884 setOriginForNaryOp(I);
3885 }
3886
3887 Value *convertBlendvToSelectMask(IRBuilder<> &IRB, Value *C) {
3888 C = CreateAppToShadowCast(IRB, C);
3889 FixedVectorType *FVT = cast<FixedVectorType>(C->getType());
3890 unsigned ElSize = FVT->getElementType()->getPrimitiveSizeInBits();
3891 C = IRB.CreateAShr(C, ElSize - 1);
3892 FVT = FixedVectorType::get(IRB.getInt1Ty(), FVT->getNumElements());
3893 return IRB.CreateTrunc(C, FVT);
3894 }
3895
3896 // `blendv(f, t, c)` is effectively `select(c[top_bit], t, f)`.
3897 void handleBlendvIntrinsic(IntrinsicInst &I) {
3898 Value *C = I.getOperand(2);
3899 Value *T = I.getOperand(1);
3900 Value *F = I.getOperand(0);
3901
3902 Value *Sc = getShadow(&I, 2);
3903 Value *Oc = MS.TrackOrigins ? getOrigin(C) : nullptr;
3904
3905 {
3906 IRBuilder<> IRB(&I);
3907 // Extract top bit from condition and its shadow.
3908 C = convertBlendvToSelectMask(IRB, C);
3909 Sc = convertBlendvToSelectMask(IRB, Sc);
3910
3911 setShadow(C, Sc);
3912 setOrigin(C, Oc);
3913 }
3914
3915 handleSelectLikeInst(I, C, T, F);
3916 }
3917
3918 // Instrument sum-of-absolute-differences intrinsic.
3919 void handleVectorSadIntrinsic(IntrinsicInst &I, bool IsMMX = false) {
3920 const unsigned SignificantBitsPerResultElement = 16;
3921 Type *ResTy = IsMMX ? IntegerType::get(*MS.C, 64) : I.getType();
3922 unsigned ZeroBitsPerResultElement =
3923 ResTy->getScalarSizeInBits() - SignificantBitsPerResultElement;
3924
3925 IRBuilder<> IRB(&I);
3926 auto *Shadow0 = getShadow(&I, 0);
3927 auto *Shadow1 = getShadow(&I, 1);
3928 Value *S = IRB.CreateOr(Shadow0, Shadow1);
3929 S = IRB.CreateBitCast(S, ResTy);
3930 S = IRB.CreateSExt(IRB.CreateICmpNE(S, Constant::getNullValue(ResTy)),
3931 ResTy);
3932 S = IRB.CreateLShr(S, ZeroBitsPerResultElement);
3933 S = IRB.CreateBitCast(S, getShadowTy(&I));
3934 setShadow(&I, S);
3935 setOriginForNaryOp(I);
3936 }
3937
3938 // Instrument dot-product / multiply-add(-accumulate)? intrinsics.
3939 //
3940 // e.g., Two operands:
3941 // <4 x i32> @llvm.x86.sse2.pmadd.wd(<8 x i16> %a, <8 x i16> %b)
3942 //
3943 // Two operands which require an EltSizeInBits override:
3944 // <1 x i64> @llvm.x86.mmx.pmadd.wd(<1 x i64> %a, <1 x i64> %b)
3945 //
3946 // Three operands:
3947 // <4 x i32> @llvm.x86.avx512.vpdpbusd.128
3948 // (<4 x i32> %s, <16 x i8> %a, <16 x i8> %b)
3949 // <2 x float> @llvm.aarch64.neon.bfdot.v2f32.v4bf16
3950 // (<2 x float> %acc, <4 x bfloat> %a, <4 x bfloat> %b)
3951 // (these are equivalent to multiply-add on %a and %b, followed by
3952 // adding/"accumulating" %s. "Accumulation" stores the result in one
3953 // of the source registers, but this accumulate vs. add distinction
3954 // is lost when dealing with LLVM intrinsics.)
3955 //
3956 // ZeroPurifies means that multiplying a known-zero with an uninitialized
3957 // value results in an initialized value. This is applicable for integer
3958 // multiplication, but not floating-point (counter-example: NaN).
3959 void handleVectorDotProductIntrinsic(IntrinsicInst &I,
3960 unsigned ReductionFactor,
3961 bool ZeroPurifies,
3962 unsigned EltSizeInBits,
3963 enum OddOrEvenLanes Lanes) {
3964 IRBuilder<> IRB(&I);
3965
3966 [[maybe_unused]] FixedVectorType *ReturnType =
3967 cast<FixedVectorType>(I.getType());
3968 assert(isa<FixedVectorType>(ReturnType));
3969
3970 // Vectors A and B, and shadows
3971 Value *Va = nullptr;
3972 Value *Vb = nullptr;
3973 Value *Sa = nullptr;
3974 Value *Sb = nullptr;
3975
3976 assert(I.arg_size() == 2 || I.arg_size() == 3);
3977 if (I.arg_size() == 2) {
3978 assert(Lanes == kBothLanes);
3979
3980 Va = I.getOperand(0);
3981 Vb = I.getOperand(1);
3982
3983 Sa = getShadow(&I, 0);
3984 Sb = getShadow(&I, 1);
3985 } else if (I.arg_size() == 3) {
3986 // Operand 0 is the accumulator. We will deal with that below.
3987 Va = I.getOperand(1);
3988 Vb = I.getOperand(2);
3989
3990 Sa = getShadow(&I, 1);
3991 Sb = getShadow(&I, 2);
3992
3993 if (Lanes == kEvenLanes || Lanes == kOddLanes) {
3994 // Convert < S0, S1, S2, S3, S4, S5, S6, S7 >
3995 // to < S0, S0, S2, S2, S4, S4, S6, S6 > (if even)
3996 // to < S1, S1, S3, S3, S5, S5, S7, S7 > (if odd)
3997 //
3998 // Note: for aarch64.neon.bfmlalb/t, the odd/even-indexed values are
3999 // zeroed, not duplicated. However, for shadow propagation, this
4000 // distinction is unimportant because Step 1 below will squeeze
4001 // each pair of elements (e.g., [S0, S0]) into a single bit, and
4002 // we only care if it is fully initialized.
4003
4004 FixedVectorType *InputShadowType = cast<FixedVectorType>(Sa->getType());
4005 unsigned Width = InputShadowType->getNumElements();
4006
4007 Sa = IRB.CreateShuffleVector(
4008 Sa, getPclmulMask(Width, /*OddElements=*/Lanes == kOddLanes));
4009 Sb = IRB.CreateShuffleVector(
4010 Sb, getPclmulMask(Width, /*OddElements=*/Lanes == kOddLanes));
4011 }
4012 }
4013
4014 FixedVectorType *ParamType = cast<FixedVectorType>(Va->getType());
4015 assert(ParamType == Vb->getType());
4016
4017 assert(ParamType->getPrimitiveSizeInBits() ==
4018 ReturnType->getPrimitiveSizeInBits());
4019
4020 if (I.arg_size() == 3) {
4021 [[maybe_unused]] auto *AccumulatorType =
4022 cast<FixedVectorType>(I.getOperand(0)->getType());
4023 assert(AccumulatorType == ReturnType);
4024 }
4025
4026 FixedVectorType *ImplicitReturnType =
4027 cast<FixedVectorType>(getShadowTy(ReturnType));
4028 // Step 1: instrument multiplication of corresponding vector elements
4029 if (EltSizeInBits) {
4030 ImplicitReturnType = cast<FixedVectorType>(
4031 getMMXVectorTy(EltSizeInBits * ReductionFactor,
4032 ParamType->getPrimitiveSizeInBits()));
4033 ParamType = cast<FixedVectorType>(
4034 getMMXVectorTy(EltSizeInBits, ParamType->getPrimitiveSizeInBits()));
4035
4036 Va = IRB.CreateBitCast(Va, ParamType);
4037 Vb = IRB.CreateBitCast(Vb, ParamType);
4038
4039 Sa = IRB.CreateBitCast(Sa, getShadowTy(ParamType));
4040 Sb = IRB.CreateBitCast(Sb, getShadowTy(ParamType));
4041 } else {
4042 assert(ParamType->getNumElements() ==
4043 ReturnType->getNumElements() * ReductionFactor);
4044 }
4045
4046 // Each element of the vector is represented by a single bit (poisoned or
4047 // not) e.g., <8 x i1>.
4048 Value *SaNonZero = IRB.CreateIsNotNull(Sa);
4049 Value *SbNonZero = IRB.CreateIsNotNull(Sb);
4050 Value *And;
4051 if (ZeroPurifies) {
4052 // Multiplying an *initialized* zero by an uninitialized element results
4053 // in an initialized zero element.
4054 //
4055 // This is analogous to bitwise AND, where "AND" of 0 and a poisoned value
4056 // results in an unpoisoned value.
4057 Value *VaInt = Va;
4058 Value *VbInt = Vb;
4059 if (!Va->getType()->isIntegerTy()) {
4060 VaInt = CreateAppToShadowCast(IRB, Va);
4061 VbInt = CreateAppToShadowCast(IRB, Vb);
4062 }
4063
4064 // We check for non-zero on a per-element basis, not per-bit.
4065 Value *VaNonZero = IRB.CreateIsNotNull(VaInt);
4066 Value *VbNonZero = IRB.CreateIsNotNull(VbInt);
4067
4068 And = handleBitwiseAnd(IRB, VaNonZero, VbNonZero, SaNonZero, SbNonZero);
4069 } else {
4070 And = IRB.CreateOr({SaNonZero, SbNonZero});
4071 }
4072
4073 // Extend <8 x i1> to <8 x i16>.
4074 // (The real pmadd intrinsic would have computed intermediate values of
4075 // <8 x i32>, but that is irrelevant for our shadow purposes because we
4076 // consider each element to be either fully initialized or fully
4077 // uninitialized.)
4078 And = IRB.CreateSExt(And, Sa->getType());
4079
4080 // Step 2: instrument horizontal add
4081 // We don't need bit-precise horizontalReduce because we only want to check
4082 // if each pair/quad of elements is fully zero.
4083 // Cast to <4 x i32>.
4084 Value *Horizontal = IRB.CreateBitCast(And, ImplicitReturnType);
4085
4086 // Compute <4 x i1>, then extend back to <4 x i32>.
4087 Value *OutShadow = IRB.CreateSExt(
4088 IRB.CreateICmpNE(Horizontal,
4089 Constant::getNullValue(Horizontal->getType())),
4090 ImplicitReturnType);
4091
4092 // Cast it back to the required fake return type (if MMX: <1 x i64>; for
4093 // AVX, it is already correct).
4094 if (EltSizeInBits)
4095 OutShadow = CreateShadowCast(IRB, OutShadow, getShadowTy(&I));
4096
4097 // Step 3 (if applicable): instrument accumulator
4098 if (I.arg_size() == 3)
4099 OutShadow = IRB.CreateOr(OutShadow, getShadow(&I, 0));
4100
4101 setShadow(&I, OutShadow);
4102 setOriginForNaryOp(I);
4103 }
4104
4105 // Instrument compare-packed intrinsic.
4106 //
4107 // x86 has the predicate as the third operand, which is ImmArg e.g.,
4108 // - <4 x double> @llvm.x86.avx.cmp.pd.256(<4 x double>, <4 x double>, i8)
4109 // - <2 x double> @llvm.x86.sse2.cmp.pd(<2 x double>, <2 x double>, i8)
4110 //
4111 // while Arm has separate intrinsics for >= and > e.g.,
4112 // - <2 x i32> @llvm.aarch64.neon.facge.v2i32.v2f32
4113 // (<2 x float> %A, <2 x float>)
4114 // - <2 x i32> @llvm.aarch64.neon.facgt.v2i32.v2f32
4115 // (<2 x float> %A, <2 x float>)
4116 //
4117 // Bonus: this also handles scalar cases e.g.,
4118 // - i32 @llvm.aarch64.neon.facgt.i32.f32(float %A, float %B)
4119 void handleVectorComparePackedIntrinsic(IntrinsicInst &I,
4120 bool PredicateAsOperand) {
4121 if (PredicateAsOperand) {
4122 assert(I.arg_size() == 3);
4123 assert(I.paramHasAttr(2, Attribute::ImmArg));
4124 } else
4125 assert(I.arg_size() == 2);
4126
4127 IRBuilder<> IRB(&I);
4128
4129 // Basically, an or followed by sext(icmp ne 0) to end up with all-zeros or
4130 // all-ones shadow.
4131 Type *ResTy = getShadowTy(&I);
4132 auto *Shadow0 = getShadow(&I, 0);
4133 auto *Shadow1 = getShadow(&I, 1);
4134 Value *S0 = IRB.CreateOr(Shadow0, Shadow1);
4135 Value *S = IRB.CreateSExt(
4136 IRB.CreateICmpNE(S0, Constant::getNullValue(ResTy)), ResTy);
4137 setShadow(&I, S);
4138 setOriginForNaryOp(I);
4139 }
4140
4141 // Instrument compare-scalar intrinsic.
4142 // This handles both cmp* intrinsics which return the result in the first
4143 // element of a vector, and comi* which return the result as i32.
4144 void handleVectorCompareScalarIntrinsic(IntrinsicInst &I) {
4145 IRBuilder<> IRB(&I);
4146 auto *Shadow0 = getShadow(&I, 0);
4147 auto *Shadow1 = getShadow(&I, 1);
4148 Value *S0 = IRB.CreateOr(Shadow0, Shadow1);
4149 Value *S = LowerElementShadowExtend(IRB, S0, getShadowTy(&I));
4150 setShadow(&I, S);
4151 setOriginForNaryOp(I);
4152 }
4153
4154 // Instrument generic vector reduction intrinsics
4155 // by ORing together all their fields.
4156 //
4157 // If AllowShadowCast is true, the return type does not need to be the same
4158 // type as the fields
4159 // e.g., declare i32 @llvm.aarch64.neon.uaddv.i32.v16i8(<16 x i8>)
4160 void handleVectorReduceIntrinsic(IntrinsicInst &I, bool AllowShadowCast) {
4161 assert(I.arg_size() == 1);
4162
4163 IRBuilder<> IRB(&I);
4164 Value *S = IRB.CreateOrReduce(getShadow(&I, 0));
4165 if (AllowShadowCast)
4166 S = CreateShadowCast(IRB, S, getShadowTy(&I));
4167 else
4168 assert(S->getType() == getShadowTy(&I));
4169 setShadow(&I, S);
4170 setOriginForNaryOp(I);
4171 }
4172
4173 // Similar to handleVectorReduceIntrinsic but with an initial starting value.
4174 // e.g., call float @llvm.vector.reduce.fadd.f32.v2f32(float %a0, <2 x float>
4175 // %a1)
4176 // shadow = shadow[a0] | shadow[a1.0] | shadow[a1.1]
4177 //
4178 // The type of the return value, initial starting value, and elements of the
4179 // vector must be identical.
4180 void handleVectorReduceWithStarterIntrinsic(IntrinsicInst &I) {
4181 assert(I.arg_size() == 2);
4182
4183 IRBuilder<> IRB(&I);
4184 Value *Shadow0 = getShadow(&I, 0);
4185 Value *Shadow1 = IRB.CreateOrReduce(getShadow(&I, 1));
4186 assert(Shadow0->getType() == Shadow1->getType());
4187 Value *S = IRB.CreateOr(Shadow0, Shadow1);
4188 assert(S->getType() == getShadowTy(&I));
4189 setShadow(&I, S);
4190 setOriginForNaryOp(I);
4191 }
4192
4193 // Instrument vector.reduce.or intrinsic.
4194 // Valid (non-poisoned) set bits in the operand pull low the
4195 // corresponding shadow bits.
4196 void handleVectorReduceOrIntrinsic(IntrinsicInst &I) {
4197 assert(I.arg_size() == 1);
4198
4199 IRBuilder<> IRB(&I);
4200 Value *OperandShadow = getShadow(&I, 0);
4201 Value *OperandUnsetBits = IRB.CreateNot(I.getOperand(0));
4202 Value *OperandUnsetOrPoison = IRB.CreateOr(OperandUnsetBits, OperandShadow);
4203 // Bit N is clean if any field's bit N is 1 and unpoison
4204 Value *OutShadowMask = IRB.CreateAndReduce(OperandUnsetOrPoison);
4205 // Otherwise, it is clean if every field's bit N is unpoison
4206 Value *OrShadow = IRB.CreateOrReduce(OperandShadow);
4207 Value *S = IRB.CreateAnd(OutShadowMask, OrShadow);
4208
4209 setShadow(&I, S);
4210 setOrigin(&I, getOrigin(&I, 0));
4211 }
4212
4213 // Instrument vector.reduce.and intrinsic.
4214 // Valid (non-poisoned) unset bits in the operand pull down the
4215 // corresponding shadow bits.
4216 void handleVectorReduceAndIntrinsic(IntrinsicInst &I) {
4217 assert(I.arg_size() == 1);
4218
4219 IRBuilder<> IRB(&I);
4220 Value *OperandShadow = getShadow(&I, 0);
4221 Value *OperandSetOrPoison = IRB.CreateOr(I.getOperand(0), OperandShadow);
4222 // Bit N is clean if any field's bit N is 0 and unpoison
4223 Value *OutShadowMask = IRB.CreateAndReduce(OperandSetOrPoison);
4224 // Otherwise, it is clean if every field's bit N is unpoison
4225 Value *OrShadow = IRB.CreateOrReduce(OperandShadow);
4226 Value *S = IRB.CreateAnd(OutShadowMask, OrShadow);
4227
4228 setShadow(&I, S);
4229 setOrigin(&I, getOrigin(&I, 0));
4230 }
4231
4232 void handleStmxcsr(IntrinsicInst &I) {
4233 IRBuilder<> IRB(&I);
4234 Value *Addr = I.getArgOperand(0);
4235 Type *Ty = IRB.getInt32Ty();
4236 Value *ShadowPtr =
4237 getShadowOriginPtr(Addr, IRB, Ty, Align(1), /*isStore*/ true).first;
4238
4239 IRB.CreateStore(getCleanShadow(Ty), ShadowPtr);
4240
4241 if (Opts.msan_check_access_address)
4242 insertCheckShadowOf(Addr, &I);
4243 }
4244
4245 void handleLdmxcsr(IntrinsicInst &I) {
4246 if (!InsertChecks)
4247 return;
4248
4249 IRBuilder<> IRB(&I);
4250 Value *Addr = I.getArgOperand(0);
4251 Type *Ty = IRB.getInt32Ty();
4252 const Align Alignment = Align(1);
4253 Value *ShadowPtr, *OriginPtr;
4254 std::tie(ShadowPtr, OriginPtr) =
4255 getShadowOriginPtr(Addr, IRB, Ty, Alignment, /*isStore*/ false);
4256
4257 if (Opts.msan_check_access_address)
4258 insertCheckShadowOf(Addr, &I);
4259
4260 Value *Shadow = IRB.CreateAlignedLoad(Ty, ShadowPtr, Alignment, "_ldmxcsr");
4261 Value *Origin = MS.TrackOrigins ? IRB.CreateLoad(MS.OriginTy, OriginPtr)
4262 : getCleanOrigin();
4263 insertCheckShadow(Shadow, Origin, &I);
4264 }
4265
4266 void handleMaskedExpandLoad(IntrinsicInst &I) {
4267 IRBuilder<> IRB(&I);
4268 Value *Ptr = I.getArgOperand(0);
4269 MaybeAlign Align = I.getParamAlign(0);
4270 Value *Mask = I.getArgOperand(1);
4271 Value *PassThru = I.getArgOperand(2);
4272
4273 if (Opts.msan_check_access_address) {
4274 insertCheckShadowOf(Ptr, &I);
4275 insertCheckShadowOf(Mask, &I);
4276 }
4277
4278 if (!PropagateShadow) {
4279 setShadow(&I, getCleanShadow(&I));
4280 setOrigin(&I, getCleanOrigin());
4281 return;
4282 }
4283
4284 Type *ShadowTy = getShadowTy(&I);
4285 Type *ElementShadowTy = cast<VectorType>(ShadowTy)->getElementType();
4286 auto [ShadowPtr, OriginPtr] =
4287 getShadowOriginPtr(Ptr, IRB, ElementShadowTy, Align, /*isStore*/ false);
4288
4289 Value *Shadow =
4290 IRB.CreateMaskedExpandLoad(ShadowTy, ShadowPtr, Align, Mask,
4291 getShadow(PassThru), "_msmaskedexpload");
4292
4293 setShadow(&I, Shadow);
4294
4295 // TODO: Store origins.
4296 setOrigin(&I, getCleanOrigin());
4297 }
4298
4299 void handleMaskedCompressStore(IntrinsicInst &I) {
4300 IRBuilder<> IRB(&I);
4301 Value *Values = I.getArgOperand(0);
4302 Value *Ptr = I.getArgOperand(1);
4303 MaybeAlign Align = I.getParamAlign(1);
4304 Value *Mask = I.getArgOperand(2);
4305
4306 if (Opts.msan_check_access_address) {
4307 insertCheckShadowOf(Ptr, &I);
4308 insertCheckShadowOf(Mask, &I);
4309 }
4310
4311 Value *Shadow = getShadow(Values);
4312 Type *ElementShadowTy =
4313 getShadowTy(cast<VectorType>(Values->getType())->getElementType());
4314 auto [ShadowPtr, OriginPtrs] =
4315 getShadowOriginPtr(Ptr, IRB, ElementShadowTy, Align, /*isStore*/ true);
4316
4317 IRB.CreateMaskedCompressStore(Shadow, ShadowPtr, Align, Mask);
4318
4319 // TODO: Store origins.
4320 }
4321
4322 void handleMaskedGather(IntrinsicInst &I) {
4323 IRBuilder<> IRB(&I);
4324 Value *Ptrs = I.getArgOperand(0);
4325 const Align Alignment = I.getParamAlign(0).valueOrOne();
4326 Value *Mask = I.getArgOperand(1);
4327 Value *PassThru = I.getArgOperand(2);
4328
4329 Type *PtrsShadowTy = getShadowTy(Ptrs);
4330 if (Opts.msan_check_access_address) {
4331 insertCheckShadowOf(Mask, &I);
4332 Value *MaskedPtrShadow = IRB.CreateSelect(
4333 Mask, getShadow(Ptrs), Constant::getNullValue((PtrsShadowTy)),
4334 "_msmaskedptrs");
4335 insertCheckShadow(MaskedPtrShadow, getOrigin(Ptrs), &I);
4336 }
4337
4338 if (!PropagateShadow) {
4339 setShadow(&I, getCleanShadow(&I));
4340 setOrigin(&I, getCleanOrigin());
4341 return;
4342 }
4343
4344 Type *ShadowTy = getShadowTy(&I);
4345 Type *ElementShadowTy = cast<VectorType>(ShadowTy)->getElementType();
4346 auto [ShadowPtrs, OriginPtrs] = getShadowOriginPtr(
4347 Ptrs, IRB, ElementShadowTy, Alignment, /*isStore*/ false);
4348
4349 Value *Shadow =
4350 IRB.CreateMaskedGather(ShadowTy, ShadowPtrs, Alignment, Mask,
4351 getShadow(PassThru), "_msmaskedgather");
4352
4353 setShadow(&I, Shadow);
4354
4355 // TODO: Store origins.
4356 setOrigin(&I, getCleanOrigin());
4357 }
4358
4359 void handleMaskedScatter(IntrinsicInst &I) {
4360 IRBuilder<> IRB(&I);
4361 Value *Values = I.getArgOperand(0);
4362 Value *Ptrs = I.getArgOperand(1);
4363 const Align Alignment = I.getParamAlign(1).valueOrOne();
4364 Value *Mask = I.getArgOperand(2);
4365
4366 Type *PtrsShadowTy = getShadowTy(Ptrs);
4367 if (Opts.msan_check_access_address) {
4368 insertCheckShadowOf(Mask, &I);
4369 Value *MaskedPtrShadow = IRB.CreateSelect(
4370 Mask, getShadow(Ptrs), Constant::getNullValue((PtrsShadowTy)),
4371 "_msmaskedptrs");
4372 insertCheckShadow(MaskedPtrShadow, getOrigin(Ptrs), &I);
4373 }
4374
4375 Value *Shadow = getShadow(Values);
4376 Type *ElementShadowTy =
4377 getShadowTy(cast<VectorType>(Values->getType())->getElementType());
4378 auto [ShadowPtrs, OriginPtrs] = getShadowOriginPtr(
4379 Ptrs, IRB, ElementShadowTy, Alignment, /*isStore*/ true);
4380
4381 IRB.CreateMaskedScatter(Shadow, ShadowPtrs, Alignment, Mask);
4382
4383 // TODO: Store origin.
4384 }
4385
4386 // Intrinsic::masked_store
4387 //
4388 // Note: handleAVXMaskedStore handles AVX/AVX2 variants, though AVX512 masked
4389 // stores are lowered to Intrinsic::masked_store.
4390 void handleMaskedStore(IntrinsicInst &I) {
4391 IRBuilder<> IRB(&I);
4392 Value *V = I.getArgOperand(0);
4393 Value *Ptr = I.getArgOperand(1);
4394 const Align Alignment = I.getParamAlign(1).valueOrOne();
4395 Value *Mask = I.getArgOperand(2);
4396 Value *Shadow = getShadow(V);
4397
4398 if (Opts.msan_check_access_address) {
4399 insertCheckShadowOf(Ptr, &I);
4400 insertCheckShadowOf(Mask, &I);
4401 }
4402
4403 Value *ShadowPtr;
4404 Value *OriginPtr;
4405 std::tie(ShadowPtr, OriginPtr) = getShadowOriginPtr(
4406 Ptr, IRB, Shadow->getType(), Alignment, /*isStore*/ true);
4407
4408 IRB.CreateMaskedStore(Shadow, ShadowPtr, Alignment, Mask);
4409
4410 if (!MS.TrackOrigins)
4411 return;
4412
4413 auto &DL = F.getDataLayout();
4414 paintOrigin(IRB, getOrigin(V), OriginPtr,
4415 DL.getTypeStoreSize(Shadow->getType()),
4416 std::max(Alignment, kMinOriginAlignment));
4417 }
4418
4419 // Intrinsic::masked_load
4420 //
4421 // Note: handleAVXMaskedLoad handles AVX/AVX2 variants, though AVX512 masked
4422 // loads are lowered to Intrinsic::masked_load.
4423 void handleMaskedLoad(IntrinsicInst &I) {
4424 IRBuilder<> IRB(&I);
4425 Value *Ptr = I.getArgOperand(0);
4426 const Align Alignment = I.getParamAlign(0).valueOrOne();
4427 Value *Mask = I.getArgOperand(1);
4428 Value *PassThru = I.getArgOperand(2);
4429
4430 if (Opts.msan_check_access_address) {
4431 insertCheckShadowOf(Ptr, &I);
4432 insertCheckShadowOf(Mask, &I);
4433 }
4434
4435 if (!PropagateShadow) {
4436 setShadow(&I, getCleanShadow(&I));
4437 setOrigin(&I, getCleanOrigin());
4438 return;
4439 }
4440
4441 Type *ShadowTy = getShadowTy(&I);
4442 Value *ShadowPtr, *OriginPtr;
4443 std::tie(ShadowPtr, OriginPtr) =
4444 getShadowOriginPtr(Ptr, IRB, ShadowTy, Alignment, /*isStore*/ false);
4445 setShadow(&I, IRB.CreateMaskedLoad(ShadowTy, ShadowPtr, Alignment, Mask,
4446 getShadow(PassThru), "_msmaskedld"));
4447
4448 if (!MS.TrackOrigins)
4449 return;
4450
4451 // Choose between PassThru's and the loaded value's origins.
4452 Value *MaskedPassThruShadow = IRB.CreateAnd(
4453 getShadow(PassThru), IRB.CreateSExt(IRB.CreateNeg(Mask), ShadowTy));
4454
4455 Value *NotNull = convertToBool(MaskedPassThruShadow, IRB, "_mscmp");
4456
4457 Value *PtrOrigin = IRB.CreateLoad(MS.OriginTy, OriginPtr);
4458 Value *Origin = IRB.CreateSelect(NotNull, getOrigin(PassThru), PtrOrigin);
4459
4460 setOrigin(&I, Origin);
4461 }
4462
4463 // e.g., <4 x i32> @llvm.masked.udiv.v4i32(<4 x i32> %dividend,
4464 // <4 x i32> %divisor,
4465 // <4 x i1> %mask)
4466 //
4467 // As handleIntegerDiv(), but per-lane: strict on the divisor and propagating
4468 // the dividend, both only on the enabled lanes. Disabled lanes cannot cause
4469 // undefined behaviour, and their result is poison.
4470 void handleMaskedIntegerDivRem(IntrinsicInst &I) {
4471 assert(I.arg_size() == 3);
4472 IRBuilder<> IRB(&I);
4473 Value *Dividend = I.getArgOperand(0);
4474 Value *Divisor = I.getArgOperand(1);
4475 Value *Mask = I.getArgOperand(2);
4476
4477 insertCheckShadowOf(Mask, &I);
4478
4479 Value *MaskedDivisorShadow = IRB.CreateSelect(
4480 Mask, getShadow(Divisor), getCleanShadow(Divisor), "_msmaskeddivisor");
4481 insertCheckShadow(MaskedDivisorShadow, getOrigin(Divisor), &I);
4482
4483 if (!PropagateShadow) {
4484 setShadow(&I, getCleanShadow(&I));
4485 setOrigin(&I, getCleanOrigin());
4486 return;
4487 }
4488
4489 setShadow(&I, IRB.CreateSelect(Mask, getShadow(Dividend),
4490 getPoisonedShadow(&I), "_msmaskeddiv"));
4491 setOrigin(&I, getOrigin(Dividend));
4492 }
4493
4494 // e.g., void @llvm.x86.avx.maskstore.ps.256(ptr, <8 x i32>, <8 x float>)
4495 // dst mask src
4496 //
4497 // AVX512 masked stores are lowered to Intrinsic::masked_load and are handled
4498 // by handleMaskedStore.
4499 //
4500 // This function handles AVX and AVX2 masked stores; these use the MSBs of a
4501 // vector of integers, unlike the LLVM masked intrinsics, which require a
4502 // vector of booleans. X86InstCombineIntrinsic.cpp::simplifyX86MaskedLoad
4503 // mentions that the x86 backend does not know how to efficiently convert
4504 // from a vector of booleans back into the AVX mask format; therefore, they
4505 // (and we) do not reduce AVX/AVX2 masked intrinsics into LLVM masked
4506 // intrinsics.
4507 void handleAVXMaskedStore(IntrinsicInst &I) {
4508 assert(I.arg_size() == 3);
4509
4510 IRBuilder<> IRB(&I);
4511
4512 Value *Dst = I.getArgOperand(0);
4513 assert(Dst->getType()->isPointerTy() && "Destination is not a pointer!");
4514
4515 Value *Mask = I.getArgOperand(1);
4516 assert(isa<VectorType>(Mask->getType()) && "Mask is not a vector!");
4517
4518 Value *Src = I.getArgOperand(2);
4519 assert(isa<VectorType>(Src->getType()) && "Source is not a vector!");
4520
4521 const Align Alignment = Align(1);
4522
4523 Value *SrcShadow = getShadow(Src);
4524
4525 if (Opts.msan_check_access_address) {
4526 insertCheckShadowOf(Dst, &I);
4527 insertCheckShadowOf(Mask, &I);
4528 }
4529
4530 Value *DstShadowPtr;
4531 Value *DstOriginPtr;
4532 std::tie(DstShadowPtr, DstOriginPtr) = getShadowOriginPtr(
4533 Dst, IRB, SrcShadow->getType(), Alignment, /*isStore*/ true);
4534
4535 SmallVector<Value *, 2> ShadowArgs;
4536 ShadowArgs.append(1, DstShadowPtr);
4537 ShadowArgs.append(1, Mask);
4538 // The intrinsic may require floating-point but shadows can be arbitrary
4539 // bit patterns, of which some would be interpreted as "invalid"
4540 // floating-point values (NaN etc.); we assume the intrinsic will happily
4541 // copy them.
4542 ShadowArgs.append(1, IRB.CreateBitCast(SrcShadow, Src->getType()));
4543
4544 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
4545 IRB.getVoidTy(), I.getIntrinsicID(), ShadowArgs);
4546 setShadow(&I, CI);
4547
4548 if (!MS.TrackOrigins)
4549 return;
4550
4551 // Approximation only
4552 auto &DL = F.getDataLayout();
4553 paintOrigin(IRB, getOrigin(Src), DstOriginPtr,
4554 DL.getTypeStoreSize(SrcShadow->getType()),
4555 std::max(Alignment, kMinOriginAlignment));
4556 }
4557
4558 // e.g., <8 x float> @llvm.x86.avx.maskload.ps.256(ptr, <8 x i32>)
4559 // return src mask
4560 //
4561 // Masked-off values are replaced with 0, which conveniently also represents
4562 // initialized memory.
4563 //
4564 // AVX512 masked stores are lowered to Intrinsic::masked_load and are handled
4565 // by handleMaskedStore.
4566 //
4567 // We do not combine this with handleMaskedLoad; see comment in
4568 // handleAVXMaskedStore for the rationale.
4569 //
4570 // This is subtly different than handleIntrinsicByApplyingToShadow(I, 1)
4571 // because we need to apply getShadowOriginPtr, not getShadow, to the first
4572 // parameter.
4573 void handleAVXMaskedLoad(IntrinsicInst &I) {
4574 assert(I.arg_size() == 2);
4575
4576 IRBuilder<> IRB(&I);
4577
4578 Value *Src = I.getArgOperand(0);
4579 assert(Src->getType()->isPointerTy() && "Source is not a pointer!");
4580
4581 Value *Mask = I.getArgOperand(1);
4582 assert(isa<VectorType>(Mask->getType()) && "Mask is not a vector!");
4583
4584 const Align Alignment = Align(1);
4585
4586 if (Opts.msan_check_access_address) {
4587 insertCheckShadowOf(Mask, &I);
4588 }
4589
4590 Type *SrcShadowTy = getShadowTy(Src);
4591 Value *SrcShadowPtr, *SrcOriginPtr;
4592 std::tie(SrcShadowPtr, SrcOriginPtr) =
4593 getShadowOriginPtr(Src, IRB, SrcShadowTy, Alignment, /*isStore*/ false);
4594
4595 SmallVector<Value *, 2> ShadowArgs;
4596 ShadowArgs.append(1, SrcShadowPtr);
4597 ShadowArgs.append(1, Mask);
4598
4599 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
4600 I.getType(), I.getIntrinsicID(), ShadowArgs);
4601 // The AVX masked load intrinsics do not have integer variants. We use the
4602 // floating-point variants, which will happily copy the shadows even if
4603 // they are interpreted as "invalid" floating-point values (NaN etc.).
4604 setShadow(&I, IRB.CreateBitCast(CI, getShadowTy(&I)));
4605
4606 if (!MS.TrackOrigins)
4607 return;
4608
4609 // The "pass-through" value is always zero (initialized). To the extent
4610 // that that results in initialized aligned 4-byte chunks, the origin value
4611 // is ignored. It is therefore correct to simply copy the origin from src.
4612 Value *PtrSrcOrigin = IRB.CreateLoad(MS.OriginTy, SrcOriginPtr);
4613 setOrigin(&I, PtrSrcOrigin);
4614 }
4615
4616 // Test whether the mask indices are initialized, only checking the bits that
4617 // are actually used.
4618 //
4619 // e.g., if Idx is <32 x i16>, only (log2(32) == 5) bits of each index are
4620 // used/checked.
4621 void maskedCheckAVXIndexShadow(IRBuilder<> &IRB, Value *Idx, Instruction *I) {
4622 assert(isFixedIntVector(Idx));
4623 auto IdxVectorSize =
4624 cast<FixedVectorType>(Idx->getType())->getNumElements();
4625 assert(isPowerOf2_64(IdxVectorSize));
4626
4627 // Compiler isn't smart enough, let's help it
4628 if (isa<Constant>(Idx))
4629 return;
4630
4631 auto *IdxShadow = getShadow(Idx);
4632 Value *Truncated = IRB.CreateTrunc(
4633 IdxShadow,
4634 FixedVectorType::get(Type::getIntNTy(*MS.C, Log2_64(IdxVectorSize)),
4635 IdxVectorSize));
4636 insertCheckShadow(Truncated, getOrigin(Idx), I);
4637 }
4638
4639 // Instrument AVX permutation intrinsic.
4640 // We apply the same permutation (argument index 1) to the shadow.
4641 void handleAVXVpermilvar(IntrinsicInst &I) {
4642 IRBuilder<> IRB(&I);
4643 Value *Shadow = getShadow(&I, 0);
4644 maskedCheckAVXIndexShadow(IRB, I.getArgOperand(1), &I);
4645
4646 // Shadows are integer-ish types but some intrinsics require a
4647 // different (e.g., floating-point) type.
4648 Shadow = IRB.CreateBitCast(Shadow, I.getArgOperand(0)->getType());
4649 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
4650 I.getType(), I.getIntrinsicID(), {Shadow, I.getArgOperand(1)});
4651
4652 setShadow(&I, IRB.CreateBitCast(CI, getShadowTy(&I)));
4653 setOriginForNaryOp(I);
4654 }
4655
4656 // Instrument AVX permutation intrinsic.
4657 // We apply the same permutation (argument index 1) to the shadows.
4658 void handleAVXVpermi2var(IntrinsicInst &I) {
4659 assert(I.arg_size() == 3);
4660 assert(isa<FixedVectorType>(I.getArgOperand(0)->getType()));
4661 assert(isa<FixedVectorType>(I.getArgOperand(1)->getType()));
4662 assert(isa<FixedVectorType>(I.getArgOperand(2)->getType()));
4663 [[maybe_unused]] auto ArgVectorSize =
4664 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getNumElements();
4665 assert(cast<FixedVectorType>(I.getArgOperand(1)->getType())
4666 ->getNumElements() == ArgVectorSize);
4667 assert(cast<FixedVectorType>(I.getArgOperand(2)->getType())
4668 ->getNumElements() == ArgVectorSize);
4669 assert(I.getArgOperand(0)->getType() == I.getArgOperand(2)->getType());
4670 assert(I.getType() == I.getArgOperand(0)->getType());
4671 assert(I.getArgOperand(1)->getType()->isIntOrIntVectorTy());
4672 IRBuilder<> IRB(&I);
4673 Value *AShadow = getShadow(&I, 0);
4674 Value *Idx = I.getArgOperand(1);
4675 Value *BShadow = getShadow(&I, 2);
4676
4677 maskedCheckAVXIndexShadow(IRB, Idx, &I);
4678
4679 // Shadows are integer-ish types but some intrinsics require a
4680 // different (e.g., floating-point) type.
4681 AShadow = IRB.CreateBitCast(AShadow, I.getArgOperand(0)->getType());
4682 BShadow = IRB.CreateBitCast(BShadow, I.getArgOperand(2)->getType());
4683 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
4684 I.getType(), I.getIntrinsicID(), {AShadow, Idx, BShadow});
4685 setShadow(&I, IRB.CreateBitCast(CI, getShadowTy(&I)));
4686 setOriginForNaryOp(I);
4687 }
4688
4689 [[maybe_unused]] static bool isFixedIntVectorTy(const Type *T) {
4690 return isa<FixedVectorType>(T) && T->isIntOrIntVectorTy();
4691 }
4692
4693 [[maybe_unused]] static bool isFixedFPVectorTy(const Type *T) {
4694 return isa<FixedVectorType>(T) && T->isFPOrFPVectorTy();
4695 }
4696
4697 [[maybe_unused]] static bool isFixedIntVector(const Value *V) {
4698 return isFixedIntVectorTy(V->getType());
4699 }
4700
4701 [[maybe_unused]] static bool isFixedFPVector(const Value *V) {
4702 return isFixedFPVectorTy(V->getType());
4703 }
4704
4705 // e.g., <16 x i32> @llvm.x86.avx512.mask.cvtps2dq.512
4706 // (<16 x float> a, <16 x i32> writethru, i16 mask,
4707 // i32 rounding)
4708 //
4709 // Inconveniently, some similar intrinsics have a different operand order:
4710 // <16 x i16> @llvm.x86.avx512.mask.vcvtps2ph.512
4711 // (<16 x float> a, i32 rounding, <16 x i16> writethru,
4712 // i16 mask)
4713 //
4714 // If the return type has more elements than A, the excess elements are
4715 // zeroed (and the corresponding shadow is initialized).
4716 // <8 x i16> @llvm.x86.avx512.mask.vcvtps2ph.128
4717 // (<4 x float> a, i32 rounding, <8 x i16> writethru,
4718 // i8 mask)
4719 //
4720 // dst[i] = mask[i] ? convert(a[i]) : writethru[i]
4721 // dst_shadow[i] = mask[i] ? all_or_nothing(a_shadow[i]) : writethru_shadow[i]
4722 // where all_or_nothing(x) is fully uninitialized if x has any
4723 // uninitialized bits
4724 void handleAVX512VectorConvertFPToInt(IntrinsicInst &I, bool LastMask) {
4725 IRBuilder<> IRB(&I);
4726
4727 assert(I.arg_size() == 4);
4728 Value *A = I.getOperand(0);
4729 Value *WriteThrough;
4730 Value *Mask;
4732 if (LastMask) {
4733 WriteThrough = I.getOperand(2);
4734 Mask = I.getOperand(3);
4735 RoundingMode = I.getOperand(1);
4736 } else {
4737 WriteThrough = I.getOperand(1);
4738 Mask = I.getOperand(2);
4739 RoundingMode = I.getOperand(3);
4740 }
4741
4742 assert(isFixedFPVector(A));
4743 assert(isFixedIntVector(WriteThrough));
4744
4745 unsigned ANumElements =
4746 cast<FixedVectorType>(A->getType())->getNumElements();
4747 [[maybe_unused]] unsigned WriteThruNumElements =
4748 cast<FixedVectorType>(WriteThrough->getType())->getNumElements();
4749 assert(ANumElements == WriteThruNumElements ||
4750 ANumElements * 2 == WriteThruNumElements);
4751
4752 assert(Mask->getType()->isIntegerTy());
4753 unsigned MaskNumElements = Mask->getType()->getScalarSizeInBits();
4754 assert(ANumElements == MaskNumElements ||
4755 ANumElements * 2 == MaskNumElements);
4756
4757 assert(WriteThruNumElements == MaskNumElements);
4758
4759 // Some bits of the mask may be unused, though it's unusual to have partly
4760 // uninitialized bits.
4761 insertCheckShadowOf(Mask, &I);
4762
4763 assert(RoundingMode->getType()->isIntegerTy());
4764 // Only some bits of the rounding mode are used, though it's very
4765 // unusual to have uninitialized bits there (more commonly, it's a
4766 // constant).
4767 insertCheckShadowOf(RoundingMode, &I);
4768
4769 assert(I.getType() == WriteThrough->getType());
4770
4771 Value *AShadow = getShadow(A);
4772 AShadow = maybeExtendVectorShadowWithZeros(AShadow, I);
4773
4774 if (ANumElements * 2 == MaskNumElements) {
4775 // Ensure that the irrelevant bits of the mask are zero, hence selecting
4776 // from the zeroed shadow instead of the writethrough's shadow.
4777 Mask =
4778 IRB.CreateTrunc(Mask, IRB.getIntNTy(ANumElements), "_ms_mask_trunc");
4779 Mask =
4780 IRB.CreateZExt(Mask, IRB.getIntNTy(MaskNumElements), "_ms_mask_zext");
4781 }
4782
4783 // Convert i16 mask to <16 x i1>
4784 Mask = IRB.CreateBitCast(
4785 Mask, FixedVectorType::get(IRB.getInt1Ty(), MaskNumElements),
4786 "_ms_mask_bitcast");
4787
4788 /// For floating-point to integer conversion, the output is:
4789 /// - fully uninitialized if *any* bit of the input is uninitialized
4790 /// - fully ininitialized if all bits of the input are ininitialized
4791 /// We apply the same principle on a per-element basis for vectors.
4792 ///
4793 /// We use the scalar width of the return type instead of A's.
4794 AShadow = IRB.CreateSExt(
4795 IRB.CreateICmpNE(AShadow, getCleanShadow(AShadow->getType())),
4796 getShadowTy(&I), "_ms_a_shadow");
4797
4798 Value *WriteThroughShadow = getShadow(WriteThrough);
4799 Value *Shadow = IRB.CreateSelect(Mask, AShadow, WriteThroughShadow,
4800 "_ms_writethru_select");
4801
4802 setShadow(&I, Shadow);
4803 setOriginForNaryOp(I);
4804 }
4805
4806 static SmallVector<int, 8> getPclmulMask(unsigned Width, bool OddElements) {
4807 SmallVector<int, 8> Mask;
4808 for (unsigned X = OddElements ? 1 : 0; X < Width; X += 2) {
4809 Mask.append(2, X);
4810 }
4811 return Mask;
4812 }
4813
4814 // Instrument pclmul intrinsics.
4815 // These intrinsics operate either on odd or on even elements of the input
4816 // vectors, depending on the constant in the 3rd argument, ignoring the rest.
4817 // Replace the unused elements with copies of the used ones, ex:
4818 // (0, 1, 2, 3) -> (0, 0, 2, 2) (even case)
4819 // or
4820 // (0, 1, 2, 3) -> (1, 1, 3, 3) (odd case)
4821 // and then apply the usual shadow combining logic.
4822 void handlePclmulIntrinsic(IntrinsicInst &I) {
4823 IRBuilder<> IRB(&I);
4824 unsigned Width =
4825 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getNumElements();
4826 assert(isa<ConstantInt>(I.getArgOperand(2)) &&
4827 "pclmul 3rd operand must be a constant");
4828 unsigned Imm = cast<ConstantInt>(I.getArgOperand(2))->getZExtValue();
4829 Value *Shuf0 = IRB.CreateShuffleVector(getShadow(&I, 0),
4830 getPclmulMask(Width, Imm & 0x01));
4831 Value *Shuf1 = IRB.CreateShuffleVector(getShadow(&I, 1),
4832 getPclmulMask(Width, Imm & 0x10));
4833 ShadowAndOriginCombiner SOC(this, IRB);
4834 SOC.Add(Shuf0, getOrigin(&I, 0));
4835 SOC.Add(Shuf1, getOrigin(&I, 1));
4836 SOC.Done(&I);
4837 }
4838
4839 // Instrument _mm_*_sd|ss intrinsics
4840 void handleUnarySdSsIntrinsic(IntrinsicInst &I) {
4841 IRBuilder<> IRB(&I);
4842 unsigned Width =
4843 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getNumElements();
4844 Value *First = getShadow(&I, 0);
4845 Value *Second = getShadow(&I, 1);
4846 // First element of second operand, remaining elements of first operand
4847 SmallVector<int, 16> Mask;
4848 Mask.push_back(Width);
4849 for (unsigned i = 1; i < Width; i++)
4850 Mask.push_back(i);
4851 Value *Shadow = IRB.CreateShuffleVector(First, Second, Mask);
4852
4853 setShadow(&I, Shadow);
4854 setOriginForNaryOp(I);
4855 }
4856
4857 void handleVtestIntrinsic(IntrinsicInst &I) {
4858 IRBuilder<> IRB(&I);
4859 Value *Shadow0 = getShadow(&I, 0);
4860 Value *Shadow1 = getShadow(&I, 1);
4861 Value *Or = IRB.CreateOr(Shadow0, Shadow1);
4862 Value *NZ = IRB.CreateICmpNE(Or, Constant::getNullValue(Or->getType()));
4863 Value *Scalar = convertShadowToScalar(NZ, IRB);
4864 Value *Shadow = IRB.CreateZExt(Scalar, getShadowTy(&I));
4865
4866 setShadow(&I, Shadow);
4867 setOriginForNaryOp(I);
4868 }
4869
4870 void handleBinarySdSsIntrinsic(IntrinsicInst &I) {
4871 IRBuilder<> IRB(&I);
4872 unsigned Width =
4873 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getNumElements();
4874 Value *First = getShadow(&I, 0);
4875 Value *Second = getShadow(&I, 1);
4876 Value *OrShadow = IRB.CreateOr(First, Second);
4877 // First element of both OR'd together, remaining elements of first operand
4878 SmallVector<int, 16> Mask;
4879 Mask.push_back(Width);
4880 for (unsigned i = 1; i < Width; i++)
4881 Mask.push_back(i);
4882 Value *Shadow = IRB.CreateShuffleVector(First, OrShadow, Mask);
4883
4884 setShadow(&I, Shadow);
4885 setOriginForNaryOp(I);
4886 }
4887
4888 // _mm_round_ps / _mm_round_ps.
4889 // Similar to maybeHandleSimpleNomemIntrinsic except
4890 // the second argument is guaranteed to be a constant integer.
4891 void handleRoundPdPsIntrinsic(IntrinsicInst &I) {
4892 assert(I.getArgOperand(0)->getType() == I.getType());
4893 assert(I.arg_size() == 2);
4894 assert(isa<ConstantInt>(I.getArgOperand(1)));
4895
4896 IRBuilder<> IRB(&I);
4897 ShadowAndOriginCombiner SC(this, IRB);
4898 SC.Add(I.getArgOperand(0));
4899 SC.Done(&I);
4900 }
4901
4902 // Instrument @llvm.abs intrinsic.
4903 //
4904 // e.g., i32 @llvm.abs.i32 (i32 <Src>, i1 <is_int_min_poison>)
4905 // <4 x i32> @llvm.abs.v4i32(<4 x i32> <Src>, i1 <is_int_min_poison>)
4906 void handleAbsIntrinsic(IntrinsicInst &I) {
4907 assert(I.arg_size() == 2);
4908 Value *Src = I.getArgOperand(0);
4909 Value *IsIntMinPoison = I.getArgOperand(1);
4910
4911 assert(I.getType()->isIntOrIntVectorTy());
4912
4913 assert(Src->getType() == I.getType());
4914
4915 assert(IsIntMinPoison->getType()->isIntegerTy());
4916 assert(IsIntMinPoison->getType()->getIntegerBitWidth() == 1);
4917
4918 IRBuilder<> IRB(&I);
4919 Value *SrcShadow = getShadow(Src);
4920
4921 APInt MinVal =
4922 APInt::getSignedMinValue(Src->getType()->getScalarSizeInBits());
4923 Value *MinValVec = ConstantInt::get(Src->getType(), MinVal);
4924 Value *SrcIsMin = IRB.CreateICmp(CmpInst::ICMP_EQ, Src, MinValVec);
4925
4926 Value *PoisonedShadow = getPoisonedShadow(Src);
4927 Value *PoisonedIfIntMinShadow =
4928 IRB.CreateSelect(SrcIsMin, PoisonedShadow, SrcShadow);
4929 Value *Shadow =
4930 IRB.CreateSelect(IsIntMinPoison, PoisonedIfIntMinShadow, SrcShadow);
4931
4932 setShadow(&I, Shadow);
4933 setOrigin(&I, getOrigin(&I, 0));
4934 }
4935
4936 void handleIsFpClass(IntrinsicInst &I) {
4937 IRBuilder<> IRB(&I);
4938 Value *Shadow = getShadow(&I, 0);
4939 setShadow(&I, IRB.CreateICmpNE(Shadow, getCleanShadow(Shadow)));
4940 setOrigin(&I, getOrigin(&I, 0));
4941 }
4942
4943 void handleArithmeticWithOverflow(IntrinsicInst &I) {
4944 IRBuilder<> IRB(&I);
4945 Value *Shadow0 = getShadow(&I, 0);
4946 Value *Shadow1 = getShadow(&I, 1);
4947 Value *ShadowElt0 = IRB.CreateOr(Shadow0, Shadow1);
4948 Value *ShadowElt1 =
4949 IRB.CreateICmpNE(ShadowElt0, getCleanShadow(ShadowElt0));
4950
4951 Value *Shadow = PoisonValue::get(getShadowTy(&I));
4952 Shadow = IRB.CreateInsertValue(Shadow, ShadowElt0, 0);
4953 Shadow = IRB.CreateInsertValue(Shadow, ShadowElt1, 1);
4954
4955 setShadow(&I, Shadow);
4956 setOriginForNaryOp(I);
4957 }
4958
4959 void handleModfOrSincos(IntrinsicInst &I) {
4960 IRBuilder<> IRB(&I);
4961 Value *ArgShadow = getShadow(&I, 0);
4962 Value *Shadow = PoisonValue::get(getShadowTy(&I));
4963 Shadow = IRB.CreateInsertValue(Shadow, ArgShadow, 0);
4964 Shadow = IRB.CreateInsertValue(Shadow, ArgShadow, 1);
4965 setShadow(&I, Shadow);
4966 setOrigin(&I, getOrigin(&I, 0));
4967 }
4968
4969 Value *extractLowerShadow(IRBuilder<> &IRB, Value *V) {
4970 assert(isa<FixedVectorType>(V->getType()));
4971 assert(cast<FixedVectorType>(V->getType())->getNumElements() > 0);
4972 Value *Shadow = getShadow(V);
4973 return IRB.CreateExtractElement(Shadow,
4974 ConstantInt::get(IRB.getInt32Ty(), 0));
4975 }
4976
4977 // Handle llvm.x86.avx512.mask.pmov{,s,us}.*.{128,256,512}
4978 //
4979 // e.g., call <16 x i8> @llvm.x86.avx512.mask.pmov.qb.512
4980 // (<8 x i64>, <16 x i8>, i8)
4981 // A WriteThru Mask
4982 //
4983 // call <16 x i8> @llvm.x86.avx512.mask.pmovs.db.512
4984 // (<16 x i32>, <16 x i8>, i16)
4985 //
4986 // Dst[i] = Mask[i] ? truncate_or_saturate(A[i]) : WriteThru[i]
4987 // Dst_shadow[i] = Mask[i] ? truncate(A_shadow[i]) : WriteThru_shadow[i]
4988 //
4989 // If Dst has more elements than A, the excess elements are zeroed (and the
4990 // corresponding shadow is initialized).
4991 //
4992 // Note: for PMOV (truncation), handleIntrinsicByApplyingToShadow is precise
4993 // and is much faster than this handler.
4994 void handleAVX512VectorDownConvert(IntrinsicInst &I) {
4995 IRBuilder<> IRB(&I);
4996
4997 assert(I.arg_size() == 3);
4998 Value *A = I.getOperand(0);
4999 Value *WriteThrough = I.getOperand(1);
5000 Value *Mask = I.getOperand(2);
5001
5002 assert(isFixedIntVector(A));
5003 assert(isFixedIntVector(WriteThrough));
5004
5005 unsigned ANumElements =
5006 cast<FixedVectorType>(A->getType())->getNumElements();
5007 unsigned OutputNumElements =
5008 cast<FixedVectorType>(WriteThrough->getType())->getNumElements();
5009 assert(ANumElements == OutputNumElements ||
5010 ANumElements * 2 == OutputNumElements);
5011 // N.B. some PMOV{,S,US} instructions have a 4x or even 8x ratio in the
5012 // number of elements e.g.,
5013 // <16 x i8> @llvm.x86.avx512.mask.pmovs.qb.256
5014 // (<4 x i64>, <16 x i8>, i8)
5015 // <16 x i8> @llvm.x86.avx512.mask.pmovs.qb.128
5016 // (<2 x i64>, <16 x i8>, i8)
5017 // However, we currently handle those elsewhere.
5018
5019 assert(Mask->getType()->isIntegerTy());
5020 insertCheckShadowOf(Mask, &I);
5021
5022 // The mask has 1 bit per element of A, but a minimum of 8 bits.
5023 if (Mask->getType()->getScalarSizeInBits() == 8 && OutputNumElements < 8)
5024 Mask = IRB.CreateTrunc(Mask, Type::getIntNTy(*MS.C, OutputNumElements));
5025 assert(Mask->getType()->getScalarSizeInBits() == ANumElements);
5026
5027 assert(I.getType() == WriteThrough->getType());
5028
5029 // Widen the mask, if necessary, to have one bit per element of the output
5030 // vector.
5031 // We want the extra bits to have '1's, so that the CreateSelect will
5032 // select the values from AShadow instead of WriteThroughShadow ("maskless"
5033 // versions of the intrinsics are sometimes implemented using an all-1's
5034 // mask and an undefined value for WriteThroughShadow). We accomplish this
5035 // by using bitwise NOT before and after the ZExt.
5036 if (ANumElements != OutputNumElements) {
5037 Mask = IRB.CreateNot(Mask);
5038 Mask = IRB.CreateZExt(Mask, Type::getIntNTy(*MS.C, OutputNumElements),
5039 "_ms_widen_mask");
5040 Mask = IRB.CreateNot(Mask);
5041 }
5042 Mask = IRB.CreateBitCast(
5043 Mask, FixedVectorType::get(IRB.getInt1Ty(), OutputNumElements));
5044
5045 Value *AShadow = getShadow(A);
5046
5047 // The return type might have more elements than the input.
5048 // Temporarily shrink the return type's number of elements.
5049 VectorType *ShadowType = maybeShrinkVectorShadowType(A, I);
5050
5051 // PMOV truncates; PMOVS/PMOVUS uses signed/unsigned saturation.
5052 // This handler treats them all as truncation, which leads to some rare
5053 // false positives in the cases where the truncated bytes could
5054 // unambiguously saturate the value e.g., if A = ??????10 ????????
5055 // (big-endian), the unsigned saturated byte conversion is 11111111 i.e.,
5056 // fully defined, but the truncated byte is ????????.
5057 //
5058 // TODO: use GetMinMaxUnsigned() to handle saturation precisely.
5059 AShadow = IRB.CreateTrunc(AShadow, ShadowType, "_ms_trunc_shadow");
5060 AShadow = maybeExtendVectorShadowWithZeros(AShadow, I);
5061
5062 Value *WriteThroughShadow = getShadow(WriteThrough);
5063
5064 Value *Shadow = IRB.CreateSelect(Mask, AShadow, WriteThroughShadow);
5065 setShadow(&I, Shadow);
5066 setOriginForNaryOp(I);
5067 }
5068
5069 // Handle llvm.x86.avx512.* instructions that take vector(s) of floating-point
5070 // values and perform an operation whose shadow propagation should be handled
5071 // as all-or-nothing [*], with masking provided by a vector and a mask
5072 // supplied as an integer.
5073 //
5074 // [*] if all bits of a vector element are initialized, the output is fully
5075 // initialized; otherwise, the output is fully uninitialized
5076 //
5077 // e.g., <16 x float> @llvm.x86.avx512.rsqrt14.ps.512
5078 // (<16 x float>, <16 x float>, i16)
5079 // A WriteThru Mask
5080 //
5081 // <2 x double> @llvm.x86.avx512.rcp14.pd.128
5082 // (<2 x double>, <2 x double>, i8)
5083 // A WriteThru Mask
5084 //
5085 // <8 x double> @llvm.x86.avx512.mask.rndscale.pd.512
5086 // (<8 x double>, i32, <8 x double>, i8, i32)
5087 // A Imm WriteThru Mask Rounding
5088 //
5089 // <16 x float> @llvm.x86.avx512.mask.scalef.ps.512
5090 // (<16 x float>, <16 x float>, <16 x float>, i16, i32)
5091 // WriteThru A B Mask Rnd
5092 //
5093 // All operands other than A, B, ..., and WriteThru (e.g., Mask, Imm,
5094 // Rounding) must be fully initialized.
5095 //
5096 // Dst[i] = Mask[i] ? some_op(A[i], B[i], ...)
5097 // : WriteThru[i]
5098 // Dst_shadow[i] = Mask[i] ? all_or_nothing(A_shadow[i] | B_shadow[i] | ...)
5099 // : WriteThru_shadow[i]
5100 void handleAVX512VectorGenericMaskedFP(IntrinsicInst &I,
5101 SmallVector<unsigned, 4> DataIndices,
5102 unsigned WriteThruIndex,
5103 unsigned MaskIndex) {
5104 IRBuilder<> IRB(&I);
5105
5106 unsigned NumArgs = I.arg_size();
5107
5108 assert(WriteThruIndex < NumArgs);
5109 assert(MaskIndex < NumArgs);
5110 assert(WriteThruIndex != MaskIndex);
5111 Value *WriteThru = I.getOperand(WriteThruIndex);
5112
5113 unsigned OutputNumElements =
5114 cast<FixedVectorType>(WriteThru->getType())->getNumElements();
5115
5116 assert(DataIndices.size() > 0);
5117
5118 bool isData[16] = {false};
5119 assert(NumArgs <= 16);
5120 for (unsigned i : DataIndices) {
5121 assert(i < NumArgs);
5122 assert(i != WriteThruIndex);
5123 assert(i != MaskIndex);
5124
5125 isData[i] = true;
5126
5127 Value *A = I.getOperand(i);
5128 assert(isFixedFPVector(A));
5129 [[maybe_unused]] unsigned ANumElements =
5130 cast<FixedVectorType>(A->getType())->getNumElements();
5131 assert(ANumElements == OutputNumElements);
5132 }
5133
5134 Value *Mask = I.getOperand(MaskIndex);
5135
5136 assert(isFixedFPVector(WriteThru));
5137
5138 for (unsigned i = 0; i < NumArgs; ++i) {
5139 if (!isData[i] && i != WriteThruIndex) {
5140 // Imm, Mask, Rounding etc. are "control" data, hence we require that
5141 // they be fully initialized.
5142 assert(I.getOperand(i)->getType()->isIntegerTy());
5143 insertCheckShadowOf(I.getOperand(i), &I);
5144 }
5145 }
5146
5147 // The mask has 1 bit per element of A, but a minimum of 8 bits.
5148 if (Mask->getType()->getScalarSizeInBits() == 8 && OutputNumElements < 8)
5149 Mask = IRB.CreateTrunc(Mask, Type::getIntNTy(*MS.C, OutputNumElements));
5150 assert(Mask->getType()->getScalarSizeInBits() == OutputNumElements);
5151
5152 assert(I.getType() == WriteThru->getType());
5153
5154 Mask = IRB.CreateBitCast(
5155 Mask, FixedVectorType::get(IRB.getInt1Ty(), OutputNumElements));
5156
5157 Value *DataShadow = nullptr;
5158 for (unsigned i : DataIndices) {
5159 Value *A = I.getOperand(i);
5160 if (DataShadow)
5161 DataShadow = IRB.CreateOr(DataShadow, getShadow(A));
5162 else
5163 DataShadow = getShadow(A);
5164 }
5165
5166 // All-or-nothing shadow
5167 DataShadow =
5168 IRB.CreateSExt(IRB.CreateICmpNE(DataShadow, getCleanShadow(DataShadow)),
5169 DataShadow->getType());
5170
5171 Value *WriteThruShadow = getShadow(WriteThru);
5172
5173 Value *Shadow = IRB.CreateSelect(Mask, DataShadow, WriteThruShadow);
5174 setShadow(&I, Shadow);
5175
5176 setOriginForNaryOp(I);
5177 }
5178
5179 // AVX512 Floating-Point Classification
5180 //
5181 // e.g.,
5182 // - < 8 x i1> @llvm.x86.avx512.fpclass.pd.512
5183 // (<8 x double> %input, i32 %classifiers)
5184 // - <16 x i1> @llvm.x86.avx512.fpclass.ps.512
5185 // (<16 x float> %input, i32 %classifiers)
5186 void handleAVX512FPClass(IntrinsicInst &I) {
5187 IRBuilder<> IRB(&I);
5188
5189 assert(I.arg_size() == 2);
5190
5191 Value *Input = I.getOperand(0);
5192 assert(isFixedFPVector(Input));
5193 [[maybe_unused]] FixedVectorType *InputType = cast<FixedVectorType>(Input->getType());
5194
5195 Value *Classifiers = I.getOperand(1);
5196 assert(isa<ConstantInt>(Classifiers));
5197 // No shadow check needed for constants
5198
5199 assert(isFixedIntVectorTy(I.getType()));
5200 FixedVectorType *OutputType = cast<FixedVectorType>(I.getType());
5201 assert(OutputType->getScalarSizeInBits() == 1);
5202
5203 assert(OutputType->getNumElements() == InputType->getNumElements());
5204
5205 Value *OutputShadow;
5206 if (cast<ConstantInt>(Classifiers)->isZero())
5207 // Each bit specifies whether a particular classifier is enabled.
5208 // If Classifiers == 0, the output is trivially known to be zero, thus
5209 // the output is fully initialized.
5210 OutputShadow = getCleanShadow(OutputType);
5211 else
5212 // Approximate each bit of the output shadow based on whether the
5213 // corresponding input element is fully initialized. It is only
5214 // approximate because some classifications do not rely on all bits of
5215 // the input element.
5216 OutputShadow = IRB.CreateICmpNE(getShadow(Input), getCleanShadow(Input));
5217
5218 setShadow(&I, OutputShadow);
5219
5220 setOriginForNaryOp(I);
5221 }
5222
5223 // For sh.* compiler intrinsics:
5224 // llvm.x86.avx512fp16.mask.{add/sub/mul/div/max/min}.sh.round
5225 // (<8 x half>, <8 x half>, <8 x half>, i8, i32)
5226 // A B WriteThru Mask RoundingMode
5227 //
5228 // DstShadow[0] = Mask[0] ? (AShadow[0] | BShadow[0]) : WriteThruShadow[0]
5229 // DstShadow[1..7] = AShadow[1..7]
5230 void visitGenericScalarHalfwordInst(IntrinsicInst &I) {
5231 IRBuilder<> IRB(&I);
5232
5233 assert(I.arg_size() == 5);
5234 Value *A = I.getOperand(0);
5235 Value *B = I.getOperand(1);
5236 Value *WriteThrough = I.getOperand(2);
5237 Value *Mask = I.getOperand(3);
5238 Value *RoundingMode = I.getOperand(4);
5239
5240 // Technically, we could probably just check whether the LSB is
5241 // initialized, but intuitively it feels like a partly uninitialized mask
5242 // is unintended, and we should warn the user immediately.
5243 insertCheckShadowOf(Mask, &I);
5244 insertCheckShadowOf(RoundingMode, &I);
5245
5246 assert(isa<FixedVectorType>(A->getType()));
5247 unsigned NumElements =
5248 cast<FixedVectorType>(A->getType())->getNumElements();
5249 assert(NumElements == 8);
5250 assert(A->getType() == B->getType());
5251 assert(B->getType() == WriteThrough->getType());
5252 assert(Mask->getType()->getPrimitiveSizeInBits() == NumElements);
5253 assert(RoundingMode->getType()->isIntegerTy());
5254
5255 Value *ALowerShadow = extractLowerShadow(IRB, A);
5256 Value *BLowerShadow = extractLowerShadow(IRB, B);
5257
5258 Value *ABLowerShadow = IRB.CreateOr(ALowerShadow, BLowerShadow);
5259
5260 Value *WriteThroughLowerShadow = extractLowerShadow(IRB, WriteThrough);
5261
5262 Mask = IRB.CreateBitCast(
5263 Mask, FixedVectorType::get(IRB.getInt1Ty(), NumElements));
5264 Value *MaskLower =
5265 IRB.CreateExtractElement(Mask, ConstantInt::get(IRB.getInt32Ty(), 0));
5266
5267 Value *AShadow = getShadow(A);
5268 Value *DstLowerShadow =
5269 IRB.CreateSelect(MaskLower, ABLowerShadow, WriteThroughLowerShadow);
5270 Value *DstShadow = IRB.CreateInsertElement(
5271 AShadow, DstLowerShadow, ConstantInt::get(IRB.getInt32Ty(), 0),
5272 "_msprop");
5273
5274 setShadow(&I, DstShadow);
5275 setOriginForNaryOp(I);
5276 }
5277
5278 // Approximately handle AVX Galois Field Affine Transformation
5279 //
5280 // e.g.,
5281 // <16 x i8> @llvm.x86.vgf2p8affineqb.128(<16 x i8>, <16 x i8>, i8)
5282 // <32 x i8> @llvm.x86.vgf2p8affineqb.256(<32 x i8>, <32 x i8>, i8)
5283 // <64 x i8> @llvm.x86.vgf2p8affineqb.512(<64 x i8>, <64 x i8>, i8)
5284 // Out A x b
5285 // where A and x are packed matrices, b is a vector,
5286 // Out = A * x + b in GF(2)
5287 //
5288 // Multiplication in GF(2) is equivalent to bitwise AND. However, the matrix
5289 // computation also includes a parity calculation.
5290 //
5291 // For the bitwise AND of bits V1 and V2, the exact shadow is:
5292 // Out_Shadow = (V1_Shadow & V2_Shadow)
5293 // | (V1 & V2_Shadow)
5294 // | (V1_Shadow & V2 )
5295 //
5296 // We approximate the shadow of gf2p8affineqb using:
5297 // Out_Shadow = gf2p8affineqb(x_Shadow, A_shadow, 0)
5298 // | gf2p8affineqb(x, A_shadow, 0)
5299 // | gf2p8affineqb(x_Shadow, A, 0)
5300 // | set1_epi8(b_Shadow)
5301 //
5302 // This approximation has false negatives: if an intermediate dot-product
5303 // contains an even number of 1's, the parity is 0.
5304 // It has no false positives.
5305 void handleAVXGF2P8Affine(IntrinsicInst &I) {
5306 IRBuilder<> IRB(&I);
5307
5308 assert(I.arg_size() == 3);
5309 Value *A = I.getOperand(0);
5310 Value *X = I.getOperand(1);
5311 Value *B = I.getOperand(2);
5312
5313 assert(isFixedIntVector(A));
5314 assert(cast<VectorType>(A->getType())
5315 ->getElementType()
5316 ->getScalarSizeInBits() == 8);
5317
5318 assert(A->getType() == X->getType());
5319
5320 assert(B->getType()->isIntegerTy());
5321 assert(B->getType()->getScalarSizeInBits() == 8);
5322
5323 assert(I.getType() == A->getType());
5324
5325 Value *AShadow = getShadow(A);
5326 Value *XShadow = getShadow(X);
5327 Value *BZeroShadow = getCleanShadow(B);
5328
5329 Value *AShadowXShadow = IRB.CreateIntrinsic(
5330 I.getType(), I.getIntrinsicID(), {XShadow, AShadow, BZeroShadow});
5331 Value *AShadowX = IRB.CreateIntrinsic(I.getType(), I.getIntrinsicID(),
5332 {X, AShadow, BZeroShadow});
5333 Value *XShadowA = IRB.CreateIntrinsic(I.getType(), I.getIntrinsicID(),
5334 {XShadow, A, BZeroShadow});
5335
5336 unsigned NumElements = cast<FixedVectorType>(I.getType())->getNumElements();
5337 Value *BShadow = getShadow(B);
5338 Value *BBroadcastShadow = getCleanShadow(AShadow);
5339 // There is no LLVM IR intrinsic for _mm512_set1_epi8.
5340 // This loop generates a lot of LLVM IR, which we expect that CodeGen will
5341 // lower appropriately (e.g., VPBROADCASTB).
5342 // Besides, b is often a constant, in which case it is fully initialized.
5343 for (unsigned i = 0; i < NumElements; i++)
5344 BBroadcastShadow = IRB.CreateInsertElement(BBroadcastShadow, BShadow, i);
5345
5346 setShadow(&I, IRB.CreateOr(
5347 {AShadowXShadow, AShadowX, XShadowA, BBroadcastShadow}));
5348 setOriginForNaryOp(I);
5349 }
5350
5351 // Handle Arm NEON vector load intrinsics (vld*).
5352 //
5353 // The WithLane instructions (ld[234]lane) are similar to:
5354 // call {<4 x i32>, <4 x i32>, <4 x i32>}
5355 // @llvm.aarch64.neon.ld3lane.v4i32.p0
5356 // (<4 x i32> %L1, <4 x i32> %L2, <4 x i32> %L3, i64 %lane, ptr
5357 // %A)
5358 //
5359 // The non-WithLane instructions (ld[234], ld1x[234], ld[234]r) are similar
5360 // to:
5361 // call {<8 x i8>, <8 x i8>} @llvm.aarch64.neon.ld2.v8i8.p0(ptr %A)
5362 void handleNEONVectorLoad(IntrinsicInst &I, bool WithLane) {
5363 unsigned int numArgs = I.arg_size();
5364
5365 // Return type is a struct of vectors of integers or floating-point
5366 assert(I.getType()->isStructTy());
5367 [[maybe_unused]] StructType *RetTy = cast<StructType>(I.getType());
5368 assert(RetTy->getNumElements() > 0);
5370 RetTy->getElementType(0)->isFPOrFPVectorTy());
5371 for (unsigned int i = 0; i < RetTy->getNumElements(); i++)
5372 assert(RetTy->getElementType(i) == RetTy->getElementType(0));
5373
5374 if (WithLane) {
5375 // 2, 3 or 4 vectors, plus lane number, plus input pointer
5376 assert(4 <= numArgs && numArgs <= 6);
5377
5378 // Return type is a struct of the input vectors
5379 assert(RetTy->getNumElements() + 2 == numArgs);
5380 for (unsigned int i = 0; i < RetTy->getNumElements(); i++)
5381 assert(I.getArgOperand(i)->getType() == RetTy->getElementType(0));
5382 } else {
5383 assert(numArgs == 1);
5384 }
5385
5386 IRBuilder<> IRB(&I);
5387
5388 SmallVector<Value *, 6> ShadowArgs;
5389 if (WithLane) {
5390 for (unsigned int i = 0; i < numArgs - 2; i++)
5391 ShadowArgs.push_back(getShadow(I.getArgOperand(i)));
5392
5393 // Lane number, passed verbatim
5394 Value *LaneNumber = I.getArgOperand(numArgs - 2);
5395 ShadowArgs.push_back(LaneNumber);
5396
5397 // TODO: blend shadow of lane number into output shadow?
5398 insertCheckShadowOf(LaneNumber, &I);
5399 }
5400
5401 Value *Src = I.getArgOperand(numArgs - 1);
5402 assert(Src->getType()->isPointerTy() && "Source is not a pointer!");
5403
5404 Type *SrcShadowTy = getShadowTy(Src);
5405 auto [SrcShadowPtr, SrcOriginPtr] =
5406 getShadowOriginPtr(Src, IRB, SrcShadowTy, Align(1), /*isStore*/ false);
5407 ShadowArgs.push_back(SrcShadowPtr);
5408
5409 // The NEON vector load instructions handled by this function all have
5410 // integer variants. It is easier to use those rather than trying to cast
5411 // a struct of vectors of floats into a struct of vectors of integers.
5412 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
5413 getShadowTy(&I), I.getIntrinsicID(), ShadowArgs);
5414 setShadow(&I, CI);
5415
5416 if (!MS.TrackOrigins)
5417 return;
5418
5419 Value *PtrSrcOrigin = IRB.CreateLoad(MS.OriginTy, SrcOriginPtr);
5420 setOrigin(&I, PtrSrcOrigin);
5421 }
5422
5423 /// Handle Arm NEON vector store intrinsics (vst{2,3,4}, vst1x_{2,3,4},
5424 /// and vst{2,3,4}lane).
5425 ///
5426 /// Arm NEON vector store intrinsics have the output address (pointer) as the
5427 /// last argument, with the initial arguments being the inputs (and lane
5428 /// number for vst{2,3,4}lane). They return void.
5429 ///
5430 /// - st4 interleaves the output e.g., st4 (inA, inB, inC, inD, outP) writes
5431 /// abcdabcdabcdabcd... into *outP
5432 /// - st1_x4 is non-interleaved e.g., st1_x4 (inA, inB, inC, inD, outP)
5433 /// writes aaaa...bbbb...cccc...dddd... into *outP
5434 /// - st4lane has arguments of (inA, inB, inC, inD, lane, outP)
5435 /// These instructions can all be instrumented with essentially the same
5436 /// MSan logic, simply by applying the corresponding intrinsic to the shadow.
5437 void handleNEONVectorStoreIntrinsic(IntrinsicInst &I, bool useLane) {
5438 IRBuilder<> IRB(&I);
5439
5440 // Don't use getNumOperands() because it includes the callee
5441 int numArgOperands = I.arg_size();
5442
5443 // The last arg operand is the output (pointer)
5444 assert(numArgOperands >= 1);
5445 Value *Addr = I.getArgOperand(numArgOperands - 1);
5446 assert(Addr->getType()->isPointerTy());
5447 int skipTrailingOperands = 1;
5448
5449 if (Opts.msan_check_access_address)
5450 insertCheckShadowOf(Addr, &I);
5451
5452 // Second-last operand is the lane number (for vst{2,3,4}lane)
5453 if (useLane) {
5454 skipTrailingOperands++;
5455 assert(numArgOperands >= static_cast<int>(skipTrailingOperands));
5457 I.getArgOperand(numArgOperands - skipTrailingOperands)->getType()));
5458 }
5459
5460 SmallVector<Value *, 8> ShadowArgs;
5461 // All the initial operands are the inputs
5462 for (int i = 0; i < numArgOperands - skipTrailingOperands; i++) {
5463 assert(isa<FixedVectorType>(I.getArgOperand(i)->getType()));
5464 Value *Shadow = getShadow(&I, i);
5465 ShadowArgs.append(1, Shadow);
5466 }
5467
5468 // MSan's GetShadowTy assumes the LHS is the type we want the shadow for
5469 // e.g., for:
5470 // [[TMP5:%.*]] = bitcast <16 x i8> [[TMP2]] to i128
5471 // we know the type of the output (and its shadow) is <16 x i8>.
5472 //
5473 // Arm NEON VST is unusual because the last argument is the output address:
5474 // define void @st2_16b(<16 x i8> %A, <16 x i8> %B, ptr %P) {
5475 // call void @llvm.aarch64.neon.st2.v16i8.p0
5476 // (<16 x i8> [[A]], <16 x i8> [[B]], ptr [[P]])
5477 // and we have no type information about P's operand. We must manually
5478 // compute the type (<16 x i8> x 2).
5479 FixedVectorType *OutputVectorTy = FixedVectorType::get(
5480 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getElementType(),
5481 cast<FixedVectorType>(I.getArgOperand(0)->getType())->getNumElements() *
5482 (numArgOperands - skipTrailingOperands));
5483 Type *OutputShadowTy = getShadowTy(OutputVectorTy);
5484
5485 if (useLane)
5486 ShadowArgs.append(1,
5487 I.getArgOperand(numArgOperands - skipTrailingOperands));
5488
5489 Value *OutputShadowPtr, *OutputOriginPtr;
5490 // AArch64 NEON does not need alignment (unless OS requires it)
5491 std::tie(OutputShadowPtr, OutputOriginPtr) = getShadowOriginPtr(
5492 Addr, IRB, OutputShadowTy, Align(1), /*isStore*/ true);
5493 ShadowArgs.append(1, OutputShadowPtr);
5494
5495 CallInst *CI = IRB.CreateIntrinsicWithoutFolding(
5496 IRB.getVoidTy(), I.getIntrinsicID(), ShadowArgs);
5497 setShadow(&I, CI);
5498
5499 if (MS.TrackOrigins) {
5500 // TODO: if we modelled the vst* instruction more precisely, we could
5501 // more accurately track the origins (e.g., if both inputs are
5502 // uninitialized for vst2, we currently blame the second input, even
5503 // though part of the output depends only on the first input).
5504 //
5505 // This is particularly imprecise for vst{2,3,4}lane, since only one
5506 // lane of each input is actually copied to the output.
5507 OriginCombiner OC(this, IRB);
5508 for (int i = 0; i < numArgOperands - skipTrailingOperands; i++)
5509 OC.Add(I.getArgOperand(i));
5510
5511 const DataLayout &DL = F.getDataLayout();
5512 OC.DoneAndStoreOrigin(DL.getTypeStoreSize(OutputVectorTy),
5513 OutputOriginPtr);
5514 }
5515 }
5516
5517 // Integer matrix multiplication:
5518 // - <4 x i32> @llvm.aarch64.neon.{s,u,us}mmla.v4i32.v16i8
5519 // (<4 x i32> %R, <16 x i8> %A, <16 x i8> %B)
5520 // - <4 x i32> is a 2x2 matrix
5521 // - <16 x i8> %A and %B are 2x8 and 8x2 matrices respectively
5522 //
5523 // Floating-point matrix multiplication:
5524 // - <4 x float> @llvm.aarch64.neon.bfmmla
5525 // (<4 x float> %R, <8 x bfloat> %A, <8 x bfloat> %B)
5526 // - <4 x float> is a 2x2 matrix
5527 // - <8 x bfloat> %A and %B are 2x4 and 4x2 matrices respectively
5528 //
5529 // The general shadow propagation approach is:
5530 // 1) get the shadows of the input matrices %A and %B
5531 // 2) map each shadow value to 0x1 if the corresponding value is fully
5532 // initialized, and 0x0 otherwise
5533 // 3) perform a matrix multiplication on the shadows of %A and %B [*].
5534 // The output will be a 2x2 matrix. For each element, a value of 0x8
5535 // (for {s,u,us}mmla) or 0x4 (for bfmmla) means all the corresponding
5536 // inputs were clean; if so, set the shadow to zero, otherwise set to -1.
5537 // 4) blend in the shadow of %R
5538 //
5539 // [*] Since shadows are integral, the obvious approach is to always apply
5540 // ummla to the shadows. Unfortunately, Armv8.2+bf16 supports bfmmla,
5541 // but not ummla. Thus, for bfmmla, our instrumentation reuses bfmmla.
5542 //
5543 // TODO: consider allowing multiplication of zero with an uninitialized value
5544 // to result in an initialized value.
5545 void handleNEONMatrixMultiply(IntrinsicInst &I) {
5546 IRBuilder<> IRB(&I);
5547
5548 assert(I.arg_size() == 3);
5549 Value *R = I.getArgOperand(0);
5550 Value *A = I.getArgOperand(1);
5551 Value *B = I.getArgOperand(2);
5552
5553 assert(I.getType() == R->getType());
5554
5555 assert(isa<FixedVectorType>(R->getType()));
5556 assert(isa<FixedVectorType>(A->getType()));
5557 assert(isa<FixedVectorType>(B->getType()));
5558
5559 FixedVectorType *RTy = cast<FixedVectorType>(R->getType());
5560 FixedVectorType *ATy = cast<FixedVectorType>(A->getType());
5561 FixedVectorType *BTy = cast<FixedVectorType>(B->getType());
5562 assert(ATy->getElementType() == BTy->getElementType());
5563
5564 if (RTy->getElementType()->isIntegerTy()) {
5565 // <4 x i32> @llvm.aarch64.neon.ummla.v4i32.v16i8
5566 // (<4 x i32> %R, <16 x i8> %X, <16 x i8> %Y)
5567 assert(RTy == FixedVectorType::get(IntegerType::get(*MS.C, 32), 4));
5568 assert(ATy == FixedVectorType::get(IntegerType::get(*MS.C, 8), 16));
5569 assert(BTy == FixedVectorType::get(IntegerType::get(*MS.C, 8), 16));
5570 } else {
5571 // <4 x float> @llvm.aarch64.neon.bfmmla
5572 // (<4 x float> %R, <8 x bfloat> %X, <8 x bfloat> %Y)
5573 assert(RTy == FixedVectorType::get(Type::getFloatTy(*MS.C), 4));
5574 assert(ATy == FixedVectorType::get(Type::getBFloatTy(*MS.C), 8));
5575 assert(BTy == FixedVectorType::get(Type::getBFloatTy(*MS.C), 8));
5576 }
5577
5578 Value *ShadowR = getShadow(&I, 0);
5579 Value *ShadowA = getShadow(&I, 1);
5580 Value *ShadowB = getShadow(&I, 2);
5581
5582 Value *ShadowAB;
5583 Value *FullyInit;
5584
5585 if (RTy->getElementType()->isIntegerTy()) {
5586 // If the value is fully initialized, the shadow will be 000...001.
5587 // Otherwise, the shadow will be all zero.
5588 // (This is the opposite of how we typically handle shadows.)
5589 ShadowA = IRB.CreateZExt(IRB.CreateICmpEQ(ShadowA, getCleanShadow(ATy)),
5590 getShadowTy(ATy));
5591 ShadowB = IRB.CreateZExt(IRB.CreateICmpEQ(ShadowB, getCleanShadow(BTy)),
5592 getShadowTy(BTy));
5593 // TODO: the CreateSelect approach used below for floating-point is more
5594 // generic than CreateZExt. Investigate whether it is worthwhile
5595 // unifying the two approaches.
5596
5597 ShadowAB = IRB.CreateIntrinsic(RTy, Intrinsic::aarch64_neon_ummla,
5598 {getCleanShadow(RTy), ShadowA, ShadowB});
5599
5600 // ummla multiplies a 2x8 matrix with an 8x2 matrix. If all entries of the
5601 // input matrices are equal to 0x1, all entries of the output matrix will
5602 // be 0x8.
5603 FullyInit = ConstantVector::getSplat(
5604 RTy->getElementCount(), ConstantInt::get(RTy->getElementType(), 0x8));
5605
5606 ShadowAB = IRB.CreateICmpNE(ShadowAB, FullyInit);
5607 } else {
5609 ATy->getElementCount(), ConstantFP::get(ATy->getElementType(), 0));
5611 ATy->getElementCount(), ConstantFP::get(ATy->getElementType(), 1));
5612
5613 // As per the integer case, if the shadow is clean, we store 0x1,
5614 // otherwise we store 0x0 (the opposite of usual shadow arithmetic).
5615 ShadowA = IRB.CreateSelect(IRB.CreateICmpEQ(ShadowA, getCleanShadow(ATy)),
5616 ABOnes, ABZeros);
5617 ShadowB = IRB.CreateSelect(IRB.CreateICmpEQ(ShadowB, getCleanShadow(BTy)),
5618 ABOnes, ABZeros);
5619
5621 RTy->getElementCount(), ConstantFP::get(RTy->getElementType(), 0));
5622
5623 ShadowAB = IRB.CreateIntrinsic(RTy, Intrinsic::aarch64_neon_bfmmla,
5624 {RZeros, ShadowA, ShadowB});
5625
5626 // bfmmla multiplies a 2x4 matrix with an 4x2 matrix. If all entries of
5627 // the input matrices are equal to 0x1, all entries of the output matrix
5628 // will be 4.0. (To avoid floating-point error, we check if each entry
5629 // < 3.5.)
5630 FullyInit = ConstantVector::getSplat(
5631 RTy->getElementCount(), ConstantFP::get(RTy->getElementType(), 3.5));
5632
5633 // FCmpULT: "yields true if either operand is a QNAN or op1 is less than"
5634 // op2"
5635 ShadowAB = IRB.CreateFCmpULT(ShadowAB, FullyInit);
5636 }
5637
5638 ShadowR = IRB.CreateICmpNE(ShadowR, getCleanShadow(RTy));
5639 ShadowR = IRB.CreateOr(ShadowAB, ShadowR);
5640
5641 setShadow(&I, IRB.CreateSExt(ShadowR, getShadowTy(RTy)));
5642
5643 setOriginForNaryOp(I);
5644 }
5645
5646 /// Handle intrinsics by applying the intrinsic to the shadows.
5647 ///
5648 /// For example, this can be applied to the Arm NEON vector table intrinsics
5649 /// (tbl{1,2,3,4}).
5650 ///
5651 /// Typically, shadowIntrinsicID will be specified by the caller to be
5652 /// I.getIntrinsicID(), but the caller can choose to replace it with another
5653 /// intrinsic of the same type.
5654 ///
5655 /// The trailing arguments are passed verbatim to the intrinsic, though any
5656 /// uninitialized trailing arguments can also taint the shadow e.g., for an
5657 /// intrinsic with one trailing verbatim argument:
5658 /// out = intrinsic(var1, var2, opType)
5659 /// we compute:
5660 /// shadow[out] =
5661 /// intrinsic(shadow[var1], shadow[var2], opType) | shadow[opType]
5662 ///
5663 /// If an intrinsic is called with floating-point arguments, we will
5664 /// typically cast the shadows to floating-point, apply the intrinsic [*],
5665 /// then cast the result back to integer/shadow.
5666 ///
5667 /// In cases where we know the intrinsic is compatible with integer
5668 /// arguments, 'forceIntegerIntrinsic' will apply the integer variant, even
5669 /// if the arguments are floating-point, thus avoiding unnecessary casts
5670 /// e.g., if I is:
5671 /// <16 x float> @llvm.x86.avx512.mask.compress
5672 /// (<16 x float>, <16 x float>, <16 x i1> %mask)
5673 /// we would prefer to compute the shadows using:
5674 /// <16 x i32> @llvm.x86.avx512.mask.compress
5675 /// (<16 x i32>, <16 x i32>, <16 x i1> %mask)
5676 ///
5677 /// [*] CAUTION: this assumes that the intrinsic will handle arbitrary
5678 /// bit-patterns (for example, if the intrinsic accepts floats
5679 /// for var1, we require that it doesn't care if inputs are
5680 /// NaNs).
5681 ///
5682 /// The origin is approximated using setOriginForNaryOp.
5683 void handleIntrinsicByApplyingToShadow(IntrinsicInst &I,
5684 Intrinsic::ID shadowIntrinsicID,
5685 unsigned int trailingVerbatimArgs,
5686 bool forceIntegerIntrinsic) {
5687 IRBuilder<> IRB(&I);
5688
5689 assert(trailingVerbatimArgs < I.arg_size());
5690
5691 SmallVector<Value *, 8> ShadowArgs;
5692 // Don't use getNumOperands() because it includes the callee
5693 for (unsigned int i = 0; i < I.arg_size() - trailingVerbatimArgs; i++) {
5694 Value *Shadow = getShadow(&I, i);
5695
5696 if (forceIntegerIntrinsic)
5697 ShadowArgs.push_back(Shadow);
5698 else
5699 ShadowArgs.push_back(
5700 IRB.CreateBitCast(Shadow, I.getArgOperand(i)->getType()));
5701 }
5702
5703 for (unsigned int i = I.arg_size() - trailingVerbatimArgs; i < I.arg_size();
5704 i++) {
5705 Value *Arg = I.getArgOperand(i);
5706 if (forceIntegerIntrinsic)
5708 ShadowArgs.push_back(Arg);
5709 }
5710
5711 Value *CombinedShadow;
5712 if (forceIntegerIntrinsic) {
5713 CombinedShadow =
5714 IRB.CreateIntrinsic(getShadowTy(&I), shadowIntrinsicID, ShadowArgs);
5715 } else {
5716 Value *CI =
5717 IRB.CreateIntrinsic(I.getType(), shadowIntrinsicID, ShadowArgs);
5718 CombinedShadow = IRB.CreateBitCast(CI, getShadowTy(&I));
5719 }
5720
5721 // Combine the computed shadow with the shadow of trailing args
5722 for (unsigned int i = I.arg_size() - trailingVerbatimArgs; i < I.arg_size();
5723 i++) {
5724 Value *Shadow =
5725 CreateShadowCast(IRB, getShadow(&I, i), CombinedShadow->getType());
5726 CombinedShadow = IRB.CreateOr(Shadow, CombinedShadow, "_msprop");
5727 }
5728
5729 setShadow(&I, CombinedShadow);
5730
5731 setOriginForNaryOp(I);
5732 }
5733
5734 // Approximation only
5735 //
5736 // e.g., <16 x i8> @llvm.aarch64.neon.pmull64(i64, i64)
5737 void handleNEONVectorMultiplyIntrinsic(IntrinsicInst &I) {
5738 assert(I.arg_size() == 2);
5739
5740 handleShadowOr(I);
5741 }
5742
5743 // Handles:
5744 // <4 x half> @llvm.aarch64.neon.fp8.fdot2.lane
5745 // (<4 x half>, <8 x i8>, <16 x i8>, i32)
5746 // accumulator A B lane
5747 //
5748 // <8 x half> @llvm.aarch64.neon.fp8.fdot2.lane
5749 // (<8 x half>, <16 x i8>, <16 x i8>, i32)
5750 // <2 x float> @llvm.aarch64.neon.fp8.fdot4.lane
5751 // (<2 x float>, <8 x i8>, <16 x i8>, i32)
5752 // <4 x float> @llvm.aarch64.neon.fp8.fdot4.lane
5753 // (<4 x float>, <16 x i8>, <16 x i8>, i32)
5754 //
5755 // The lane specifies which pair (fdot2) or quad (fdot4) of numbers to
5756 // extract from B, which is then splatted before being used in the dot
5757 // products e.g., for
5758 // <4 x half> @llvm.aarch64.neon.fp8.fdot2.lane:
5759 // (<4 x half>, <8 x i8>, <16 x i8>, 1)
5760 //
5761 // acc[0] acc[1] acc[2] acc[3]
5762 // + + + + + + + +
5763 // A[0] A[1] A[2] A[3] A[4] A[5] A[6] A[7]
5764 // * * * * * * * *
5765 // B[2] B[3] B[2] B[3] B[2] B[3] B[2] B[3]
5766 //
5767 // Notice that if any bit of B[2] or B[3] is uninitialized, every accumulator
5768 // value will become tainted; we approximate this by marking the output as
5769 // fully uninitialized. This permits a 'Select' optimization.
5770 //
5771 // This function is separate from handleVectorDotProductIntrinsic(), because
5772 // the non-overlapping features (e.g., odd/even lanes vs. numbered lanes,
5773 // ZeroPurifies, EltSizeInBits) and optimizations make clean code reuse
5774 // difficult.
5775 void handleNEONDotProductLaneIntrinsic(IntrinsicInst &I,
5776 unsigned ReductionFactor) {
5777 IRBuilder<> IRB(&I);
5778 assert(I.arg_size() == 4);
5779
5780 [[maybe_unused]] Value *VAcc = I.getOperand(0);
5781 [[maybe_unused]] Value *Va = I.getOperand(1);
5782 [[maybe_unused]] Value *Vb = I.getOperand(2);
5783 Value *Lane = I.getOperand(3);
5784
5786 assert(VAcc->getType() == I.getType());
5787
5790 I.getType()->getPrimitiveSizeInBits());
5791
5792 assert(cast<FixedVectorType>(Va->getType())->getNumElements() ==
5793 cast<FixedVectorType>(I.getType())->getNumElements() *
5794 ReductionFactor);
5795
5797 // Deliberately not strict equality
5799 I.getType()->getPrimitiveSizeInBits());
5800
5801 assert(Lane->getType()->isIntegerTy());
5802
5803 // (<4 x 16>, <8 x i8>, <16 x i8>)
5804 // SAcc Sa Sb
5805 Value *SAcc = getShadow(&I, 0);
5806 Value *Sa = getShadow(&I, 1);
5807 Value *Sb = getShadow(&I, 2);
5808
5809 // Cast the shadows to:
5810 // (<4 x i16>, <4 x i16>, <8 x i16>)
5811 // SAcc Sa Sb
5812 Sa = IRB.CreateBitCast(Sa, SAcc->getType());
5813 Sb = IRB.CreateBitCast(
5814 Sb, FixedVectorType::getWithSizeAndScalar(
5816 cast<FixedVectorType>(SAcc->getType())->getElementType()));
5817
5818 // All-or-nothing shadows
5819 Sa =
5820 IRB.CreateSExt(IRB.CreateICmpNE(Sa, getCleanShadow(Sa)), Sa->getType());
5821
5822 // Extract the specific lane from Sb to get i16, then turn it into a single
5823 // bit representing if it is fully initialized.
5824 Sb = IRB.CreateExtractElement(Sb, Lane);
5825 Value *SbClean = IRB.CreateIsNull(Sb);
5826
5827 Value *SOutput = IRB.CreateOr(SAcc, Sa);
5828
5829 // Select is cheaper than broadcasting Sb into <4 x i16>.
5830 SOutput = IRB.CreateSelect(SbClean, SOutput, getPoisonedShadow(SOutput));
5831
5832 setShadow(&I, SOutput);
5833 setOriginForNaryOp(I);
5834 }
5835
5836 bool maybeHandleCrossPlatformIntrinsic(IntrinsicInst &I) {
5837 switch (I.getIntrinsicID()) {
5838 case Intrinsic::uadd_with_overflow:
5839 case Intrinsic::sadd_with_overflow:
5840 case Intrinsic::usub_with_overflow:
5841 case Intrinsic::ssub_with_overflow:
5842 case Intrinsic::umul_with_overflow:
5843 case Intrinsic::smul_with_overflow:
5844 handleArithmeticWithOverflow(I);
5845 break;
5846 case Intrinsic::modf:
5847 case Intrinsic::sincos:
5848 case Intrinsic::sincospi:
5849 handleModfOrSincos(I);
5850 break;
5851 case Intrinsic::abs:
5852 handleAbsIntrinsic(I);
5853 break;
5854 case Intrinsic::bitreverse:
5855 handleIntrinsicByApplyingToShadow(I, I.getIntrinsicID(),
5856 /*trailingVerbatimArgs=*/0,
5857 /*forceIntegerIntrinsic=*/false);
5858 break;
5859 case Intrinsic::is_fpclass:
5860 handleIsFpClass(I);
5861 break;
5862 case Intrinsic::lifetime_start:
5863 handleLifetimeStart(I);
5864 break;
5865 case Intrinsic::launder_invariant_group:
5866 handleInvariantGroup(I);
5867 break;
5868 case Intrinsic::bswap:
5869 handleBswap(I);
5870 break;
5871 case Intrinsic::ctlz:
5872 case Intrinsic::cttz:
5873 handleCountLeadingTrailingZeros(I);
5874 break;
5875 case Intrinsic::masked_compressstore:
5876 handleMaskedCompressStore(I);
5877 break;
5878 case Intrinsic::masked_expandload:
5879 handleMaskedExpandLoad(I);
5880 break;
5881 case Intrinsic::masked_gather:
5882 handleMaskedGather(I);
5883 break;
5884 case Intrinsic::masked_scatter:
5885 handleMaskedScatter(I);
5886 break;
5887 case Intrinsic::masked_store:
5888 handleMaskedStore(I);
5889 break;
5890 case Intrinsic::masked_load:
5891 handleMaskedLoad(I);
5892 break;
5893 case Intrinsic::masked_udiv:
5894 case Intrinsic::masked_sdiv:
5895 case Intrinsic::masked_urem:
5896 case Intrinsic::masked_srem:
5897 handleMaskedIntegerDivRem(I);
5898 break;
5899 case Intrinsic::vector_reduce_and:
5900 handleVectorReduceAndIntrinsic(I);
5901 break;
5902 case Intrinsic::vector_reduce_or:
5903 handleVectorReduceOrIntrinsic(I);
5904 break;
5905
5906 case Intrinsic::vector_reduce_add:
5907 case Intrinsic::vector_reduce_xor:
5908 case Intrinsic::vector_reduce_mul:
5909 // Signed/Unsigned Min/Max
5910 // TODO: handling similarly to AND/OR may be more precise.
5911 case Intrinsic::vector_reduce_smax:
5912 case Intrinsic::vector_reduce_smin:
5913 case Intrinsic::vector_reduce_umax:
5914 case Intrinsic::vector_reduce_umin:
5915 // TODO: this has no false positives, but arguably we should check that all
5916 // the bits are initialized.
5917 case Intrinsic::vector_reduce_fmax:
5918 case Intrinsic::vector_reduce_fmin:
5919 handleVectorReduceIntrinsic(I, /*AllowShadowCast=*/false);
5920 break;
5921
5922 case Intrinsic::vector_reduce_fadd:
5923 case Intrinsic::vector_reduce_fmul:
5924 handleVectorReduceWithStarterIntrinsic(I);
5925 break;
5926
5927 case Intrinsic::scmp:
5928 case Intrinsic::ucmp: {
5929 handleShadowOr(I);
5930 break;
5931 }
5932
5933 case Intrinsic::fshl:
5934 case Intrinsic::fshr:
5935 handleFunnelShift(I);
5936 break;
5937
5938 case Intrinsic::pdep:
5939 case Intrinsic::pext:
5940 handleGenericBitManipulation(I);
5941 break;
5942
5943 case Intrinsic::is_constant:
5944 // The result of llvm.is.constant() is always defined.
5945 setShadow(&I, getCleanShadow(&I));
5946 setOrigin(&I, getCleanOrigin());
5947 break;
5948
5949 // The non-saturating versions are handled by visitFPTo[US]IInst().
5950 //
5951 // N.B. some platform-specific intrinsics, such as AArch64 fcvtz[us], are
5952 // lowered to these cross-platform intrinsics.
5953 case Intrinsic::fptosi_sat:
5954 case Intrinsic::fptoui_sat:
5955 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
5956 break;
5957
5958 // e.g.,
5959 // notail call void (...) @llvm.fake.use(i64 %x)
5960 // notail call void (...) @llvm.fake.use(i32 %y)
5961 // notail call void (...) @llvm.fake.use(ptr %z)
5962 case Intrinsic::fake_use:
5963 assert(I.getType()->isVoidTy());
5964 // fake_uses aren't real, they can't hurt you. If the use isn't real, it
5965 // can't be a real use-of-uninitialized memory. Silently skip over
5966 // fake_use.
5967 return true;
5968
5969 default:
5970 return false;
5971 }
5972
5973 return true;
5974 }
5975
5976 bool maybeHandleX86SIMDIntrinsic(IntrinsicInst &I) {
5977 switch (I.getIntrinsicID()) {
5978 case Intrinsic::x86_sse_stmxcsr:
5979 handleStmxcsr(I);
5980 break;
5981 case Intrinsic::x86_sse_ldmxcsr:
5982 handleLdmxcsr(I);
5983 break;
5984
5985 // Convert Scalar Double Precision Floating-Point Value
5986 // to Unsigned Doubleword Integer
5987 // etc.
5988 case Intrinsic::x86_avx512_vcvtsd2usi64:
5989 case Intrinsic::x86_avx512_vcvtsd2usi32:
5990 case Intrinsic::x86_avx512_vcvtss2usi64:
5991 case Intrinsic::x86_avx512_vcvtss2usi32:
5992 case Intrinsic::x86_avx512_cvttss2usi64:
5993 case Intrinsic::x86_avx512_cvttss2usi:
5994 case Intrinsic::x86_avx512_cvttsd2usi64:
5995 case Intrinsic::x86_avx512_cvttsd2usi:
5996 case Intrinsic::x86_avx512_cvtusi2ss:
5997 case Intrinsic::x86_avx512_cvtusi642sd:
5998 case Intrinsic::x86_avx512_cvtusi642ss:
5999 handleSSEVectorConvertIntrinsic(I, 1, true);
6000 break;
6001 case Intrinsic::x86_sse2_cvtsd2si64:
6002 case Intrinsic::x86_sse2_cvtsd2si:
6003 case Intrinsic::x86_sse2_cvtsd2ss:
6004 case Intrinsic::x86_sse2_cvttsd2si64:
6005 case Intrinsic::x86_sse2_cvttsd2si:
6006 case Intrinsic::x86_sse_cvtss2si64:
6007 case Intrinsic::x86_sse_cvtss2si:
6008 case Intrinsic::x86_sse_cvttss2si64:
6009 case Intrinsic::x86_sse_cvttss2si:
6010 handleSSEVectorConvertIntrinsic(I, 1);
6011 break;
6012 case Intrinsic::x86_sse_cvtps2pi:
6013 case Intrinsic::x86_sse_cvttps2pi:
6014 handleSSEVectorConvertIntrinsic(I, 2);
6015 break;
6016
6017 // TODO:
6018 // <1 x i64> @llvm.x86.sse.cvtpd2pi(<2 x double>)
6019 // <2 x double> @llvm.x86.sse.cvtpi2pd(<1 x i64>)
6020 // <4 x float> @llvm.x86.sse.cvtpi2ps(<4 x float>, <1 x i64>)
6021
6022 case Intrinsic::x86_vcvtps2ph_128:
6023 case Intrinsic::x86_vcvtps2ph_256: {
6024 handleSSEVectorConvertIntrinsicByProp(I, /*HasRoundingMode=*/true);
6025 break;
6026 }
6027
6028 // Convert Packed Single Precision Floating-Point Values
6029 // to Packed Signed Doubleword Integer Values
6030 //
6031 // <16 x i32> @llvm.x86.avx512.mask.cvtps2dq.512
6032 // (<16 x float>, <16 x i32>, i16, i32)
6033 case Intrinsic::x86_avx512_mask_cvtps2dq_512:
6034 handleAVX512VectorConvertFPToInt(I, /*LastMask=*/false);
6035 break;
6036
6037 // Convert Packed Double Precision Floating-Point Values
6038 // to Packed Single Precision Floating-Point Values
6039 case Intrinsic::x86_sse2_cvtpd2ps:
6040 case Intrinsic::x86_sse2_cvtps2dq:
6041 case Intrinsic::x86_sse2_cvtpd2dq:
6042 case Intrinsic::x86_sse2_cvttps2dq:
6043 case Intrinsic::x86_sse2_cvttpd2dq:
6044 case Intrinsic::x86_avx_cvt_pd2_ps_256:
6045 case Intrinsic::x86_avx_cvt_ps2dq_256:
6046 case Intrinsic::x86_avx_cvt_pd2dq_256:
6047 case Intrinsic::x86_avx_cvtt_ps2dq_256:
6048 case Intrinsic::x86_avx_cvtt_pd2dq_256: {
6049 handleSSEVectorConvertIntrinsicByProp(I, /*HasRoundingMode=*/false);
6050 break;
6051 }
6052
6053 // Convert Single-Precision FP Value to 16-bit FP Value
6054 // <16 x i16> @llvm.x86.avx512.mask.vcvtps2ph.512
6055 // (<16 x float>, i32, <16 x i16>, i16)
6056 // <8 x i16> @llvm.x86.avx512.mask.vcvtps2ph.128
6057 // (<4 x float>, i32, <8 x i16>, i8)
6058 // <8 x i16> @llvm.x86.avx512.mask.vcvtps2ph.256
6059 // (<8 x float>, i32, <8 x i16>, i8)
6060 case Intrinsic::x86_avx512_mask_vcvtps2ph_512:
6061 case Intrinsic::x86_avx512_mask_vcvtps2ph_256:
6062 case Intrinsic::x86_avx512_mask_vcvtps2ph_128:
6063 handleAVX512VectorConvertFPToInt(I, /*LastMask=*/true);
6064 break;
6065
6066 // Shift Packed Data (Left Logical, Right Arithmetic, Right Logical)
6067 case Intrinsic::x86_avx512_psll_w_512:
6068 case Intrinsic::x86_avx512_psll_d_512:
6069 case Intrinsic::x86_avx512_psll_q_512:
6070 case Intrinsic::x86_avx512_pslli_w_512:
6071 case Intrinsic::x86_avx512_pslli_d_512:
6072 case Intrinsic::x86_avx512_pslli_q_512:
6073 case Intrinsic::x86_avx512_psrl_w_512:
6074 case Intrinsic::x86_avx512_psrl_d_512:
6075 case Intrinsic::x86_avx512_psrl_q_512:
6076 case Intrinsic::x86_avx512_psra_w_512:
6077 case Intrinsic::x86_avx512_psra_d_512:
6078 case Intrinsic::x86_avx512_psra_q_512:
6079 case Intrinsic::x86_avx512_psrli_w_512:
6080 case Intrinsic::x86_avx512_psrli_d_512:
6081 case Intrinsic::x86_avx512_psrli_q_512:
6082 case Intrinsic::x86_avx512_psrai_w_512:
6083 case Intrinsic::x86_avx512_psrai_d_512:
6084 case Intrinsic::x86_avx512_psrai_q_512:
6085 case Intrinsic::x86_avx512_psra_q_256:
6086 case Intrinsic::x86_avx512_psra_q_128:
6087 case Intrinsic::x86_avx512_psrai_q_256:
6088 case Intrinsic::x86_avx512_psrai_q_128:
6089 case Intrinsic::x86_avx2_psll_w:
6090 case Intrinsic::x86_avx2_psll_d:
6091 case Intrinsic::x86_avx2_psll_q:
6092 case Intrinsic::x86_avx2_pslli_w:
6093 case Intrinsic::x86_avx2_pslli_d:
6094 case Intrinsic::x86_avx2_pslli_q:
6095 case Intrinsic::x86_avx2_psrl_w:
6096 case Intrinsic::x86_avx2_psrl_d:
6097 case Intrinsic::x86_avx2_psrl_q:
6098 case Intrinsic::x86_avx2_psra_w:
6099 case Intrinsic::x86_avx2_psra_d:
6100 case Intrinsic::x86_avx2_psrli_w:
6101 case Intrinsic::x86_avx2_psrli_d:
6102 case Intrinsic::x86_avx2_psrli_q:
6103 case Intrinsic::x86_avx2_psrai_w:
6104 case Intrinsic::x86_avx2_psrai_d:
6105 case Intrinsic::x86_sse2_psll_w:
6106 case Intrinsic::x86_sse2_psll_d:
6107 case Intrinsic::x86_sse2_psll_q:
6108 case Intrinsic::x86_sse2_pslli_w:
6109 case Intrinsic::x86_sse2_pslli_d:
6110 case Intrinsic::x86_sse2_pslli_q:
6111 case Intrinsic::x86_sse2_psrl_w:
6112 case Intrinsic::x86_sse2_psrl_d:
6113 case Intrinsic::x86_sse2_psrl_q:
6114 case Intrinsic::x86_sse2_psra_w:
6115 case Intrinsic::x86_sse2_psra_d:
6116 case Intrinsic::x86_sse2_psrli_w:
6117 case Intrinsic::x86_sse2_psrli_d:
6118 case Intrinsic::x86_sse2_psrli_q:
6119 case Intrinsic::x86_sse2_psrai_w:
6120 case Intrinsic::x86_sse2_psrai_d:
6121 case Intrinsic::x86_mmx_psll_w:
6122 case Intrinsic::x86_mmx_psll_d:
6123 case Intrinsic::x86_mmx_psll_q:
6124 case Intrinsic::x86_mmx_pslli_w:
6125 case Intrinsic::x86_mmx_pslli_d:
6126 case Intrinsic::x86_mmx_pslli_q:
6127 case Intrinsic::x86_mmx_psrl_w:
6128 case Intrinsic::x86_mmx_psrl_d:
6129 case Intrinsic::x86_mmx_psrl_q:
6130 case Intrinsic::x86_mmx_psra_w:
6131 case Intrinsic::x86_mmx_psra_d:
6132 case Intrinsic::x86_mmx_psrli_w:
6133 case Intrinsic::x86_mmx_psrli_d:
6134 case Intrinsic::x86_mmx_psrli_q:
6135 case Intrinsic::x86_mmx_psrai_w:
6136 case Intrinsic::x86_mmx_psrai_d:
6137 handleVectorShiftIntrinsic(I, /* Variable */ false);
6138 break;
6139 case Intrinsic::x86_avx2_psllv_d:
6140 case Intrinsic::x86_avx2_psllv_d_256:
6141 case Intrinsic::x86_avx512_psllv_d_512:
6142 case Intrinsic::x86_avx2_psllv_q:
6143 case Intrinsic::x86_avx2_psllv_q_256:
6144 case Intrinsic::x86_avx512_psllv_q_512:
6145 case Intrinsic::x86_avx2_psrlv_d:
6146 case Intrinsic::x86_avx2_psrlv_d_256:
6147 case Intrinsic::x86_avx512_psrlv_d_512:
6148 case Intrinsic::x86_avx2_psrlv_q:
6149 case Intrinsic::x86_avx2_psrlv_q_256:
6150 case Intrinsic::x86_avx512_psrlv_q_512:
6151 case Intrinsic::x86_avx2_psrav_d:
6152 case Intrinsic::x86_avx2_psrav_d_256:
6153 case Intrinsic::x86_avx512_psrav_d_512:
6154 case Intrinsic::x86_avx512_psrav_q_128:
6155 case Intrinsic::x86_avx512_psrav_q_256:
6156 case Intrinsic::x86_avx512_psrav_q_512:
6157 handleVectorShiftIntrinsic(I, /* Variable */ true);
6158 break;
6159
6160 // Pack with Signed/Unsigned Saturation
6161 case Intrinsic::x86_sse2_packsswb_128:
6162 case Intrinsic::x86_sse2_packssdw_128:
6163 case Intrinsic::x86_sse2_packuswb_128:
6164 case Intrinsic::x86_sse41_packusdw:
6165 case Intrinsic::x86_avx2_packsswb:
6166 case Intrinsic::x86_avx2_packssdw:
6167 case Intrinsic::x86_avx2_packuswb:
6168 case Intrinsic::x86_avx2_packusdw:
6169 // e.g., <64 x i8> @llvm.x86.avx512.packsswb.512
6170 // (<32 x i16> %a, <32 x i16> %b)
6171 // <32 x i16> @llvm.x86.avx512.packssdw.512
6172 // (<16 x i32> %a, <16 x i32> %b)
6173 // Note: AVX512 masked variants are auto-upgraded by LLVM.
6174 case Intrinsic::x86_avx512_packsswb_512:
6175 case Intrinsic::x86_avx512_packssdw_512:
6176 case Intrinsic::x86_avx512_packuswb_512:
6177 case Intrinsic::x86_avx512_packusdw_512:
6178 handleVectorPackIntrinsic(I);
6179 break;
6180
6181 case Intrinsic::x86_sse41_pblendvb:
6182 case Intrinsic::x86_sse41_blendvpd:
6183 case Intrinsic::x86_sse41_blendvps:
6184 case Intrinsic::x86_avx_blendv_pd_256:
6185 case Intrinsic::x86_avx_blendv_ps_256:
6186 case Intrinsic::x86_avx2_pblendvb:
6187 handleBlendvIntrinsic(I);
6188 break;
6189
6190 case Intrinsic::x86_avx_dp_ps_256:
6191 case Intrinsic::x86_sse41_dppd:
6192 case Intrinsic::x86_sse41_dpps:
6193 handleDppIntrinsic(I);
6194 break;
6195
6196 case Intrinsic::x86_mmx_packsswb:
6197 case Intrinsic::x86_mmx_packuswb:
6198 handleVectorPackIntrinsic(I, 16);
6199 break;
6200
6201 case Intrinsic::x86_mmx_packssdw:
6202 handleVectorPackIntrinsic(I, 32);
6203 break;
6204
6205 case Intrinsic::x86_mmx_psad_bw:
6206 handleVectorSadIntrinsic(I, true);
6207 break;
6208 case Intrinsic::x86_sse2_psad_bw:
6209 case Intrinsic::x86_avx2_psad_bw:
6210 handleVectorSadIntrinsic(I);
6211 break;
6212
6213 // Multiply and Add Packed Words
6214 // < 4 x i32> @llvm.x86.sse2.pmadd.wd(<8 x i16>, <8 x i16>)
6215 // < 8 x i32> @llvm.x86.avx2.pmadd.wd(<16 x i16>, <16 x i16>)
6216 // <16 x i32> @llvm.x86.avx512.pmaddw.d.512(<32 x i16>, <32 x i16>)
6217 //
6218 // Multiply and Add Packed Signed and Unsigned Bytes
6219 // < 8 x i16> @llvm.x86.ssse3.pmadd.ub.sw.128(<16 x i8>, <16 x i8>)
6220 // <16 x i16> @llvm.x86.avx2.pmadd.ub.sw(<32 x i8>, <32 x i8>)
6221 // <32 x i16> @llvm.x86.avx512.pmaddubs.w.512(<64 x i8>, <64 x i8>)
6222 //
6223 // These intrinsics are auto-upgraded into non-masked forms:
6224 // < 4 x i32> @llvm.x86.avx512.mask.pmaddw.d.128
6225 // (<8 x i16>, <8 x i16>, <4 x i32>, i8)
6226 // < 8 x i32> @llvm.x86.avx512.mask.pmaddw.d.256
6227 // (<16 x i16>, <16 x i16>, <8 x i32>, i8)
6228 // <16 x i32> @llvm.x86.avx512.mask.pmaddw.d.512
6229 // (<32 x i16>, <32 x i16>, <16 x i32>, i16)
6230 // < 8 x i16> @llvm.x86.avx512.mask.pmaddubs.w.128
6231 // (<16 x i8>, <16 x i8>, <8 x i16>, i8)
6232 // <16 x i16> @llvm.x86.avx512.mask.pmaddubs.w.256
6233 // (<32 x i8>, <32 x i8>, <16 x i16>, i16)
6234 // <32 x i16> @llvm.x86.avx512.mask.pmaddubs.w.512
6235 // (<64 x i8>, <64 x i8>, <32 x i16>, i32)
6236 case Intrinsic::x86_sse2_pmadd_wd:
6237 case Intrinsic::x86_avx2_pmadd_wd:
6238 case Intrinsic::x86_avx512_pmaddw_d_512:
6239 case Intrinsic::x86_ssse3_pmadd_ub_sw_128:
6240 case Intrinsic::x86_avx2_pmadd_ub_sw:
6241 case Intrinsic::x86_avx512_pmaddubs_w_512:
6242 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6243 /*ZeroPurifies=*/true,
6244 /*EltSizeInBits=*/0,
6245 /*Lanes=*/kBothLanes);
6246 break;
6247
6248 // <1 x i64> @llvm.x86.ssse3.pmadd.ub.sw(<1 x i64>, <1 x i64>)
6249 case Intrinsic::x86_ssse3_pmadd_ub_sw:
6250 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6251 /*ZeroPurifies=*/true,
6252 /*EltSizeInBits=*/8,
6253 /*Lanes=*/kBothLanes);
6254 break;
6255
6256 // <1 x i64> @llvm.x86.mmx.pmadd.wd(<1 x i64>, <1 x i64>)
6257 case Intrinsic::x86_mmx_pmadd_wd:
6258 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6259 /*ZeroPurifies=*/true,
6260 /*EltSizeInBits=*/16,
6261 /*Lanes=*/kBothLanes);
6262 break;
6263
6264 // BFloat16 multiply-add to single-precision
6265 // <4 x float> llvm.aarch64.neon.bfmlalt
6266 // (<4 x float>, <8 x bfloat>, <8 x bfloat>)
6267 case Intrinsic::aarch64_neon_bfmlalt:
6268 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6269 /*ZeroPurifies=*/false,
6270 /*EltSizeInBits=*/0,
6271 /*Lanes=*/kOddLanes);
6272 break;
6273
6274 // <4 x float> llvm.aarch64.neon.bfmlalb
6275 // (<4 x float>, <8 x bfloat>, <8 x bfloat>)
6276 case Intrinsic::aarch64_neon_bfmlalb:
6277 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6278 /*ZeroPurifies=*/false,
6279 /*EltSizeInBits=*/0,
6280 /*Lanes=*/kEvenLanes);
6281 break;
6282
6283 // AVX Vector Neural Network Instructions: bytes
6284 //
6285 // Multiply and Add Signed Bytes
6286 // < 4 x i32> @llvm.x86.avx2.vpdpbssd.128
6287 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6288 // < 8 x i32> @llvm.x86.avx2.vpdpbssd.256
6289 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6290 // <16 x i32> @llvm.x86.avx10.vpdpbssd.512
6291 // (<16 x i32>, <64 x i8>, <64 x i8>)
6292 //
6293 // Multiply and Add Signed Bytes With Saturation
6294 // < 4 x i32> @llvm.x86.avx2.vpdpbssds.128
6295 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6296 // < 8 x i32> @llvm.x86.avx2.vpdpbssds.256
6297 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6298 // <16 x i32> @llvm.x86.avx10.vpdpbssds.512
6299 // (<16 x i32>, <64 x i8>, <64 x i8>)
6300 //
6301 // Multiply and Add Signed and Unsigned Bytes
6302 // < 4 x i32> @llvm.x86.avx2.vpdpbsud.128
6303 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6304 // < 8 x i32> @llvm.x86.avx2.vpdpbsud.256
6305 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6306 // <16 x i32> @llvm.x86.avx10.vpdpbsud.512
6307 // (<16 x i32>, <64 x i8>, <64 x i8>)
6308 //
6309 // Multiply and Add Signed and Unsigned Bytes With Saturation
6310 // < 4 x i32> @llvm.x86.avx2.vpdpbsuds.128
6311 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6312 // < 8 x i32> @llvm.x86.avx2.vpdpbsuds.256
6313 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6314 // <16 x i32> @llvm.x86.avx512.vpdpbusds.512
6315 // (<16 x i32>, <64 x i8>, <64 x i8>)
6316 //
6317 // Multiply and Add Unsigned and Signed Bytes
6318 // < 4 x i32> @llvm.x86.avx512.vpdpbusd.128
6319 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6320 // < 8 x i32> @llvm.x86.avx512.vpdpbusd.256
6321 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6322 // <16 x i32> @llvm.x86.avx512.vpdpbusd.512
6323 // (<16 x i32>, <64 x i8>, <64 x i8>)
6324 //
6325 // Multiply and Add Unsigned and Signed Bytes With Saturation
6326 // < 4 x i32> @llvm.x86.avx512.vpdpbusds.128
6327 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6328 // < 8 x i32> @llvm.x86.avx512.vpdpbusds.256
6329 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6330 // <16 x i32> @llvm.x86.avx10.vpdpbsuds.512
6331 // (<16 x i32>, <64 x i8>, <64 x i8>)
6332 //
6333 // Multiply and Add Unsigned Bytes
6334 // < 4 x i32> @llvm.x86.avx2.vpdpbuud.128
6335 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6336 // < 8 x i32> @llvm.x86.avx2.vpdpbuud.256
6337 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6338 // <16 x i32> @llvm.x86.avx10.vpdpbuud.512
6339 // (<16 x i32>, <64 x i8>, <64 x i8>)
6340 //
6341 // Multiply and Add Unsigned Bytes With Saturation
6342 // < 4 x i32> @llvm.x86.avx2.vpdpbuuds.128
6343 // (< 4 x i32>, <16 x i8>, <16 x i8>)
6344 // < 8 x i32> @llvm.x86.avx2.vpdpbuuds.256
6345 // (< 8 x i32>, <32 x i8>, <32 x i8>)
6346 // <16 x i32> @llvm.x86.avx10.vpdpbuuds.512
6347 // (<16 x i32>, <64 x i8>, <64 x i8>)
6348 //
6349 // These intrinsics are auto-upgraded into non-masked forms:
6350 // <4 x i32> @llvm.x86.avx512.mask.vpdpbusd.128
6351 // (<4 x i32>, <16 x i8>, <16 x i8>, i8)
6352 // <4 x i32> @llvm.x86.avx512.maskz.vpdpbusd.128
6353 // (<4 x i32>, <16 x i8>, <16 x i8>, i8)
6354 // <8 x i32> @llvm.x86.avx512.mask.vpdpbusd.256
6355 // (<8 x i32>, <32 x i8>, <32 x i8>, i8)
6356 // <8 x i32> @llvm.x86.avx512.maskz.vpdpbusd.256
6357 // (<8 x i32>, <32 x i8>, <32 x i8>, i8)
6358 // <16 x i32> @llvm.x86.avx512.mask.vpdpbusd.512
6359 // (<16 x i32>, <64 x i8>, <64 x i8>, i16)
6360 // <16 x i32> @llvm.x86.avx512.maskz.vpdpbusd.512
6361 // (<16 x i32>, <64 x i8>, <64 x i8>, i16)
6362 //
6363 // <4 x i32> @llvm.x86.avx512.mask.vpdpbusds.128
6364 // (<4 x i32>, <16 x i8>, <16 x i8>, i8)
6365 // <4 x i32> @llvm.x86.avx512.maskz.vpdpbusds.128
6366 // (<4 x i32>, <16 x i8>, <16 x i8>, i8)
6367 // <8 x i32> @llvm.x86.avx512.mask.vpdpbusds.256
6368 // (<8 x i32>, <32 x i8>, <32 x i8>, i8)
6369 // <8 x i32> @llvm.x86.avx512.maskz.vpdpbusds.256
6370 // (<8 x i32>, <32 x i8>, <32 x i8>, i8)
6371 // <16 x i32> @llvm.x86.avx512.mask.vpdpbusds.512
6372 // (<16 x i32>, <64 x i8>, <64 x i8>, i16)
6373 // <16 x i32> @llvm.x86.avx512.maskz.vpdpbusds.512
6374 // (<16 x i32>, <64 x i8>, <64 x i8>, i16)
6375 case Intrinsic::x86_avx512_vpdpbusd_128:
6376 case Intrinsic::x86_avx512_vpdpbusd_256:
6377 case Intrinsic::x86_avx512_vpdpbusd_512:
6378 case Intrinsic::x86_avx512_vpdpbusds_128:
6379 case Intrinsic::x86_avx512_vpdpbusds_256:
6380 case Intrinsic::x86_avx512_vpdpbusds_512:
6381 case Intrinsic::x86_avx2_vpdpbssd_128:
6382 case Intrinsic::x86_avx2_vpdpbssd_256:
6383 case Intrinsic::x86_avx10_vpdpbssd_512:
6384 case Intrinsic::x86_avx2_vpdpbssds_128:
6385 case Intrinsic::x86_avx2_vpdpbssds_256:
6386 case Intrinsic::x86_avx10_vpdpbssds_512:
6387 case Intrinsic::x86_avx2_vpdpbsud_128:
6388 case Intrinsic::x86_avx2_vpdpbsud_256:
6389 case Intrinsic::x86_avx10_vpdpbsud_512:
6390 case Intrinsic::x86_avx2_vpdpbsuds_128:
6391 case Intrinsic::x86_avx2_vpdpbsuds_256:
6392 case Intrinsic::x86_avx10_vpdpbsuds_512:
6393 case Intrinsic::x86_avx2_vpdpbuud_128:
6394 case Intrinsic::x86_avx2_vpdpbuud_256:
6395 case Intrinsic::x86_avx10_vpdpbuud_512:
6396 case Intrinsic::x86_avx2_vpdpbuuds_128:
6397 case Intrinsic::x86_avx2_vpdpbuuds_256:
6398 case Intrinsic::x86_avx10_vpdpbuuds_512:
6399 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/4,
6400 /*ZeroPurifies=*/true,
6401 /*EltSizeInBits=*/0,
6402 /*Lanes=*/kBothLanes);
6403 break;
6404
6405 // AVX Vector Neural Network Instructions: words
6406 //
6407 // Multiply and Add Signed Word Integers
6408 // < 4 x i32> @llvm.x86.avx512.vpdpwssd.128
6409 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6410 // < 8 x i32> @llvm.x86.avx512.vpdpwssd.256
6411 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6412 // <16 x i32> @llvm.x86.avx512.vpdpwssd.512
6413 // (<16 x i32>, <32 x i16>, <32 x i16>)
6414 //
6415 // Multiply and Add Signed Word Integers With Saturation
6416 // < 4 x i32> @llvm.x86.avx512.vpdpwssds.128
6417 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6418 // < 8 x i32> @llvm.x86.avx512.vpdpwssds.256
6419 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6420 // <16 x i32> @llvm.x86.avx512.vpdpwssds.512
6421 // (<16 x i32>, <32 x i16>, <32 x i16>)
6422 //
6423 // Multiply and Add Signed and Unsigned Word Integers
6424 // < 4 x i32> @llvm.x86.avx2.vpdpwsud.128
6425 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6426 // < 8 x i32> @llvm.x86.avx2.vpdpwsud.256
6427 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6428 // <16 x i32> @llvm.x86.avx10.vpdpwsud.512
6429 // (<16 x i32>, <32 x i16>, <32 x i16>)
6430 //
6431 // Multiply and Add Signed and Unsigned Word Integers With Saturation
6432 // < 4 x i32> @llvm.x86.avx2.vpdpwsuds.128
6433 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6434 // < 8 x i32> @llvm.x86.avx2.vpdpwsuds.256
6435 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6436 // <16 x i32> @llvm.x86.avx10.vpdpwsuds.512
6437 // (<16 x i32>, <32 x i16>, <32 x i16>)
6438 //
6439 // Multiply and Add Unsigned and Signed Word Integers
6440 // < 4 x i32> @llvm.x86.avx2.vpdpwusd.128
6441 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6442 // < 8 x i32> @llvm.x86.avx2.vpdpwusd.256
6443 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6444 // <16 x i32> @llvm.x86.avx10.vpdpwusd.512
6445 // (<16 x i32>, <32 x i16>, <32 x i16>)
6446 //
6447 // Multiply and Add Unsigned and Signed Word Integers With Saturation
6448 // < 4 x i32> @llvm.x86.avx2.vpdpwusds.128
6449 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6450 // < 8 x i32> @llvm.x86.avx2.vpdpwusds.256
6451 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6452 // <16 x i32> @llvm.x86.avx10.vpdpwusds.512
6453 // (<16 x i32>, <32 x i16>, <32 x i16>)
6454 //
6455 // Multiply and Add Unsigned and Unsigned Word Integers
6456 // < 4 x i32> @llvm.x86.avx2.vpdpwuud.128
6457 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6458 // < 8 x i32> @llvm.x86.avx2.vpdpwuud.256
6459 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6460 // <16 x i32> @llvm.x86.avx10.vpdpwuud.512
6461 // (<16 x i32>, <32 x i16>, <32 x i16>)
6462 //
6463 // Multiply and Add Unsigned and Unsigned Word Integers With Saturation
6464 // < 4 x i32> @llvm.x86.avx2.vpdpwuuds.128
6465 // (< 4 x i32>, < 8 x i16>, < 8 x i16>)
6466 // < 8 x i32> @llvm.x86.avx2.vpdpwuuds.256
6467 // (< 8 x i32>, <16 x i16>, <16 x i16>)
6468 // <16 x i32> @llvm.x86.avx10.vpdpwuuds.512
6469 // (<16 x i32>, <32 x i16>, <32 x i16>)
6470 //
6471 // These intrinsics are auto-upgraded into non-masked forms:
6472 // <4 x i32> @llvm.x86.avx512.mask.vpdpwssd.128
6473 // (<4 x i32>, <8 x i16>, <8 x i16>, i8)
6474 // <4 x i32> @llvm.x86.avx512.maskz.vpdpwssd.128
6475 // (<4 x i32>, <8 x i16>, <8 x i16>, i8)
6476 // <8 x i32> @llvm.x86.avx512.mask.vpdpwssd.256
6477 // (<8 x i32>, <16 x i16>, <16 x i16>, i8)
6478 // <8 x i32> @llvm.x86.avx512.maskz.vpdpwssd.256
6479 // (<8 x i32>, <16 x i16>, <16 x i16>, i8)
6480 // <16 x i32> @llvm.x86.avx512.mask.vpdpwssd.512
6481 // (<16 x i32>, <32 x i16>, <32 x i16>, i16)
6482 // <16 x i32> @llvm.x86.avx512.maskz.vpdpwssd.512
6483 // (<16 x i32>, <32 x i16>, <32 x i16>, i16)
6484 //
6485 // <4 x i32> @llvm.x86.avx512.mask.vpdpwssds.128
6486 // (<4 x i32>, <8 x i16>, <8 x i16>, i8)
6487 // <4 x i32> @llvm.x86.avx512.maskz.vpdpwssds.128
6488 // (<4 x i32>, <8 x i16>, <8 x i16>, i8)
6489 // <8 x i32> @llvm.x86.avx512.mask.vpdpwssds.256
6490 // (<8 x i32>, <16 x i16>, <16 x i16>, i8)
6491 // <8 x i32> @llvm.x86.avx512.maskz.vpdpwssds.256
6492 // (<8 x i32>, <16 x i16>, <16 x i16>, i8)
6493 // <16 x i32> @llvm.x86.avx512.mask.vpdpwssds.512
6494 // (<16 x i32>, <32 x i16>, <32 x i16>, i16)
6495 // <16 x i32> @llvm.x86.avx512.maskz.vpdpwssds.512
6496 // (<16 x i32>, <32 x i16>, <32 x i16>, i16)
6497 case Intrinsic::x86_avx512_vpdpwssd_128:
6498 case Intrinsic::x86_avx512_vpdpwssd_256:
6499 case Intrinsic::x86_avx512_vpdpwssd_512:
6500 case Intrinsic::x86_avx512_vpdpwssds_128:
6501 case Intrinsic::x86_avx512_vpdpwssds_256:
6502 case Intrinsic::x86_avx512_vpdpwssds_512:
6503 case Intrinsic::x86_avx2_vpdpwsud_128:
6504 case Intrinsic::x86_avx2_vpdpwsud_256:
6505 case Intrinsic::x86_avx10_vpdpwsud_512:
6506 case Intrinsic::x86_avx2_vpdpwsuds_128:
6507 case Intrinsic::x86_avx2_vpdpwsuds_256:
6508 case Intrinsic::x86_avx10_vpdpwsuds_512:
6509 case Intrinsic::x86_avx2_vpdpwusd_128:
6510 case Intrinsic::x86_avx2_vpdpwusd_256:
6511 case Intrinsic::x86_avx10_vpdpwusd_512:
6512 case Intrinsic::x86_avx2_vpdpwusds_128:
6513 case Intrinsic::x86_avx2_vpdpwusds_256:
6514 case Intrinsic::x86_avx10_vpdpwusds_512:
6515 case Intrinsic::x86_avx2_vpdpwuud_128:
6516 case Intrinsic::x86_avx2_vpdpwuud_256:
6517 case Intrinsic::x86_avx10_vpdpwuud_512:
6518 case Intrinsic::x86_avx2_vpdpwuuds_128:
6519 case Intrinsic::x86_avx2_vpdpwuuds_256:
6520 case Intrinsic::x86_avx10_vpdpwuuds_512:
6521 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6522 /*ZeroPurifies=*/true,
6523 /*EltSizeInBits=*/0,
6524 /*Lanes=*/kBothLanes);
6525 break;
6526
6527 // Dot Product of BF16 Pairs Accumulated Into Packed Single
6528 // Precision
6529 // <4 x float> @llvm.x86.avx512bf16.dpbf16ps.128
6530 // (<4 x float>, <8 x bfloat>, <8 x bfloat>)
6531 // <8 x float> @llvm.x86.avx512bf16.dpbf16ps.256
6532 // (<8 x float>, <16 x bfloat>, <16 x bfloat>)
6533 // <16 x float> @llvm.x86.avx512bf16.dpbf16ps.512
6534 // (<16 x float>, <32 x bfloat>, <32 x bfloat>)
6535 case Intrinsic::x86_avx512bf16_dpbf16ps_128:
6536 case Intrinsic::x86_avx512bf16_dpbf16ps_256:
6537 case Intrinsic::x86_avx512bf16_dpbf16ps_512:
6538 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
6539 /*ZeroPurifies=*/false,
6540 /*EltSizeInBits=*/0,
6541 /*Lanes=*/kBothLanes);
6542 break;
6543
6544 case Intrinsic::x86_sse_cmp_ss:
6545 case Intrinsic::x86_sse2_cmp_sd:
6546 case Intrinsic::x86_sse_comieq_ss:
6547 case Intrinsic::x86_sse_comilt_ss:
6548 case Intrinsic::x86_sse_comile_ss:
6549 case Intrinsic::x86_sse_comigt_ss:
6550 case Intrinsic::x86_sse_comige_ss:
6551 case Intrinsic::x86_sse_comineq_ss:
6552 case Intrinsic::x86_sse_ucomieq_ss:
6553 case Intrinsic::x86_sse_ucomilt_ss:
6554 case Intrinsic::x86_sse_ucomile_ss:
6555 case Intrinsic::x86_sse_ucomigt_ss:
6556 case Intrinsic::x86_sse_ucomige_ss:
6557 case Intrinsic::x86_sse_ucomineq_ss:
6558 case Intrinsic::x86_sse2_comieq_sd:
6559 case Intrinsic::x86_sse2_comilt_sd:
6560 case Intrinsic::x86_sse2_comile_sd:
6561 case Intrinsic::x86_sse2_comigt_sd:
6562 case Intrinsic::x86_sse2_comige_sd:
6563 case Intrinsic::x86_sse2_comineq_sd:
6564 case Intrinsic::x86_sse2_ucomieq_sd:
6565 case Intrinsic::x86_sse2_ucomilt_sd:
6566 case Intrinsic::x86_sse2_ucomile_sd:
6567 case Intrinsic::x86_sse2_ucomigt_sd:
6568 case Intrinsic::x86_sse2_ucomige_sd:
6569 case Intrinsic::x86_sse2_ucomineq_sd:
6570 handleVectorCompareScalarIntrinsic(I);
6571 break;
6572
6573 case Intrinsic::x86_avx_cmp_pd_256:
6574 case Intrinsic::x86_avx_cmp_ps_256:
6575 case Intrinsic::x86_sse2_cmp_pd:
6576 case Intrinsic::x86_sse_cmp_ps:
6577 handleVectorComparePackedIntrinsic(I, /*PredicateAsOperand=*/true);
6578 break;
6579
6580 case Intrinsic::x86_bmi_bextr_32:
6581 case Intrinsic::x86_bmi_bextr_64:
6582 case Intrinsic::x86_bmi_bzhi_32:
6583 case Intrinsic::x86_bmi_bzhi_64:
6584 handleGenericBitManipulation(I);
6585 break;
6586
6587 case Intrinsic::x86_pclmulqdq:
6588 case Intrinsic::x86_pclmulqdq_256:
6589 case Intrinsic::x86_pclmulqdq_512:
6590 handlePclmulIntrinsic(I);
6591 break;
6592
6593 case Intrinsic::x86_avx_round_pd_256:
6594 case Intrinsic::x86_avx_round_ps_256:
6595 case Intrinsic::x86_sse41_round_pd:
6596 case Intrinsic::x86_sse41_round_ps:
6597 handleRoundPdPsIntrinsic(I);
6598 break;
6599
6600 case Intrinsic::x86_sse41_round_sd:
6601 case Intrinsic::x86_sse41_round_ss:
6602 handleUnarySdSsIntrinsic(I);
6603 break;
6604
6605 case Intrinsic::x86_sse2_max_sd:
6606 case Intrinsic::x86_sse_max_ss:
6607 case Intrinsic::x86_sse2_min_sd:
6608 case Intrinsic::x86_sse_min_ss:
6609 handleBinarySdSsIntrinsic(I);
6610 break;
6611
6612 case Intrinsic::x86_avx_vtestc_pd:
6613 case Intrinsic::x86_avx_vtestc_pd_256:
6614 case Intrinsic::x86_avx_vtestc_ps:
6615 case Intrinsic::x86_avx_vtestc_ps_256:
6616 case Intrinsic::x86_avx_vtestnzc_pd:
6617 case Intrinsic::x86_avx_vtestnzc_pd_256:
6618 case Intrinsic::x86_avx_vtestnzc_ps:
6619 case Intrinsic::x86_avx_vtestnzc_ps_256:
6620 case Intrinsic::x86_avx_vtestz_pd:
6621 case Intrinsic::x86_avx_vtestz_pd_256:
6622 case Intrinsic::x86_avx_vtestz_ps:
6623 case Intrinsic::x86_avx_vtestz_ps_256:
6624 case Intrinsic::x86_avx_ptestc_256:
6625 case Intrinsic::x86_avx_ptestnzc_256:
6626 case Intrinsic::x86_avx_ptestz_256:
6627 case Intrinsic::x86_sse41_ptestc:
6628 case Intrinsic::x86_sse41_ptestnzc:
6629 case Intrinsic::x86_sse41_ptestz:
6630 handleVtestIntrinsic(I);
6631 break;
6632
6633 // Packed Horizontal Add/Subtract
6634 case Intrinsic::x86_ssse3_phadd_w:
6635 case Intrinsic::x86_ssse3_phadd_w_128:
6636 case Intrinsic::x86_ssse3_phsub_w:
6637 case Intrinsic::x86_ssse3_phsub_w_128:
6638 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/1,
6639 /*ReinterpretElemWidth=*/16);
6640 break;
6641
6642 case Intrinsic::x86_avx2_phadd_w:
6643 case Intrinsic::x86_avx2_phsub_w:
6644 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/2,
6645 /*ReinterpretElemWidth=*/16);
6646 break;
6647
6648 // Packed Horizontal Add/Subtract
6649 case Intrinsic::x86_ssse3_phadd_d:
6650 case Intrinsic::x86_ssse3_phadd_d_128:
6651 case Intrinsic::x86_ssse3_phsub_d:
6652 case Intrinsic::x86_ssse3_phsub_d_128:
6653 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/1,
6654 /*ReinterpretElemWidth=*/32);
6655 break;
6656
6657 case Intrinsic::x86_avx2_phadd_d:
6658 case Intrinsic::x86_avx2_phsub_d:
6659 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/2,
6660 /*ReinterpretElemWidth=*/32);
6661 break;
6662
6663 // Packed Horizontal Add/Subtract and Saturate
6664 case Intrinsic::x86_ssse3_phadd_sw:
6665 case Intrinsic::x86_ssse3_phadd_sw_128:
6666 case Intrinsic::x86_ssse3_phsub_sw:
6667 case Intrinsic::x86_ssse3_phsub_sw_128:
6668 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/1,
6669 /*ReinterpretElemWidth=*/16);
6670 break;
6671
6672 case Intrinsic::x86_avx2_phadd_sw:
6673 case Intrinsic::x86_avx2_phsub_sw:
6674 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/2,
6675 /*ReinterpretElemWidth=*/16);
6676 break;
6677
6678 // Packed Single/Double Precision Floating-Point Horizontal Add
6679 case Intrinsic::x86_sse3_hadd_ps:
6680 case Intrinsic::x86_sse3_hadd_pd:
6681 case Intrinsic::x86_sse3_hsub_ps:
6682 case Intrinsic::x86_sse3_hsub_pd:
6683 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/1);
6684 break;
6685
6686 case Intrinsic::x86_avx_hadd_pd_256:
6687 case Intrinsic::x86_avx_hadd_ps_256:
6688 case Intrinsic::x86_avx_hsub_pd_256:
6689 case Intrinsic::x86_avx_hsub_ps_256:
6690 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/2);
6691 break;
6692
6693 case Intrinsic::x86_avx_maskstore_ps:
6694 case Intrinsic::x86_avx_maskstore_pd:
6695 case Intrinsic::x86_avx_maskstore_ps_256:
6696 case Intrinsic::x86_avx_maskstore_pd_256:
6697 case Intrinsic::x86_avx2_maskstore_d:
6698 case Intrinsic::x86_avx2_maskstore_q:
6699 case Intrinsic::x86_avx2_maskstore_d_256:
6700 case Intrinsic::x86_avx2_maskstore_q_256: {
6701 handleAVXMaskedStore(I);
6702 break;
6703 }
6704
6705 case Intrinsic::x86_avx_maskload_ps:
6706 case Intrinsic::x86_avx_maskload_pd:
6707 case Intrinsic::x86_avx_maskload_ps_256:
6708 case Intrinsic::x86_avx_maskload_pd_256:
6709 case Intrinsic::x86_avx2_maskload_d:
6710 case Intrinsic::x86_avx2_maskload_q:
6711 case Intrinsic::x86_avx2_maskload_d_256:
6712 case Intrinsic::x86_avx2_maskload_q_256: {
6713 handleAVXMaskedLoad(I);
6714 break;
6715 }
6716
6717 // Packed
6718 case Intrinsic::x86_avx512fp16_add_ph_512:
6719 case Intrinsic::x86_avx512fp16_sub_ph_512:
6720 case Intrinsic::x86_avx512fp16_mul_ph_512:
6721 case Intrinsic::x86_avx512fp16_div_ph_512:
6722 case Intrinsic::x86_avx512fp16_max_ph_512:
6723 case Intrinsic::x86_avx512fp16_min_ph_512:
6724 case Intrinsic::x86_avx512_min_ps_512:
6725 case Intrinsic::x86_avx512_min_pd_512:
6726 case Intrinsic::x86_avx512_max_ps_512:
6727 case Intrinsic::x86_avx512_max_pd_512: {
6728 // These AVX512 variants contain the rounding mode as a trailing flag.
6729 // Earlier variants do not have a trailing flag and are already handled
6730 // by maybeHandleSimpleNomemIntrinsic(I, 0) via
6731 // maybeHandleUnknownIntrinsic.
6732 [[maybe_unused]] bool Success =
6733 maybeHandleSimpleNomemIntrinsic(I, /*trailingFlags=*/1);
6734 assert(Success);
6735 break;
6736 }
6737
6738 case Intrinsic::x86_avx_vpermilvar_pd:
6739 case Intrinsic::x86_avx_vpermilvar_pd_256:
6740 case Intrinsic::x86_avx512_vpermilvar_pd_512:
6741 case Intrinsic::x86_avx_vpermilvar_ps:
6742 case Intrinsic::x86_avx_vpermilvar_ps_256:
6743 case Intrinsic::x86_avx512_vpermilvar_ps_512: {
6744 handleAVXVpermilvar(I);
6745 break;
6746 }
6747
6748 case Intrinsic::x86_avx512_vpermi2var_d_128:
6749 case Intrinsic::x86_avx512_vpermi2var_d_256:
6750 case Intrinsic::x86_avx512_vpermi2var_d_512:
6751 case Intrinsic::x86_avx512_vpermi2var_hi_128:
6752 case Intrinsic::x86_avx512_vpermi2var_hi_256:
6753 case Intrinsic::x86_avx512_vpermi2var_hi_512:
6754 case Intrinsic::x86_avx512_vpermi2var_pd_128:
6755 case Intrinsic::x86_avx512_vpermi2var_pd_256:
6756 case Intrinsic::x86_avx512_vpermi2var_pd_512:
6757 case Intrinsic::x86_avx512_vpermi2var_ps_128:
6758 case Intrinsic::x86_avx512_vpermi2var_ps_256:
6759 case Intrinsic::x86_avx512_vpermi2var_ps_512:
6760 case Intrinsic::x86_avx512_vpermi2var_q_128:
6761 case Intrinsic::x86_avx512_vpermi2var_q_256:
6762 case Intrinsic::x86_avx512_vpermi2var_q_512:
6763 case Intrinsic::x86_avx512_vpermi2var_qi_128:
6764 case Intrinsic::x86_avx512_vpermi2var_qi_256:
6765 case Intrinsic::x86_avx512_vpermi2var_qi_512:
6766 handleAVXVpermi2var(I);
6767 break;
6768
6769 // Packed Shuffle
6770 // llvm.x86.sse.pshuf.w(<1 x i64>, i8)
6771 // llvm.x86.ssse3.pshuf.b(<1 x i64>, <1 x i64>)
6772 // llvm.x86.ssse3.pshuf.b.128(<16 x i8>, <16 x i8>)
6773 // llvm.x86.avx2.pshuf.b(<32 x i8>, <32 x i8>)
6774 // llvm.x86.avx512.pshuf.b.512(<64 x i8>, <64 x i8>)
6775 //
6776 // The following intrinsics are auto-upgraded:
6777 // llvm.x86.sse2.pshuf.d(<4 x i32>, i8)
6778 // llvm.x86.sse2.gpshufh.w(<8 x i16>, i8)
6779 // llvm.x86.sse2.pshufl.w(<8 x i16>, i8)
6780 case Intrinsic::x86_avx2_pshuf_b:
6781 case Intrinsic::x86_sse_pshuf_w:
6782 case Intrinsic::x86_ssse3_pshuf_b_128:
6783 case Intrinsic::x86_ssse3_pshuf_b:
6784 case Intrinsic::x86_avx512_pshuf_b_512:
6785 handleIntrinsicByApplyingToShadow(I, I.getIntrinsicID(),
6786 /*trailingVerbatimArgs=*/1,
6787 /*forceIntegerIntrinsic=*/false);
6788 break;
6789
6790 // AVX512 PMOV: Packed MOV, with truncation
6791 // Precisely handled by applying the same intrinsic to the shadow
6792 case Intrinsic::x86_avx512_mask_pmov_dw_128:
6793 case Intrinsic::x86_avx512_mask_pmov_db_128:
6794 case Intrinsic::x86_avx512_mask_pmov_qb_128:
6795 case Intrinsic::x86_avx512_mask_pmov_qw_128:
6796 case Intrinsic::x86_avx512_mask_pmov_qd_128:
6797 case Intrinsic::x86_avx512_mask_pmov_wb_128:
6798 case Intrinsic::x86_avx512_mask_pmov_dw_256:
6799 case Intrinsic::x86_avx512_mask_pmov_db_256:
6800 case Intrinsic::x86_avx512_mask_pmov_qb_256:
6801 case Intrinsic::x86_avx512_mask_pmov_qw_256:
6802 case Intrinsic::x86_avx512_mask_pmov_dw_512:
6803 case Intrinsic::x86_avx512_mask_pmov_db_512:
6804 case Intrinsic::x86_avx512_mask_pmov_qb_512:
6805 case Intrinsic::x86_avx512_mask_pmov_qw_512: {
6806 // Intrinsic::x86_avx512_mask_pmov_{qd,wb}_{256,512} were removed in
6807 // f608dc1f5775ee880e8ea30e2d06ab5a4a935c22
6808 handleIntrinsicByApplyingToShadow(I, I.getIntrinsicID(),
6809 /*trailingVerbatimArgs=*/1,
6810 /*forceIntegerIntrinsic=*/false);
6811 break;
6812 }
6813
6814 // AVX512 PMOV{S,US}: Packed MOV, with signed/unsigned saturation
6815 // Approximately handled using the corresponding truncation intrinsic
6816 // TODO: improve handleAVX512VectorDownConvert to precisely model saturation
6817 case Intrinsic::x86_avx512_mask_pmovs_dw_512:
6818 case Intrinsic::x86_avx512_mask_pmovus_dw_512: {
6819 handleIntrinsicByApplyingToShadow(
6820 I, Intrinsic::x86_avx512_mask_pmov_dw_512,
6821 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6822 break;
6823 }
6824
6825 case Intrinsic::x86_avx512_mask_pmovs_dw_256:
6826 case Intrinsic::x86_avx512_mask_pmovus_dw_256:
6827 handleIntrinsicByApplyingToShadow(
6828 I, Intrinsic::x86_avx512_mask_pmov_dw_256,
6829 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6830 break;
6831
6832 case Intrinsic::x86_avx512_mask_pmovs_dw_128:
6833 case Intrinsic::x86_avx512_mask_pmovus_dw_128:
6834 handleIntrinsicByApplyingToShadow(
6835 I, Intrinsic::x86_avx512_mask_pmov_dw_128,
6836 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6837 break;
6838
6839 case Intrinsic::x86_avx512_mask_pmovs_db_512:
6840 case Intrinsic::x86_avx512_mask_pmovus_db_512: {
6841 handleIntrinsicByApplyingToShadow(
6842 I, Intrinsic::x86_avx512_mask_pmov_db_512,
6843 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6844 break;
6845 }
6846
6847 case Intrinsic::x86_avx512_mask_pmovs_db_256:
6848 case Intrinsic::x86_avx512_mask_pmovus_db_256:
6849 handleIntrinsicByApplyingToShadow(
6850 I, Intrinsic::x86_avx512_mask_pmov_db_256,
6851 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6852 break;
6853
6854 case Intrinsic::x86_avx512_mask_pmovs_db_128:
6855 case Intrinsic::x86_avx512_mask_pmovus_db_128:
6856 handleIntrinsicByApplyingToShadow(
6857 I, Intrinsic::x86_avx512_mask_pmov_db_128,
6858 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6859 break;
6860
6861 case Intrinsic::x86_avx512_mask_pmovs_qb_512:
6862 case Intrinsic::x86_avx512_mask_pmovus_qb_512: {
6863 handleIntrinsicByApplyingToShadow(
6864 I, Intrinsic::x86_avx512_mask_pmov_qb_512,
6865 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6866 break;
6867 }
6868
6869 case Intrinsic::x86_avx512_mask_pmovs_qb_256:
6870 case Intrinsic::x86_avx512_mask_pmovus_qb_256:
6871 handleIntrinsicByApplyingToShadow(
6872 I, Intrinsic::x86_avx512_mask_pmov_qb_256,
6873 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6874 break;
6875
6876 case Intrinsic::x86_avx512_mask_pmovs_qb_128:
6877 case Intrinsic::x86_avx512_mask_pmovus_qb_128:
6878 handleIntrinsicByApplyingToShadow(
6879 I, Intrinsic::x86_avx512_mask_pmov_qb_128,
6880 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6881 break;
6882
6883 case Intrinsic::x86_avx512_mask_pmovs_qw_512:
6884 case Intrinsic::x86_avx512_mask_pmovus_qw_512: {
6885 handleIntrinsicByApplyingToShadow(
6886 I, Intrinsic::x86_avx512_mask_pmov_qw_512,
6887 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6888 break;
6889 }
6890
6891 case Intrinsic::x86_avx512_mask_pmovs_qw_256:
6892 case Intrinsic::x86_avx512_mask_pmovus_qw_256:
6893 handleIntrinsicByApplyingToShadow(
6894 I, Intrinsic::x86_avx512_mask_pmov_qw_256,
6895 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6896 break;
6897
6898 case Intrinsic::x86_avx512_mask_pmovs_qw_128:
6899 case Intrinsic::x86_avx512_mask_pmovus_qw_128:
6900 handleIntrinsicByApplyingToShadow(
6901 I, Intrinsic::x86_avx512_mask_pmov_qw_128,
6902 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6903 break;
6904
6905 case Intrinsic::x86_avx512_mask_pmovs_qd_128:
6906 case Intrinsic::x86_avx512_mask_pmovus_qd_128:
6907 handleIntrinsicByApplyingToShadow(
6908 I, Intrinsic::x86_avx512_mask_pmov_qd_128,
6909 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6910 break;
6911
6912 case Intrinsic::x86_avx512_mask_pmovs_wb_128:
6913 case Intrinsic::x86_avx512_mask_pmovus_wb_128:
6914 handleIntrinsicByApplyingToShadow(
6915 I, Intrinsic::x86_avx512_mask_pmov_wb_128,
6916 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
6917 break;
6918
6919 case Intrinsic::x86_avx512_mask_pmovs_qd_256:
6920 case Intrinsic::x86_avx512_mask_pmovus_qd_256:
6921 case Intrinsic::x86_avx512_mask_pmovs_wb_256:
6922 case Intrinsic::x86_avx512_mask_pmovus_wb_256:
6923 case Intrinsic::x86_avx512_mask_pmovs_qd_512:
6924 case Intrinsic::x86_avx512_mask_pmovus_qd_512:
6925 case Intrinsic::x86_avx512_mask_pmovs_wb_512:
6926 case Intrinsic::x86_avx512_mask_pmovus_wb_512: {
6927 // Since Intrinsic::x86_avx512_mask_pmov_{qd,wb}_{256,512} do not exist,
6928 // we cannot use handleIntrinsicByApplyingToShadow. Instead, we call the
6929 // slow-path handler.
6930 handleAVX512VectorDownConvert(I);
6931 break;
6932 }
6933
6934 // e.g.,
6935 // <16 x float> @llvm.x86.avx512.mask.compress
6936 // (<16 x float> %data, <16 x float> %passthru,
6937 // <16 x i1> %mask)
6938 // <16 x i32> @llvm.x86.avx512.mask.compress
6939 // (<16 x i32> %data, <16 x i32> %passthru,
6940 // <16 x i1> %mask)
6941 case Intrinsic::x86_avx512_mask_compress:
6942 handleIntrinsicByApplyingToShadow(I, I.getIntrinsicID(),
6943 /*trailingVerbatimArgs=*/1,
6944 /*forceIntegerIntrinsic=*/true);
6945 break;
6946
6947 // AVX512/AVX10 Reciprocal
6948 // <16 x float> @llvm.x86.avx512.rsqrt14.ps.512
6949 // (<16 x float>, <16 x float>, i16)
6950 // <8 x float> @llvm.x86.avx512.rsqrt14.ps.256
6951 // (<8 x float>, <8 x float>, i8)
6952 // <4 x float> @llvm.x86.avx512.rsqrt14.ps.128
6953 // (<4 x float>, <4 x float>, i8)
6954 //
6955 // <8 x double> @llvm.x86.avx512.rsqrt14.pd.512
6956 // (<8 x double>, <8 x double>, i8)
6957 // <4 x double> @llvm.x86.avx512.rsqrt14.pd.256
6958 // (<4 x double>, <4 x double>, i8)
6959 // <2 x double> @llvm.x86.avx512.rsqrt14.pd.128
6960 // (<2 x double>, <2 x double>, i8)
6961 //
6962 // <32 x bfloat> @llvm.x86.avx10.mask.rsqrt.bf16.512
6963 // (<32 x bfloat>, <32 x bfloat>, i32)
6964 // <16 x bfloat> @llvm.x86.avx10.mask.rsqrt.bf16.256
6965 // (<16 x bfloat>, <16 x bfloat>, i16)
6966 // <8 x bfloat> @llvm.x86.avx10.mask.rsqrt.bf16.128
6967 // (<8 x bfloat>, <8 x bfloat>, i8)
6968 //
6969 // <32 x half> @llvm.x86.avx512fp16.mask.rsqrt.ph.512
6970 // (<32 x half>, <32 x half>, i32)
6971 // <16 x half> @llvm.x86.avx512fp16.mask.rsqrt.ph.256
6972 // (<16 x half>, <16 x half>, i16)
6973 // <8 x half> @llvm.x86.avx512fp16.mask.rsqrt.ph.128
6974 // (<8 x half>, <8 x half>, i8)
6975 //
6976 // TODO: 3-operand variants are not handled:
6977 // <2 x double> @llvm.x86.avx512.rsqrt14.sd
6978 // (<2 x double>, <2 x double>, <2 x double>, i8)
6979 // <4 x float> @llvm.x86.avx512.rsqrt14.ss
6980 // (<4 x float>, <4 x float>, <4 x float>, i8)
6981 // <8 x half> @llvm.x86.avx512fp16.mask.rsqrt.sh
6982 // (<8 x half>, <8 x half>, <8 x half>, i8)
6983 case Intrinsic::x86_avx512_rsqrt14_ps_512:
6984 case Intrinsic::x86_avx512_rsqrt14_ps_256:
6985 case Intrinsic::x86_avx512_rsqrt14_ps_128:
6986 case Intrinsic::x86_avx512_rsqrt14_pd_512:
6987 case Intrinsic::x86_avx512_rsqrt14_pd_256:
6988 case Intrinsic::x86_avx512_rsqrt14_pd_128:
6989 case Intrinsic::x86_avx10_mask_rsqrt_bf16_512:
6990 case Intrinsic::x86_avx10_mask_rsqrt_bf16_256:
6991 case Intrinsic::x86_avx10_mask_rsqrt_bf16_128:
6992 case Intrinsic::x86_avx512fp16_mask_rsqrt_ph_512:
6993 case Intrinsic::x86_avx512fp16_mask_rsqrt_ph_256:
6994 case Intrinsic::x86_avx512fp16_mask_rsqrt_ph_128:
6995 handleAVX512VectorGenericMaskedFP(I, /*DataIndices=*/{0},
6996 /*WriteThruIndex=*/1,
6997 /*MaskIndex=*/2);
6998 break;
6999
7000 // AVX512/AVX10 Reciprocal Square Root
7001 // <16 x float> @llvm.x86.avx512.rcp14.ps.512
7002 // (<16 x float>, <16 x float>, i16)
7003 // <8 x float> @llvm.x86.avx512.rcp14.ps.256
7004 // (<8 x float>, <8 x float>, i8)
7005 // <4 x float> @llvm.x86.avx512.rcp14.ps.128
7006 // (<4 x float>, <4 x float>, i8)
7007 //
7008 // <8 x double> @llvm.x86.avx512.rcp14.pd.512
7009 // (<8 x double>, <8 x double>, i8)
7010 // <4 x double> @llvm.x86.avx512.rcp14.pd.256
7011 // (<4 x double>, <4 x double>, i8)
7012 // <2 x double> @llvm.x86.avx512.rcp14.pd.128
7013 // (<2 x double>, <2 x double>, i8)
7014 //
7015 // <32 x bfloat> @llvm.x86.avx10.mask.rcp.bf16.512
7016 // (<32 x bfloat>, <32 x bfloat>, i32)
7017 // <16 x bfloat> @llvm.x86.avx10.mask.rcp.bf16.256
7018 // (<16 x bfloat>, <16 x bfloat>, i16)
7019 // <8 x bfloat> @llvm.x86.avx10.mask.rcp.bf16.128
7020 // (<8 x bfloat>, <8 x bfloat>, i8)
7021 //
7022 // <32 x half> @llvm.x86.avx512fp16.mask.rcp.ph.512
7023 // (<32 x half>, <32 x half>, i32)
7024 // <16 x half> @llvm.x86.avx512fp16.mask.rcp.ph.256
7025 // (<16 x half>, <16 x half>, i16)
7026 // <8 x half> @llvm.x86.avx512fp16.mask.rcp.ph.128
7027 // (<8 x half>, <8 x half>, i8)
7028 //
7029 // TODO: 3-operand variants are not handled:
7030 // <2 x double> @llvm.x86.avx512.rcp14.sd
7031 // (<2 x double>, <2 x double>, <2 x double>, i8)
7032 // <4 x float> @llvm.x86.avx512.rcp14.ss
7033 // (<4 x float>, <4 x float>, <4 x float>, i8)
7034 // <8 x half> @llvm.x86.avx512fp16.mask.rcp.sh
7035 // (<8 x half>, <8 x half>, <8 x half>, i8)
7036 case Intrinsic::x86_avx512_rcp14_ps_512:
7037 case Intrinsic::x86_avx512_rcp14_ps_256:
7038 case Intrinsic::x86_avx512_rcp14_ps_128:
7039 case Intrinsic::x86_avx512_rcp14_pd_512:
7040 case Intrinsic::x86_avx512_rcp14_pd_256:
7041 case Intrinsic::x86_avx512_rcp14_pd_128:
7042 case Intrinsic::x86_avx10_mask_rcp_bf16_512:
7043 case Intrinsic::x86_avx10_mask_rcp_bf16_256:
7044 case Intrinsic::x86_avx10_mask_rcp_bf16_128:
7045 case Intrinsic::x86_avx512fp16_mask_rcp_ph_512:
7046 case Intrinsic::x86_avx512fp16_mask_rcp_ph_256:
7047 case Intrinsic::x86_avx512fp16_mask_rcp_ph_128:
7048 handleAVX512VectorGenericMaskedFP(I, /*DataIndices=*/{0},
7049 /*WriteThruIndex=*/1,
7050 /*MaskIndex=*/2);
7051 break;
7052
7053 // <32 x half> @llvm.x86.avx512fp16.mask.rndscale.ph.512
7054 // (<32 x half>, i32, <32 x half>, i32, i32)
7055 // <16 x half> @llvm.x86.avx512fp16.mask.rndscale.ph.256
7056 // (<16 x half>, i32, <16 x half>, i32, i16)
7057 // <8 x half> @llvm.x86.avx512fp16.mask.rndscale.ph.128
7058 // (<8 x half>, i32, <8 x half>, i32, i8)
7059 //
7060 // <16 x float> @llvm.x86.avx512.mask.rndscale.ps.512
7061 // (<16 x float>, i32, <16 x float>, i16, i32)
7062 // <8 x float> @llvm.x86.avx512.mask.rndscale.ps.256
7063 // (<8 x float>, i32, <8 x float>, i8)
7064 // <4 x float> @llvm.x86.avx512.mask.rndscale.ps.128
7065 // (<4 x float>, i32, <4 x float>, i8)
7066 //
7067 // <8 x double> @llvm.x86.avx512.mask.rndscale.pd.512
7068 // (<8 x double>, i32, <8 x double>, i8, i32)
7069 // A Imm WriteThru Mask Rounding
7070 // <4 x double> @llvm.x86.avx512.mask.rndscale.pd.256
7071 // (<4 x double>, i32, <4 x double>, i8)
7072 // <2 x double> @llvm.x86.avx512.mask.rndscale.pd.128
7073 // (<2 x double>, i32, <2 x double>, i8)
7074 // A Imm WriteThru Mask
7075 //
7076 // <32 x bfloat> @llvm.x86.avx10.mask.rndscale.bf16.512
7077 // (<32 x bfloat>, i32, <32 x bfloat>, i32)
7078 // <16 x bfloat> @llvm.x86.avx10.mask.rndscale.bf16.256
7079 // (<16 x bfloat>, i32, <16 x bfloat>, i16)
7080 // <8 x bfloat> @llvm.x86.avx10.mask.rndscale.bf16.128
7081 // (<8 x bfloat>, i32, <8 x bfloat>, i8)
7082 //
7083 // Not supported: three vectors
7084 // - <8 x half> @llvm.x86.avx512fp16.mask.rndscale.sh
7085 // (<8 x half>, <8 x half>,<8 x half>, i8, i32, i32)
7086 // - <4 x float> @llvm.x86.avx512.mask.rndscale.ss
7087 // (<4 x float>, <4 x float>, <4 x float>, i8, i32, i32)
7088 // - <2 x double> @llvm.x86.avx512.mask.rndscale.sd
7089 // (<2 x double>, <2 x double>, <2 x double>, i8, i32,
7090 // i32)
7091 // A B WriteThru Mask Imm
7092 // Rounding
7093 case Intrinsic::x86_avx512fp16_mask_rndscale_ph_512:
7094 case Intrinsic::x86_avx512fp16_mask_rndscale_ph_256:
7095 case Intrinsic::x86_avx512fp16_mask_rndscale_ph_128:
7096 case Intrinsic::x86_avx512_mask_rndscale_ps_512:
7097 case Intrinsic::x86_avx512_mask_rndscale_ps_256:
7098 case Intrinsic::x86_avx512_mask_rndscale_ps_128:
7099 case Intrinsic::x86_avx512_mask_rndscale_pd_512:
7100 case Intrinsic::x86_avx512_mask_rndscale_pd_256:
7101 case Intrinsic::x86_avx512_mask_rndscale_pd_128:
7102 case Intrinsic::x86_avx10_mask_rndscale_bf16_512:
7103 case Intrinsic::x86_avx10_mask_rndscale_bf16_256:
7104 case Intrinsic::x86_avx10_mask_rndscale_bf16_128:
7105 handleAVX512VectorGenericMaskedFP(I, /*DataIndices=*/{0},
7106 /*WriteThruIndex=*/2,
7107 /*MaskIndex=*/3);
7108 break;
7109
7110 // AVX512 Vector Scale Float* Packed
7111 //
7112 // < 8 x double> @llvm.x86.avx512.mask.scalef.pd.512
7113 // (<8 x double>, <8 x double>, <8 x double>, i8, i32)
7114 // A B WriteThru Msk Round
7115 // < 4 x double> @llvm.x86.avx512.mask.scalef.pd.256
7116 // (<4 x double>, <4 x double>, <4 x double>, i8)
7117 // < 2 x double> @llvm.x86.avx512.mask.scalef.pd.128
7118 // (<2 x double>, <2 x double>, <2 x double>, i8)
7119 //
7120 // <16 x float> @llvm.x86.avx512.mask.scalef.ps.512
7121 // (<16 x float>, <16 x float>, <16 x float>, i16, i32)
7122 // < 8 x float> @llvm.x86.avx512.mask.scalef.ps.256
7123 // (<8 x float>, <8 x float>, <8 x float>, i8)
7124 // < 4 x float> @llvm.x86.avx512.mask.scalef.ps.128
7125 // (<4 x float>, <4 x float>, <4 x float>, i8)
7126 //
7127 // <32 x half> @llvm.x86.avx512fp16.mask.scalef.ph.512
7128 // (<32 x half>, <32 x half>, <32 x half>, i32, i32)
7129 // <16 x half> @llvm.x86.avx512fp16.mask.scalef.ph.256
7130 // (<16 x half>, <16 x half>, <16 x half>, i16)
7131 // < 8 x half> @llvm.x86.avx512fp16.mask.scalef.ph.128
7132 // (<8 x half>, <8 x half>, <8 x half>, i8)
7133 //
7134 // TODO: AVX10
7135 // <32 x bfloat> @llvm.x86.avx10.mask.scalef.bf16.512
7136 // (<32 x bfloat>, <32 x bfloat>, <32 x bfloat>, i32)
7137 // <16 x bfloat> @llvm.x86.avx10.mask.scalef.bf16.256
7138 // (<16 x bfloat>, <16 x bfloat>, <16 x bfloat>, i16)
7139 // < 8 x bfloat> @llvm.x86.avx10.mask.scalef.bf16.128
7140 // (<8 x bfloat>, <8 x bfloat>, <8 x bfloat>, i8)
7141 case Intrinsic::x86_avx512_mask_scalef_pd_512:
7142 case Intrinsic::x86_avx512_mask_scalef_pd_256:
7143 case Intrinsic::x86_avx512_mask_scalef_pd_128:
7144 case Intrinsic::x86_avx512_mask_scalef_ps_512:
7145 case Intrinsic::x86_avx512_mask_scalef_ps_256:
7146 case Intrinsic::x86_avx512_mask_scalef_ps_128:
7147 case Intrinsic::x86_avx512fp16_mask_scalef_ph_512:
7148 case Intrinsic::x86_avx512fp16_mask_scalef_ph_256:
7149 case Intrinsic::x86_avx512fp16_mask_scalef_ph_128:
7150 // The AVX512 512-bit operand variants have an extra operand (the
7151 // Rounding mode). The extra operand, if present, will be
7152 // automatically checked by the handler.
7153 handleAVX512VectorGenericMaskedFP(I, /*DataIndices=*/{0, 1},
7154 /*WriteThruIndex=*/2,
7155 /*MaskIndex=*/3);
7156 break;
7157
7158 // TODO: AVX512 Vector Scale Float* Scalar
7159 //
7160 // This is different from the Packed variant, because some bits are copied,
7161 // and some bits are zeroed.
7162 //
7163 // < 4 x float> @llvm.x86.avx512.mask.scalef.ss
7164 // (<4 x float>, <4 x float>, <4 x float>, i8, i32)
7165 //
7166 // < 2 x double> @llvm.x86.avx512.mask.scalef.sd
7167 // (<2 x double>, <2 x double>, <2 x double>, i8, i32)
7168 //
7169 // < 8 x half> @llvm.x86.avx512fp16.mask.scalef.sh
7170 // (<8 x half>, <8 x half>, <8 x half>, i8, i32)
7171
7172 // AVX512 FP16 Arithmetic
7173 case Intrinsic::x86_avx512fp16_mask_add_sh_round:
7174 case Intrinsic::x86_avx512fp16_mask_sub_sh_round:
7175 case Intrinsic::x86_avx512fp16_mask_mul_sh_round:
7176 case Intrinsic::x86_avx512fp16_mask_div_sh_round:
7177 case Intrinsic::x86_avx512fp16_mask_max_sh_round:
7178 case Intrinsic::x86_avx512fp16_mask_min_sh_round: {
7179 visitGenericScalarHalfwordInst(I);
7180 break;
7181 }
7182
7183 // AVX512 Floating-Point Classification
7184 // - <8 x i1> @llvm.x86.avx512.fpclass.pd.512(<8 x double>, i32)
7185 // - <16 x i1> @llvm.x86.avx512.fpclass.ps.512(<16 x float>, i32)
7186 case Intrinsic::x86_avx512_fpclass_pd_512:
7187 case Intrinsic::x86_avx512_fpclass_ps_512:
7188 handleAVX512FPClass(I);
7189 break;
7190
7191 // AVX Galois Field New Instructions
7192 case Intrinsic::x86_vgf2p8affineqb_128:
7193 case Intrinsic::x86_vgf2p8affineqb_256:
7194 case Intrinsic::x86_vgf2p8affineqb_512:
7195 handleAVXGF2P8Affine(I);
7196 break;
7197
7198 default:
7199 return false;
7200 }
7201
7202 return true;
7203 }
7204
7205 bool maybeHandleArmSIMDIntrinsic(IntrinsicInst &I) {
7206 switch (I.getIntrinsicID()) {
7207 // Two operands e.g.,
7208 // - <8 x i8> @llvm.aarch64.neon.rshrn.v8i8 (<8 x i16>, i32)
7209 // - <4 x i16> @llvm.aarch64.neon.uqrshl.v4i16(<4 x i16>, <4 x i16>)
7210 case Intrinsic::aarch64_neon_rshrn:
7211 case Intrinsic::aarch64_neon_sqrshl:
7212 case Intrinsic::aarch64_neon_sqrshrn:
7213 case Intrinsic::aarch64_neon_sqrshrun:
7214 case Intrinsic::aarch64_neon_sqshl:
7215 case Intrinsic::aarch64_neon_sqshlu:
7216 case Intrinsic::aarch64_neon_sqshrn:
7217 case Intrinsic::aarch64_neon_sqshrun:
7218 case Intrinsic::aarch64_neon_srshl:
7219 case Intrinsic::aarch64_neon_sshl:
7220 case Intrinsic::aarch64_neon_uqrshl:
7221 case Intrinsic::aarch64_neon_uqrshrn:
7222 case Intrinsic::aarch64_neon_uqshl:
7223 case Intrinsic::aarch64_neon_uqshrn:
7224 case Intrinsic::aarch64_neon_urshl:
7225 case Intrinsic::aarch64_neon_ushl:
7226 handleVectorShiftIntrinsic(I, /* Variable */ false);
7227 break;
7228
7229 // Vector Shift Left/Right and Insert
7230 //
7231 // Three operands e.g.,
7232 // - <4 x i16> @llvm.aarch64.neon.vsli.v4i16
7233 // (<4 x i16> %a, <4 x i16> %b, i32 %n)
7234 // - <16 x i8> @llvm.aarch64.neon.vsri.v16i8
7235 // (<16 x i8> %a, <16 x i8> %b, i32 %n)
7236 //
7237 // %b is shifted by %n bits, and the "missing" bits are filled in with %a
7238 // (instead of zero-extending/sign-extending).
7239 case Intrinsic::aarch64_neon_vsli:
7240 case Intrinsic::aarch64_neon_vsri:
7241 handleIntrinsicByApplyingToShadow(I, I.getIntrinsicID(),
7242 /*trailingVerbatimArgs=*/1,
7243 /*forceIntegerIntrinsic=*/false);
7244 break;
7245
7246 // TODO: handling max/min similarly to AND/OR may be more precise
7247 // Floating-Point Maximum/Minimum Pairwise
7248 case Intrinsic::aarch64_neon_fmaxp:
7249 case Intrinsic::aarch64_neon_fminp:
7250 // Floating-Point Maximum/Minimum Number Pairwise
7251 case Intrinsic::aarch64_neon_fmaxnmp:
7252 case Intrinsic::aarch64_neon_fminnmp:
7253 // Signed/Unsigned Maximum/Minimum Pairwise
7254 case Intrinsic::aarch64_neon_smaxp:
7255 case Intrinsic::aarch64_neon_sminp:
7256 case Intrinsic::aarch64_neon_umaxp:
7257 case Intrinsic::aarch64_neon_uminp:
7258 // Add Pairwise
7259 case Intrinsic::aarch64_neon_addp:
7260 // Floating-point Add Pairwise
7261 case Intrinsic::aarch64_neon_faddp:
7262 // Add Long Pairwise
7263 case Intrinsic::aarch64_neon_saddlp:
7264 case Intrinsic::aarch64_neon_uaddlp: {
7265 handlePairwiseShadowOrIntrinsic(I, /*Shards=*/1);
7266 break;
7267 }
7268
7269 // Floating-point Convert to integer, rounding to nearest with ties to Away
7270 case Intrinsic::aarch64_neon_fcvtas:
7271 case Intrinsic::aarch64_neon_fcvtau:
7272 // Floating-point convert to integer, rounding toward minus infinity
7273 case Intrinsic::aarch64_neon_fcvtms:
7274 case Intrinsic::aarch64_neon_fcvtmu:
7275 // Floating-point convert to integer, rounding to nearest with ties to even
7276 case Intrinsic::aarch64_neon_fcvtns:
7277 case Intrinsic::aarch64_neon_fcvtnu:
7278 // Floating-point convert to integer, rounding toward plus infinity
7279 case Intrinsic::aarch64_neon_fcvtps:
7280 case Intrinsic::aarch64_neon_fcvtpu:
7281 // Floating-point Convert to integer, rounding toward Zero
7282 case Intrinsic::aarch64_neon_fcvtzs:
7283 case Intrinsic::aarch64_neon_fcvtzu:
7284 // Floating-point convert to lower precision narrow, rounding to odd
7285 case Intrinsic::aarch64_neon_fcvtxn:
7286 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/false);
7287 break;
7288
7289 // Vector Conversions Between Fixed-Point and Floating-Point
7290 case Intrinsic::aarch64_neon_vcvtfxs2fp:
7291 case Intrinsic::aarch64_neon_vcvtfp2fxs:
7292 case Intrinsic::aarch64_neon_vcvtfxu2fp:
7293 case Intrinsic::aarch64_neon_vcvtfp2fxu:
7294 handleGenericVectorConvertIntrinsic(I, /*FixedPoint=*/true);
7295 break;
7296
7297 // TODO: bfloat conversions
7298 // - bfloat @llvm.aarch64.neon.bfcvt(float)
7299 // - <8 x bfloat> @llvm.aarch64.neon.bfcvtn(<4 x float>)
7300 // - <8 x bfloat> @llvm.aarch64.neon.bfcvtn2(<8 x bfloat>, <4 x float>)
7301
7302 // Add reduction to scalar
7303 case Intrinsic::aarch64_neon_faddv:
7304 case Intrinsic::aarch64_neon_saddv:
7305 case Intrinsic::aarch64_neon_uaddv:
7306 // Signed/Unsigned min/max (Vector)
7307 // TODO: handling similarly to AND/OR may be more precise.
7308 case Intrinsic::aarch64_neon_smaxv:
7309 case Intrinsic::aarch64_neon_sminv:
7310 case Intrinsic::aarch64_neon_umaxv:
7311 case Intrinsic::aarch64_neon_uminv:
7312 // Floating-point min/max (vector)
7313 // The f{min,max}"nm"v variants handle NaN differently than f{min,max}v,
7314 // but our shadow propagation is the same.
7315 case Intrinsic::aarch64_neon_fmaxv:
7316 case Intrinsic::aarch64_neon_fminv:
7317 case Intrinsic::aarch64_neon_fmaxnmv:
7318 case Intrinsic::aarch64_neon_fminnmv:
7319 // Sum long across vector
7320 case Intrinsic::aarch64_neon_saddlv:
7321 case Intrinsic::aarch64_neon_uaddlv:
7322 handleVectorReduceIntrinsic(I, /*AllowShadowCast=*/true);
7323 break;
7324
7325 case Intrinsic::aarch64_neon_ld1x2:
7326 case Intrinsic::aarch64_neon_ld1x3:
7327 case Intrinsic::aarch64_neon_ld1x4:
7328 case Intrinsic::aarch64_neon_ld2:
7329 case Intrinsic::aarch64_neon_ld3:
7330 case Intrinsic::aarch64_neon_ld4:
7331 case Intrinsic::aarch64_neon_ld2r:
7332 case Intrinsic::aarch64_neon_ld3r:
7333 case Intrinsic::aarch64_neon_ld4r: {
7334 handleNEONVectorLoad(I, /*WithLane=*/false);
7335 break;
7336 }
7337
7338 case Intrinsic::aarch64_neon_ld2lane:
7339 case Intrinsic::aarch64_neon_ld3lane:
7340 case Intrinsic::aarch64_neon_ld4lane: {
7341 handleNEONVectorLoad(I, /*WithLane=*/true);
7342 break;
7343 }
7344
7345 // Saturating extract narrow
7346 case Intrinsic::aarch64_neon_sqxtn:
7347 case Intrinsic::aarch64_neon_sqxtun:
7348 case Intrinsic::aarch64_neon_uqxtn:
7349 // These only have one argument, but we (ab)use handleShadowOr because it
7350 // does work on single argument intrinsics and will typecast the shadow
7351 // (and update the origin).
7352 handleShadowOr(I);
7353 break;
7354
7355 case Intrinsic::aarch64_neon_st1x2:
7356 case Intrinsic::aarch64_neon_st1x3:
7357 case Intrinsic::aarch64_neon_st1x4:
7358 case Intrinsic::aarch64_neon_st2:
7359 case Intrinsic::aarch64_neon_st3:
7360 case Intrinsic::aarch64_neon_st4: {
7361 handleNEONVectorStoreIntrinsic(I, false);
7362 break;
7363 }
7364
7365 case Intrinsic::aarch64_neon_st2lane:
7366 case Intrinsic::aarch64_neon_st3lane:
7367 case Intrinsic::aarch64_neon_st4lane: {
7368 handleNEONVectorStoreIntrinsic(I, true);
7369 break;
7370 }
7371
7372 // Arm NEON vector table intrinsics have the source/table register(s) as
7373 // arguments, followed by the index register. They return the output.
7374 //
7375 // 'TBL writes a zero if an index is out-of-range, while TBX leaves the
7376 // original value unchanged in the destination register.'
7377 // Conveniently, zero denotes a clean shadow, which means out-of-range
7378 // indices for TBL will initialize the user data with zero and also clean
7379 // the shadow. (For TBX, neither the user data nor the shadow will be
7380 // updated, which is also correct.)
7381 case Intrinsic::aarch64_neon_tbl1:
7382 case Intrinsic::aarch64_neon_tbl2:
7383 case Intrinsic::aarch64_neon_tbl3:
7384 case Intrinsic::aarch64_neon_tbl4:
7385 case Intrinsic::aarch64_neon_tbx1:
7386 case Intrinsic::aarch64_neon_tbx2:
7387 case Intrinsic::aarch64_neon_tbx3:
7388 case Intrinsic::aarch64_neon_tbx4: {
7389 // The last trailing argument (index register) should be handled verbatim
7390 handleIntrinsicByApplyingToShadow(
7391 I, /*shadowIntrinsicID=*/I.getIntrinsicID(),
7392 /*trailingVerbatimArgs=*/1, /*forceIntegerIntrinsic=*/false);
7393 break;
7394 }
7395
7396 case Intrinsic::aarch64_neon_fmulx:
7397 case Intrinsic::aarch64_neon_pmul:
7398 case Intrinsic::aarch64_neon_pmull:
7399 case Intrinsic::aarch64_neon_smull:
7400 case Intrinsic::aarch64_neon_pmull64:
7401 case Intrinsic::aarch64_neon_umull: {
7402 handleNEONVectorMultiplyIntrinsic(I);
7403 break;
7404 }
7405
7406 case Intrinsic::aarch64_neon_smmla:
7407 case Intrinsic::aarch64_neon_ummla:
7408 case Intrinsic::aarch64_neon_usmmla:
7409 case Intrinsic::aarch64_neon_bfmmla:
7410 handleNEONMatrixMultiply(I);
7411 break;
7412
7413 // <2 x i32> @llvm.aarch64.neon.{u,s,us}dot.v2i32.v8i8
7414 // (<2 x i32> %acc, <8 x i8> %a, <8 x i8> %b)
7415 // <4 x i32> @llvm.aarch64.neon.{u,s,us}dot.v4i32.v16i8
7416 // (<4 x i32> %acc, <16 x i8> %a, <16 x i8> %b)
7417 case Intrinsic::aarch64_neon_sdot:
7418 case Intrinsic::aarch64_neon_udot:
7419 case Intrinsic::aarch64_neon_usdot:
7420 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/4,
7421 /*ZeroPurifies=*/true,
7422 /*EltSizeInBits=*/0,
7423 /*Lanes=*/kBothLanes);
7424 break;
7425
7426 // <2 x float> @llvm.aarch64.neon.bfdot.v2f32.v4bf16
7427 // (<2 x float> %acc, <4 x bfloat> %a, <4 x bfloat> %b)
7428 // <4 x float> @llvm.aarch64.neon.bfdot.v4f32.v8bf16
7429 // (<4 x float> %acc, <8 x bfloat> %a, <8 x bfloat> %b)
7430 case Intrinsic::aarch64_neon_bfdot:
7431 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
7432 /*ZeroPurifies=*/false,
7433 /*EltSizeInBits=*/0,
7434 /*Lanes=*/kBothLanes);
7435 break;
7436
7437 // <4 x half > @llvm.aarch64.neon.fp8.fdot2
7438 // (<4 x half >, < 8 x i8>, < 8 x i8>)
7439 // <8 x half > @llvm.aarch64.neon.fp8.fdot2
7440 // (<8 x half >, <16 x i8>, <16 x i8>)
7441 //
7442 // N.B. although the multiplicands are i8, they are actually fp8, thus
7443 // ZeroPurifies is not applicable.
7444 case Intrinsic::aarch64_neon_fp8_fdot2:
7445 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/2,
7446 /*ZeroPurifies=*/false,
7447 /*EltSizeInBits=*/0,
7448 /*Lanes=*/kBothLanes);
7449 break;
7450
7451 // <2 x float> @llvm.aarch64.neon.fp8.fdot4
7452 // (<2 x float>, < 8 x i8>, < 8 x i8>)
7453 // <4 x float> @llvm.aarch64.neon.fp8.fdot4
7454 // (<4 x float>, <16 x i8>, <16 x i8>)
7455 //
7456 // N.B. although the multiplicands are i8, they are actually fp8, thus
7457 // ZeroPurifies is not applicable.
7458 case Intrinsic::aarch64_neon_fp8_fdot4:
7459 handleVectorDotProductIntrinsic(I, /*ReductionFactor=*/4,
7460 /*ZeroPurifies=*/false,
7461 /*EltSizeInBits=*/0,
7462 /*Lanes=*/kBothLanes);
7463 break;
7464
7465 // <4 x half> @llvm.aarch64.neon.fp8.fdot2.lane
7466 // (<4 x half>, <8 x i8>, <16 x i8>, i32)
7467 // <8 x half> @llvm.aarch64.neon.fp8.fdot2.lane
7468 // (<8 x half>, <16 x i8>, <16 x i8>, i32)
7469 case Intrinsic::aarch64_neon_fp8_fdot2_lane:
7470 handleNEONDotProductLaneIntrinsic(I, /*ReductionFactor=*/2);
7471 break;
7472
7473 // <2 x float> @llvm.aarch64.neon.fp8.fdot4.lane
7474 // (<2 x float>, <8 x i8>, <16 x i8>, i32)
7475 // <4 x float> @llvm.aarch64.neon.fp8.fdot4.lane
7476 // (<4 x float>, <16 x i8>, <16 x i8>, i32)
7477 case Intrinsic::aarch64_neon_fp8_fdot4_lane:
7478 handleNEONDotProductLaneIntrinsic(I, /*ReductionFactor=*/4);
7479 break;
7480
7481 // Floating-Point Absolute Compare Greater Than/Equal
7482 case Intrinsic::aarch64_neon_facge:
7483 case Intrinsic::aarch64_neon_facgt:
7484 handleVectorComparePackedIntrinsic(I, /*PredicateAsOperand=*/false);
7485 break;
7486
7487 default:
7488 return false;
7489 }
7490
7491 return true;
7492 }
7493
7494 void visitIntrinsicInst(IntrinsicInst &I) {
7495 if (maybeHandleCrossPlatformIntrinsic(I))
7496 return;
7497
7498 if (maybeHandleX86SIMDIntrinsic(I))
7499 return;
7500
7501 if (maybeHandleArmSIMDIntrinsic(I))
7502 return;
7503
7504 if (maybeHandleUnknownIntrinsic(I))
7505 return;
7506
7507 visitInstruction(I);
7508 }
7509
7510 void visitLibAtomicLoad(CallBase &CB) {
7511 // Since we use getNextNode here, we can't have CB terminate the BB.
7512 assert(isa<CallInst>(CB));
7513
7514 IRBuilder<> IRB(&CB);
7515 Value *Size = CB.getArgOperand(0);
7516 Value *SrcPtr = CB.getArgOperand(1);
7517 Value *DstPtr = CB.getArgOperand(2);
7518 Value *Ordering = CB.getArgOperand(3);
7519 // Convert the call to have at least Acquire ordering to make sure
7520 // the shadow operations aren't reordered before it.
7521 Value *NewOrdering =
7522 IRB.CreateExtractElement(makeAddAcquireOrderingTable(IRB), Ordering);
7523 CB.setArgOperand(3, NewOrdering);
7524
7525 NextNodeIRBuilder NextIRB(&CB);
7526 Value *SrcShadowPtr, *SrcOriginPtr;
7527 std::tie(SrcShadowPtr, SrcOriginPtr) =
7528 getShadowOriginPtr(SrcPtr, NextIRB, NextIRB.getInt8Ty(), Align(1),
7529 /*isStore*/ false);
7530 Value *DstShadowPtr =
7531 getShadowOriginPtr(DstPtr, NextIRB, NextIRB.getInt8Ty(), Align(1),
7532 /*isStore*/ true)
7533 .first;
7534
7535 NextIRB.CreateMemCpy(DstShadowPtr, Align(1), SrcShadowPtr, Align(1), Size);
7536 if (MS.TrackOrigins) {
7537 Value *SrcOrigin = NextIRB.CreateAlignedLoad(MS.OriginTy, SrcOriginPtr,
7539 Value *NewOrigin = updateOrigin(SrcOrigin, NextIRB);
7540 NextIRB.CreateCall(MS.MsanSetOriginFn, {DstPtr, Size, NewOrigin});
7541 }
7542 }
7543
7544 void visitLibAtomicStore(CallBase &CB) {
7545 IRBuilder<> IRB(&CB);
7546 Value *Size = CB.getArgOperand(0);
7547 Value *DstPtr = CB.getArgOperand(2);
7548 Value *Ordering = CB.getArgOperand(3);
7549 // Convert the call to have at least Release ordering to make sure
7550 // the shadow operations aren't reordered after it.
7551 Value *NewOrdering =
7552 IRB.CreateExtractElement(makeAddReleaseOrderingTable(IRB), Ordering);
7553 CB.setArgOperand(3, NewOrdering);
7554
7555 Value *DstShadowPtr =
7556 getShadowOriginPtr(DstPtr, IRB, IRB.getInt8Ty(), Align(1),
7557 /*isStore*/ true)
7558 .first;
7559
7560 // Atomic store always paints clean shadow/origin. See file header.
7561 IRB.CreateMemSet(DstShadowPtr, getCleanShadow(IRB.getInt8Ty()), Size,
7562 Align(1));
7563 }
7564
7565 void visitCallBase(CallBase &CB) {
7566 assert(!CB.getMetadata(LLVMContext::MD_nosanitize));
7567 if (CB.isInlineAsm()) {
7568 // For inline asm (either a call to asm function, or callbr instruction),
7569 // do the usual thing: check argument shadow and mark all outputs as
7570 // clean. Note that any side effects of the inline asm that are not
7571 // immediately visible in its constraints are not handled.
7572 if (Opts.msan_handle_asm_conservative)
7573 visitAsmInstruction(CB);
7574 else
7575 visitInstruction(CB);
7576 return;
7577 }
7578 LibFunc LF = TLI->getLibFunc(CB);
7579 if (LF != NotLibFunc) {
7580 // libatomic.a functions need to have special handling because there isn't
7581 // a good way to intercept them or compile the library with
7582 // instrumentation.
7583 switch (LF) {
7584 case LibFunc_atomic_load:
7585 if (!isa<CallInst>(CB)) {
7586 llvm::errs() << "MSAN -- cannot instrument invoke of libatomic load."
7587 "Ignoring!\n";
7588 break;
7589 }
7590 visitLibAtomicLoad(CB);
7591 return;
7592 case LibFunc_atomic_store:
7593 visitLibAtomicStore(CB);
7594 return;
7595 default:
7596 break;
7597 }
7598 }
7599
7600 if (auto *Call = dyn_cast<CallInst>(&CB)) {
7601 assert(!isa<IntrinsicInst>(Call) && "intrinsics are handled elsewhere");
7602
7603 // We are going to insert code that relies on the fact that the callee
7604 // will become a non-readonly function after it is instrumented by us. To
7605 // prevent this code from being optimized out, mark that function
7606 // non-readonly in advance.
7607 // TODO: We can likely do better than dropping memory() completely here.
7608 AttributeMask B;
7609 B.addAttribute(Attribute::Memory).addAttribute(Attribute::Speculatable);
7610
7612 if (Function *Func = Call->getCalledFunction()) {
7613 Func->removeFnAttrs(B);
7614 }
7615
7617 }
7618 IRBuilder<> IRB(&CB);
7619 bool MayCheckCall = MS.EagerChecks;
7620 if (Function *Func = CB.getCalledFunction()) {
7621 // __sanitizer_unaligned_{load,store} functions may be called by users
7622 // and always expects shadows in the TLS. So don't check them.
7623 MayCheckCall &= !Func->getName().starts_with("__sanitizer_unaligned_");
7624 }
7625
7626 unsigned ArgOffset = 0;
7627 LLVM_DEBUG(dbgs() << " CallSite: " << CB << "\n");
7628 for (const auto &[i, A] : llvm::enumerate(CB.args())) {
7629 if (!A->getType()->isSized()) {
7630 LLVM_DEBUG(dbgs() << "Arg " << i << " is not sized: " << CB << "\n");
7631 continue;
7632 }
7633
7634 if (A->getType()->isScalableTy()) {
7635 LLVM_DEBUG(dbgs() << "Arg " << i << " is vscale: " << CB << "\n");
7636 // Handle as noundef, but don't reserve tls slots.
7637 insertCheckShadowOf(A, &CB);
7638 continue;
7639 }
7640
7641 unsigned Size = 0;
7642 const DataLayout &DL = F.getDataLayout();
7643
7644 bool ByVal = CB.isByValArgument(i);
7645 bool NoUndef = CB.paramHasAttr(i, Attribute::NoUndef);
7646 bool EagerCheck = MayCheckCall && !ByVal && NoUndef;
7647
7648 if (EagerCheck) {
7649 insertCheckShadowOf(A, &CB);
7650 Size = DL.getTypeAllocSize(A->getType());
7651 } else {
7652 [[maybe_unused]] Value *Store = nullptr;
7653 // Compute the Shadow for arg even if it is ByVal, because
7654 // in that case getShadow() will copy the actual arg shadow to
7655 // __msan_param_tls.
7656 Value *ArgShadow = getShadow(A);
7657 Value *ArgShadowBase = getShadowPtrForArgument(IRB, ArgOffset);
7658 LLVM_DEBUG(dbgs() << " Arg#" << i << ": " << *A
7659 << " Shadow: " << *ArgShadow << "\n");
7660 if (ByVal) {
7661 // ByVal requires some special handling as it's too big for a single
7662 // load
7663 assert(A->getType()->isPointerTy() &&
7664 "ByVal argument is not a pointer!");
7665 Size = DL.getTypeAllocSize(CB.getParamByValType(i));
7666 if (ArgOffset + Size > kParamTLSSize)
7667 break;
7668 const MaybeAlign ParamAlignment(CB.getParamAlign(i));
7669 MaybeAlign Alignment = std::nullopt;
7670 if (ParamAlignment)
7671 Alignment = std::min(*ParamAlignment, kShadowTLSAlignment);
7672 Value *AShadowPtr, *AOriginPtr;
7673 std::tie(AShadowPtr, AOriginPtr) =
7674 getShadowOriginPtr(A, IRB, IRB.getInt8Ty(), Alignment,
7675 /*isStore*/ false);
7676 if (!PropagateShadow) {
7677 Store = IRB.CreateMemSet(ArgShadowBase,
7679 Size, Alignment);
7680 } else {
7681 Store = IRB.CreateMemCpy(ArgShadowBase, Alignment, AShadowPtr,
7682 Alignment, Size);
7683 if (MS.TrackOrigins) {
7684 Value *ArgOriginBase = getOriginPtrForArgument(IRB, ArgOffset);
7685 // FIXME: OriginSize should be:
7686 // alignTo(A % kMinOriginAlignment + Size, kMinOriginAlignment)
7687 unsigned OriginSize = alignTo(Size, kMinOriginAlignment);
7688 IRB.CreateMemCpy(
7689 ArgOriginBase,
7690 /* by origin_tls[ArgOffset] */ kMinOriginAlignment,
7691 AOriginPtr,
7692 /* by getShadowOriginPtr */ kMinOriginAlignment, OriginSize);
7693 }
7694 }
7695 } else {
7696 // Any other parameters mean we need bit-grained tracking of uninit
7697 // data
7698 Size = DL.getTypeAllocSize(A->getType());
7699 if (ArgOffset + Size > kParamTLSSize)
7700 break;
7701 Store = IRB.CreateAlignedStore(ArgShadow, ArgShadowBase,
7703 Constant *Cst = dyn_cast<Constant>(ArgShadow);
7704 if (MS.TrackOrigins && !(Cst && Cst->isNullValue())) {
7705 IRB.CreateStore(getOrigin(A),
7706 getOriginPtrForArgument(IRB, ArgOffset));
7707 }
7708 }
7709 assert(Store != nullptr);
7710 LLVM_DEBUG(dbgs() << " Param:" << *Store << "\n");
7711 }
7712 assert(Size != 0);
7713 ArgOffset += alignTo(Size, kShadowTLSAlignment);
7714 }
7715 LLVM_DEBUG(dbgs() << " done with call args\n");
7716
7717 FunctionType *FT = CB.getFunctionType();
7718 if (FT->isVarArg()) {
7719 VAHelper->visitCallBase(CB, IRB);
7720 }
7721
7722 // Now, get the shadow for the RetVal.
7723 if (!CB.getType()->isSized())
7724 return;
7725 // Don't emit the epilogue for musttail call returns.
7726 if (isa<CallInst>(CB) && cast<CallInst>(CB).isMustTailCall())
7727 return;
7728
7729 if (MayCheckCall && CB.hasRetAttr(Attribute::NoUndef)) {
7730 setShadow(&CB, getCleanShadow(&CB));
7731 setOrigin(&CB, getCleanOrigin());
7732 return;
7733 }
7734
7735 IRBuilder<> IRBBefore(&CB);
7736 // Until we have full dynamic coverage, make sure the retval shadow is 0.
7737 Value *Base = getShadowPtrForRetval(IRBBefore);
7738 IRBBefore.CreateAlignedStore(getCleanShadow(&CB), Base,
7740 BasicBlock::iterator NextInsn;
7741 if (isa<CallInst>(CB)) {
7742 NextInsn = ++CB.getIterator();
7743 assert(NextInsn != CB.getParent()->end());
7744 } else {
7745 BasicBlock *NormalDest = cast<InvokeInst>(CB).getNormalDest();
7746 if (!NormalDest->getSinglePredecessor()) {
7747 // FIXME: this case is tricky, so we are just conservative here.
7748 // Perhaps we need to split the edge between this BB and NormalDest,
7749 // but a naive attempt to use SplitEdge leads to a crash.
7750 setShadow(&CB, getCleanShadow(&CB));
7751 setOrigin(&CB, getCleanOrigin());
7752 return;
7753 }
7754 // FIXME: NextInsn is likely in a basic block that has not been visited
7755 // yet. Anything inserted there will be instrumented by MSan later!
7756 NextInsn = NormalDest->getFirstInsertionPt();
7757 assert(NextInsn != NormalDest->end() &&
7758 "Could not find insertion point for retval shadow load");
7759 }
7760 IRBuilder<> IRBAfter(&*NextInsn);
7761 Value *RetvalShadow = IRBAfter.CreateAlignedLoad(
7762 getShadowTy(&CB), getShadowPtrForRetval(IRBAfter), kShadowTLSAlignment,
7763 "_msret");
7764 setShadow(&CB, RetvalShadow);
7765 if (MS.TrackOrigins)
7766 setOrigin(&CB, IRBAfter.CreateLoad(MS.OriginTy, getOriginPtrForRetval()));
7767 }
7768
7769 bool isAMustTailRetVal(Value *RetVal) {
7770 if (auto *I = dyn_cast<BitCastInst>(RetVal)) {
7771 RetVal = I->getOperand(0);
7772 }
7773 if (auto *I = dyn_cast<CallInst>(RetVal)) {
7774 return I->isMustTailCall();
7775 }
7776 return false;
7777 }
7778
7779 void visitReturnInst(ReturnInst &I) {
7780 IRBuilder<> IRB(&I);
7781 Value *RetVal = I.getReturnValue();
7782 if (!RetVal)
7783 return;
7784 // Don't emit the epilogue for musttail call returns.
7785 if (isAMustTailRetVal(RetVal))
7786 return;
7787 Value *ShadowPtr = getShadowPtrForRetval(IRB);
7788 bool HasNoUndef = F.hasRetAttribute(Attribute::NoUndef);
7789 bool StoreShadow = !(MS.EagerChecks && HasNoUndef);
7790 // FIXME: Consider using SpecialCaseList to specify a list of functions that
7791 // must always return fully initialized values. For now, we hardcode "main".
7792 bool EagerCheck = (MS.EagerChecks && HasNoUndef) || (F.getName() == "main");
7793
7794 Value *Shadow = getShadow(RetVal);
7795 bool StoreOrigin = true;
7796 if (EagerCheck) {
7797 insertCheckShadowOf(RetVal, &I);
7798 Shadow = getCleanShadow(RetVal);
7799 StoreOrigin = false;
7800 }
7801
7802 // The caller may still expect information passed over TLS if we pass our
7803 // check
7804 if (StoreShadow) {
7805 IRB.CreateAlignedStore(Shadow, ShadowPtr, kShadowTLSAlignment);
7806 if (MS.TrackOrigins && StoreOrigin)
7807 IRB.CreateStore(getOrigin(RetVal), getOriginPtrForRetval());
7808 }
7809 }
7810
7811 void visitPHINode(PHINode &I) {
7812 IRBuilder<> IRB(&I);
7813 if (!PropagateShadow) {
7814 setShadow(&I, getCleanShadow(&I));
7815 setOrigin(&I, getCleanOrigin());
7816 return;
7817 }
7818
7819 ShadowPHINodes.push_back(&I);
7820 setShadow(&I, IRB.CreatePHI(getShadowTy(&I), I.getNumIncomingValues(),
7821 "_msphi_s"));
7822 if (MS.TrackOrigins)
7823 setOrigin(
7824 &I, IRB.CreatePHI(MS.OriginTy, I.getNumIncomingValues(), "_msphi_o"));
7825 }
7826
7827 Value *getLocalVarIdptr(AllocaInst &I) {
7828 ConstantInt *IntConst =
7829 ConstantInt::get(Type::getInt32Ty((*F.getParent()).getContext()), 0);
7830 return new GlobalVariable(*F.getParent(), IntConst->getType(),
7831 /*isConstant=*/false, GlobalValue::PrivateLinkage,
7832 IntConst);
7833 }
7834
7835 Value *getLocalVarDescription(AllocaInst &I) {
7836 return createPrivateConstGlobalForString(*F.getParent(), I.getName());
7837 }
7838
7839 void poisonAllocaUserspace(AllocaInst &I, IRBuilder<> &IRB, Value *Len) {
7840 if (PoisonStack && Opts.msan_poison_stack_with_call) {
7841 IRB.CreateCall(MS.MsanPoisonStackFn, {&I, Len});
7842 } else {
7843 Value *ShadowBase, *OriginBase;
7844 std::tie(ShadowBase, OriginBase) = getShadowOriginPtr(
7845 &I, IRB, IRB.getInt8Ty(), Align(1), /*isStore*/ true);
7846
7847 Value *PoisonValue =
7848 IRB.getInt8(PoisonStack ? Opts.msan_poison_stack_pattern : 0);
7849 IRB.CreateMemSet(ShadowBase, PoisonValue, Len, I.getAlign());
7850 }
7851
7852 if (PoisonStack && MS.TrackOrigins) {
7853 Value *Idptr = getLocalVarIdptr(I);
7854 if (Opts.msan_print_stack_names) {
7855 Value *Descr = getLocalVarDescription(I);
7856 IRB.CreateCall(MS.MsanSetAllocaOriginWithDescriptionFn,
7857 {&I, Len, Idptr, Descr});
7858 } else {
7859 IRB.CreateCall(MS.MsanSetAllocaOriginNoDescriptionFn, {&I, Len, Idptr});
7860 }
7861 }
7862 }
7863
7864 void poisonAllocaKmsan(AllocaInst &I, IRBuilder<> &IRB, Value *Len) {
7865 Value *Descr = getLocalVarDescription(I);
7866 if (PoisonStack) {
7867 IRB.CreateCall(MS.MsanPoisonAllocaFn, {&I, Len, Descr});
7868 } else {
7869 IRB.CreateCall(MS.MsanUnpoisonAllocaFn, {&I, Len});
7870 }
7871 }
7872
7873 void instrumentAlloca(AllocaInst &I, Instruction *InsPoint = nullptr) {
7874 if (!InsPoint)
7875 InsPoint = &I;
7876 NextNodeIRBuilder IRB(InsPoint);
7877 Value *Len = IRB.CreateAllocationSize(MS.IntptrTy, &I);
7878
7879 if (MS.CompileKernel)
7880 poisonAllocaKmsan(I, IRB, Len);
7881 else
7882 poisonAllocaUserspace(I, IRB, Len);
7883 }
7884
7885 void visitAllocaInst(AllocaInst &I) {
7886 setShadow(&I, getCleanShadow(&I));
7887 setOrigin(&I, getCleanOrigin());
7888 // We'll get to this alloca later unless it's poisoned at the corresponding
7889 // llvm.lifetime.start.
7890 AllocaSet.insert(&I);
7891 }
7892
7893 void visitSelectInst(SelectInst &I) {
7894 // a = select b, c, d
7895 Value *B = I.getCondition();
7896 Value *C = I.getTrueValue();
7897 Value *D = I.getFalseValue();
7898
7899 handleSelectLikeInst(I, B, C, D);
7900 }
7901
7902 void handleSelectLikeInst(Instruction &I, Value *B, Value *C, Value *D) {
7903 IRBuilder<> IRB(&I);
7904
7905 Value *Sb = getShadow(B);
7906 Value *Sc = getShadow(C);
7907 Value *Sd = getShadow(D);
7908
7909 Value *Ob = MS.TrackOrigins ? getOrigin(B) : nullptr;
7910 Value *Oc = MS.TrackOrigins ? getOrigin(C) : nullptr;
7911 Value *Od = MS.TrackOrigins ? getOrigin(D) : nullptr;
7912
7913 // Result shadow if condition shadow is 0.
7914 Value *Sa0 = IRB.CreateSelect(B, Sc, Sd);
7915 Value *Sa1;
7916 if (I.getType()->isAggregateType()) {
7917 // To avoid "sign extending" i1 to an arbitrary aggregate type, we just do
7918 // an extra "select". This results in much more compact IR.
7919 // Sa = select Sb, poisoned, (select b, Sc, Sd)
7920 Sa1 = getPoisonedShadow(getShadowTy(I.getType()));
7921 } else if (isScalableNonVectorType(I.getType())) {
7922 // This is intended to handle target("aarch64.svcount"), which can't be
7923 // handled in the else branch because of incompatibility with CreateXor
7924 // ("The supported LLVM operations on this type are limited to load,
7925 // store, phi, select and alloca instructions").
7926
7927 // TODO: this currently underapproximates. Use Arm SVE EOR in the else
7928 // branch as needed instead.
7929 Sa1 = getCleanShadow(getShadowTy(I.getType()));
7930 } else {
7931 // Sa = select Sb, [ (c^d) | Sc | Sd ], [ b ? Sc : Sd ]
7932 // If Sb (condition is poisoned), look for bits in c and d that are equal
7933 // and both unpoisoned.
7934 // If !Sb (condition is unpoisoned), simply pick one of Sc and Sd.
7935
7936 // Cast arguments to shadow-compatible type.
7937 C = CreateAppToShadowCast(IRB, C);
7938 D = CreateAppToShadowCast(IRB, D);
7939
7940 // Result shadow if condition shadow is 1.
7941 Sa1 = IRB.CreateOr({IRB.CreateXor(C, D), Sc, Sd});
7942 }
7943 Value *Sa = IRB.CreateSelect(Sb, Sa1, Sa0, "_msprop_select");
7944 setShadow(&I, Sa);
7945 if (MS.TrackOrigins) {
7946 // Origins are always i32, so any vector conditions must be flattened.
7947 // FIXME: consider tracking vector origins for app vectors?
7948 if (B->getType()->isVectorTy()) {
7949 B = convertToBool(B, IRB);
7950 Sb = convertToBool(Sb, IRB);
7951 }
7952 // a = select b, c, d
7953 // Oa = Sb ? Ob : (b ? Oc : Od)
7954 setOrigin(&I, IRB.CreateSelect(Sb, Ob, IRB.CreateSelect(B, Oc, Od)));
7955 }
7956 }
7957
7958 void visitLandingPadInst(LandingPadInst &I) {
7959 // Do nothing.
7960 // See https://github.com/google/sanitizers/issues/504
7961 setShadow(&I, getCleanShadow(&I));
7962 setOrigin(&I, getCleanOrigin());
7963 }
7964
7965 void visitCatchSwitchInst(CatchSwitchInst &I) {
7966 setShadow(&I, getCleanShadow(&I));
7967 setOrigin(&I, getCleanOrigin());
7968 }
7969
7970 void visitFuncletPadInst(FuncletPadInst &I) {
7971 setShadow(&I, getCleanShadow(&I));
7972 setOrigin(&I, getCleanOrigin());
7973 }
7974
7975 void visitGetElementPtrInst(GetElementPtrInst &I) { handleShadowOr(I); }
7976
7977 void visitExtractValueInst(ExtractValueInst &I) {
7978 IRBuilder<> IRB(&I);
7979 Value *Agg = I.getAggregateOperand();
7980 LLVM_DEBUG(dbgs() << "ExtractValue: " << I << "\n");
7981 Value *AggShadow = getShadow(Agg);
7982 LLVM_DEBUG(dbgs() << " AggShadow: " << *AggShadow << "\n");
7983 Value *ResShadow = IRB.CreateExtractValue(AggShadow, I.getIndices());
7984 LLVM_DEBUG(dbgs() << " ResShadow: " << *ResShadow << "\n");
7985 setShadow(&I, ResShadow);
7986 setOriginForNaryOp(I);
7987 }
7988
7989 void visitInsertValueInst(InsertValueInst &I) {
7990 IRBuilder<> IRB(&I);
7991 LLVM_DEBUG(dbgs() << "InsertValue: " << I << "\n");
7992 Value *AggShadow = getShadow(I.getAggregateOperand());
7993 Value *InsShadow = getShadow(I.getInsertedValueOperand());
7994 LLVM_DEBUG(dbgs() << " AggShadow: " << *AggShadow << "\n");
7995 LLVM_DEBUG(dbgs() << " InsShadow: " << *InsShadow << "\n");
7996 Value *Res = IRB.CreateInsertValue(AggShadow, InsShadow, I.getIndices());
7997 LLVM_DEBUG(dbgs() << " Res: " << *Res << "\n");
7998 setShadow(&I, Res);
7999 setOriginForNaryOp(I);
8000 }
8001
8002 void dumpInst(Instruction &I, const Twine &Prefix) {
8003 // Instruction name only
8004 // For intrinsics, the full/overloaded name is used
8005 //
8006 // e.g., "call llvm.aarch64.neon.uqsub.v16i8"
8007 if (CallInst *CI = dyn_cast<CallInst>(&I)) {
8008 errs() << "ZZZ:" << Prefix << " call "
8009 << CI->getCalledFunction()->getName() << "\n";
8010 } else {
8011 errs() << "ZZZ:" << Prefix << " " << I.getOpcodeName() << "\n";
8012 }
8013
8014 // Instruction prototype (including return type and parameter types)
8015 // For intrinsics, we use the base/non-overloaded name
8016 //
8017 // e.g., "call <16 x i8> @llvm.aarch64.neon.uqsub(<16 x i8>, <16 x i8>)"
8018 unsigned NumOperands = I.getNumOperands();
8019 if (CallInst *CI = dyn_cast<CallInst>(&I)) {
8020 errs() << "YYY:" << Prefix << " call " << *I.getType() << " @";
8021
8022 if (IntrinsicInst *II = dyn_cast<IntrinsicInst>(CI))
8023 errs() << Intrinsic::getBaseName(II->getIntrinsicID());
8024 else
8025 errs() << CI->getCalledFunction()->getName();
8026
8027 errs() << "(";
8028
8029 // The last operand of a CallInst is the function itself.
8030 NumOperands--;
8031 } else
8032 errs() << "YYY:" << Prefix << " " << *I.getType() << " "
8033 << I.getOpcodeName() << "(";
8034
8035 for (size_t i = 0; i < NumOperands; i++) {
8036 if (i > 0)
8037 errs() << ", ";
8038
8039 errs() << *(I.getOperand(i)->getType());
8040 }
8041
8042 errs() << ")\n";
8043
8044 // Full instruction, including types and operand values
8045 // For intrinsics, the full/overloaded name is used
8046 //
8047 // e.g., "%vqsubq_v.i15 = call noundef <16 x i8>
8048 // @llvm.aarch64.neon.uqsub.v16i8(<16 x i8> %vext21.i,
8049 // <16 x i8> splat (i8 1)), !dbg !66"
8050 errs() << "QQQ:" << Prefix << " " << I << "\n";
8051 }
8052
8053 void visitResumeInst(ResumeInst &I) {
8054 LLVM_DEBUG(dbgs() << "Resume: " << I << "\n");
8055 // Nothing to do here.
8056 }
8057
8058 void visitCleanupReturnInst(CleanupReturnInst &CRI) {
8059 LLVM_DEBUG(dbgs() << "CleanupReturn: " << CRI << "\n");
8060 // Nothing to do here.
8061 }
8062
8063 void visitCatchReturnInst(CatchReturnInst &CRI) {
8064 LLVM_DEBUG(dbgs() << "CatchReturn: " << CRI << "\n");
8065 // Nothing to do here.
8066 }
8067
8068 void instrumentAsmArgument(Value *Operand, Type *ElemTy, Instruction &I,
8069 IRBuilder<> &IRB, const DataLayout &DL,
8070 bool isOutput) {
8071 // For each assembly argument, we check its value for being initialized.
8072 // If the argument is a pointer, we assume it points to a single element
8073 // of the corresponding type (or to a 8-byte word, if the type is unsized).
8074 // Each such pointer is instrumented with a call to the runtime library.
8075 Type *OpType = Operand->getType();
8076 // Check the operand value itself.
8077 insertCheckShadowOf(Operand, &I);
8078 if (!OpType->isPointerTy() || !isOutput) {
8079 assert(!isOutput);
8080 return;
8081 }
8082 if (!ElemTy->isSized())
8083 return;
8084 auto Size = DL.getTypeStoreSize(ElemTy);
8085 Value *SizeVal = IRB.CreateTypeSize(MS.IntptrTy, Size);
8086 if (MS.CompileKernel) {
8087 IRB.CreateCall(MS.MsanInstrumentAsmStoreFn, {Operand, SizeVal});
8088 } else {
8089 // ElemTy, derived from elementtype(), does not encode the alignment of
8090 // the pointer. Conservatively assume that the shadow memory is unaligned.
8091 // When Size is large, avoid StoreInst as it would expand to many
8092 // instructions.
8093 auto [ShadowPtr, _] =
8094 getShadowOriginPtrUserspace(Operand, IRB, IRB.getInt8Ty(), Align(1));
8095 if (Size <= 32)
8096 IRB.CreateAlignedStore(getCleanShadow(ElemTy), ShadowPtr, Align(1));
8097 else
8098 IRB.CreateMemSet(ShadowPtr, ConstantInt::getNullValue(IRB.getInt8Ty()),
8099 SizeVal, Align(1));
8100 }
8101 }
8102
8103 /// Get the number of output arguments returned by pointers.
8104 int getNumOutputArgs(InlineAsm *IA, CallBase *CB) {
8105 int NumRetOutputs = 0;
8106 int NumOutputs = 0;
8107 Type *RetTy = cast<Value>(CB)->getType();
8108 if (!RetTy->isVoidTy()) {
8109 // Register outputs are returned via the CallInst return value.
8110 auto *ST = dyn_cast<StructType>(RetTy);
8111 if (ST)
8112 NumRetOutputs = ST->getNumElements();
8113 else
8114 NumRetOutputs = 1;
8115 }
8116 InlineAsm::ConstraintInfoVector Constraints = IA->ParseConstraints();
8117 for (const InlineAsm::ConstraintInfo &Info : Constraints) {
8118 switch (Info.Type) {
8120 NumOutputs++;
8121 break;
8122 default:
8123 break;
8124 }
8125 }
8126 return NumOutputs - NumRetOutputs;
8127 }
8128
8129 void visitAsmInstruction(Instruction &I) {
8130 // Conservative inline assembly handling: check for poisoned shadow of
8131 // asm() arguments, then unpoison the result and all the memory locations
8132 // pointed to by those arguments.
8133 // An inline asm() statement in C++ contains lists of input and output
8134 // arguments used by the assembly code. These are mapped to operands of the
8135 // CallInst as follows:
8136 // - nR register outputs ("=r) are returned by value in a single structure
8137 // (SSA value of the CallInst);
8138 // - nO other outputs ("=m" and others) are returned by pointer as first
8139 // nO operands of the CallInst;
8140 // - nI inputs ("r", "m" and others) are passed to CallInst as the
8141 // remaining nI operands.
8142 // The total number of asm() arguments in the source is nR+nO+nI, and the
8143 // corresponding CallInst has nO+nI+1 operands (the last operand is the
8144 // function to be called).
8145 const DataLayout &DL = F.getDataLayout();
8146 CallBase *CB = cast<CallBase>(&I);
8147 IRBuilder<> IRB(&I);
8148 InlineAsm *IA = cast<InlineAsm>(CB->getCalledOperand());
8149 int OutputArgs = getNumOutputArgs(IA, CB);
8150 // The last operand of a CallInst is the function itself.
8151 int NumOperands = CB->getNumOperands() - 1;
8152
8153 // Check input arguments. Doing so before unpoisoning output arguments, so
8154 // that we won't overwrite uninit values before checking them.
8155 for (int i = OutputArgs; i < NumOperands; i++) {
8156 Value *Operand = CB->getOperand(i);
8157 instrumentAsmArgument(Operand, CB->getParamElementType(i), I, IRB, DL,
8158 /*isOutput*/ false);
8159 }
8160 // Unpoison output arguments. This must happen before the actual InlineAsm
8161 // call, so that the shadow for memory published in the asm() statement
8162 // remains valid.
8163 for (int i = 0; i < OutputArgs; i++) {
8164 Value *Operand = CB->getOperand(i);
8165 instrumentAsmArgument(Operand, CB->getParamElementType(i), I, IRB, DL,
8166 /*isOutput*/ true);
8167 }
8168
8169 setShadow(&I, getCleanShadow(&I));
8170 setOrigin(&I, getCleanOrigin());
8171 }
8172
8173 void visitFreezeInst(FreezeInst &I) {
8174 // Freeze always returns a fully defined value.
8175 setShadow(&I, getCleanShadow(&I));
8176 setOrigin(&I, getCleanOrigin());
8177 }
8178
8179 void visitInstruction(Instruction &I) {
8180 // Everything else: stop propagating and check for poisoned shadow.
8181 if (Opts.msan_dump_strict_instructions)
8182 dumpInst(I, "Strict");
8183 LLVM_DEBUG(dbgs() << "DEFAULT: " << I << "\n");
8184 for (size_t i = 0, n = I.getNumOperands(); i < n; i++) {
8185 Value *Operand = I.getOperand(i);
8186 if (Operand->getType()->isSized())
8187 insertCheckShadowOf(Operand, &I);
8188 }
8189 setShadow(&I, getCleanShadow(&I));
8190 setOrigin(&I, getCleanOrigin());
8191 }
8192};
8193
8194struct VarArgHelperBase : public VarArgHelper {
8195 Function &F;
8196 MemorySanitizer &MS;
8197 MemorySanitizerVisitor &MSV;
8198 SmallVector<CallInst *, 16> VAStartInstrumentationList;
8199 const unsigned VAListTagSize;
8200
8201 VarArgHelperBase(Function &F, MemorySanitizer &MS,
8202 MemorySanitizerVisitor &MSV, unsigned VAListTagSize)
8203 : F(F), MS(MS), MSV(MSV), VAListTagSize(VAListTagSize) {}
8204
8205 Value *getShadowAddrForVAArgument(IRBuilder<> &IRB, unsigned ArgOffset) {
8206 Value *Base = IRB.CreatePointerCast(MS.VAArgTLS, MS.IntptrTy);
8207 return IRB.CreateAdd(Base, ConstantInt::get(MS.IntptrTy, ArgOffset));
8208 }
8209
8210 /// Compute the shadow address for a given va_arg.
8211 Value *getShadowPtrForVAArgument(IRBuilder<> &IRB, unsigned ArgOffset) {
8212 return IRB.CreatePtrAdd(
8213 MS.VAArgTLS, ConstantInt::get(MS.IntptrTy, ArgOffset), "_msarg_va_s");
8214 }
8215
8216 /// Compute the shadow address for a given va_arg.
8217 Value *getShadowPtrForVAArgument(IRBuilder<> &IRB, unsigned ArgOffset,
8218 unsigned ArgSize) {
8219 // Make sure we don't overflow __msan_va_arg_tls.
8220 if (ArgOffset + ArgSize > kParamTLSSize)
8221 return nullptr;
8222 return getShadowPtrForVAArgument(IRB, ArgOffset);
8223 }
8224
8225 /// Compute the origin address for a given va_arg.
8226 Value *getOriginPtrForVAArgument(IRBuilder<> &IRB, int ArgOffset) {
8227 // getOriginPtrForVAArgument() is always called after
8228 // getShadowPtrForVAArgument(), so __msan_va_arg_origin_tls can never
8229 // overflow.
8230 return IRB.CreatePtrAdd(MS.VAArgOriginTLS,
8231 ConstantInt::get(MS.IntptrTy, ArgOffset),
8232 "_msarg_va_o");
8233 }
8234
8235 void CleanUnusedTLS(IRBuilder<> &IRB, Value *ShadowBase,
8236 unsigned BaseOffset) {
8237 // The tails of __msan_va_arg_tls is not large enough to fit full
8238 // value shadow, but it will be copied to backup anyway. Make it
8239 // clean.
8240 if (BaseOffset >= kParamTLSSize)
8241 return;
8242 Value *TailSize =
8243 ConstantInt::getSigned(IRB.getInt32Ty(), kParamTLSSize - BaseOffset);
8244 IRB.CreateMemSet(ShadowBase, ConstantInt::getNullValue(IRB.getInt8Ty()),
8245 TailSize, Align(8));
8246 }
8247
8248 void unpoisonVAListTagForInst(IntrinsicInst &I) {
8249 IRBuilder<> IRB(&I);
8250 Value *VAListTag = I.getArgOperand(0);
8251 const Align Alignment = Align(8);
8252 auto [ShadowPtr, OriginPtr] = MSV.getShadowOriginPtr(
8253 VAListTag, IRB, IRB.getInt8Ty(), Alignment, /*isStore*/ true);
8254 // Unpoison the whole __va_list_tag.
8255 IRB.CreateMemSet(ShadowPtr, Constant::getNullValue(IRB.getInt8Ty()),
8256 VAListTagSize, Alignment, false);
8257 }
8258
8259 void visitVAStartInst(VAStartInst &I) override {
8260 if (F.getCallingConv() == CallingConv::Win64)
8261 return;
8262 VAStartInstrumentationList.push_back(&I);
8263 unpoisonVAListTagForInst(I);
8264 }
8265
8266 void visitVACopyInst(VACopyInst &I) override {
8267 if (F.getCallingConv() == CallingConv::Win64)
8268 return;
8269 unpoisonVAListTagForInst(I);
8270 }
8271};
8272
8273/// AMD64-specific implementation of VarArgHelper.
8274struct VarArgAMD64Helper : public VarArgHelperBase {
8275 // An unfortunate workaround for asymmetric lowering of va_arg stuff.
8276 // See a comment in visitCallBase for more details.
8277 static const unsigned AMD64GpEndOffset = 48; // AMD64 ABI Draft 0.99.6 p3.5.7
8278 static const unsigned AMD64FpEndOffsetSSE = 176;
8279 // If SSE is disabled, fp_offset in va_list is zero.
8280 static const unsigned AMD64FpEndOffsetNoSSE = AMD64GpEndOffset;
8281
8282 unsigned AMD64FpEndOffset;
8283 AllocaInst *VAArgTLSCopy = nullptr;
8284 AllocaInst *VAArgTLSOriginCopy = nullptr;
8285 Value *VAArgOverflowSize = nullptr;
8286
8287 enum ArgKind { AK_GeneralPurpose, AK_FloatingPoint, AK_Memory };
8288
8289 VarArgAMD64Helper(Function &F, MemorySanitizer &MS,
8290 MemorySanitizerVisitor &MSV)
8291 : VarArgHelperBase(F, MS, MSV, /*VAListTagSize=*/24) {
8292 AMD64FpEndOffset = AMD64FpEndOffsetSSE;
8293 for (const auto &Attr : F.getAttributes().getFnAttrs()) {
8294 if (Attr.isStringAttribute() &&
8295 (Attr.getKindAsString() == "target-features")) {
8296 if (Attr.getValueAsString().contains("-sse"))
8297 AMD64FpEndOffset = AMD64FpEndOffsetNoSSE;
8298 break;
8299 }
8300 }
8301 }
8302
8303 ArgKind classifyArgument(Value *arg) {
8304 // A very rough approximation of X86_64 argument classification rules.
8305 Type *T = arg->getType();
8306 if (T->isX86_FP80Ty())
8307 return AK_Memory;
8308 if (T->isFPOrFPVectorTy())
8309 return AK_FloatingPoint;
8310 if (T->isIntegerTy() && T->getPrimitiveSizeInBits() <= 64)
8311 return AK_GeneralPurpose;
8312 if (T->isPointerTy())
8313 return AK_GeneralPurpose;
8314 return AK_Memory;
8315 }
8316
8317 // For VarArg functions, store the argument shadow in an ABI-specific format
8318 // that corresponds to va_list layout.
8319 // We do this because Clang lowers va_arg in the frontend, and this pass
8320 // only sees the low level code that deals with va_list internals.
8321 // A much easier alternative (provided that Clang emits va_arg instructions)
8322 // would have been to associate each live instance of va_list with a copy of
8323 // MSanParamTLS, and extract shadow on va_arg() call in the argument list
8324 // order.
8325 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
8326 unsigned GpOffset = 0;
8327 unsigned FpOffset = AMD64GpEndOffset;
8328 unsigned OverflowOffset = AMD64FpEndOffset;
8329 const DataLayout &DL = F.getDataLayout();
8330
8331 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
8332 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
8333 bool IsByVal = CB.isByValArgument(ArgNo);
8334 if (IsByVal) {
8335 // ByVal arguments always go to the overflow area.
8336 // Fixed arguments passed through the overflow area will be stepped
8337 // over by va_start, so don't count them towards the offset.
8338 if (IsFixed)
8339 continue;
8340 assert(A->getType()->isPointerTy());
8341 Type *RealTy = CB.getParamByValType(ArgNo);
8342 uint64_t ArgSize = DL.getTypeAllocSize(RealTy);
8343 uint64_t AlignedSize = alignTo(ArgSize, 8);
8344 unsigned BaseOffset = OverflowOffset;
8345 Value *ShadowBase = getShadowPtrForVAArgument(IRB, OverflowOffset);
8346 Value *OriginBase = nullptr;
8347 if (MS.TrackOrigins)
8348 OriginBase = getOriginPtrForVAArgument(IRB, OverflowOffset);
8349 OverflowOffset += AlignedSize;
8350
8351 if (OverflowOffset > kParamTLSSize) {
8352 CleanUnusedTLS(IRB, ShadowBase, BaseOffset);
8353 continue; // We have no space to copy shadow there.
8354 }
8355
8356 Value *ShadowPtr, *OriginPtr;
8357 std::tie(ShadowPtr, OriginPtr) =
8358 MSV.getShadowOriginPtr(A, IRB, IRB.getInt8Ty(), kShadowTLSAlignment,
8359 /*isStore*/ false);
8360 IRB.CreateMemCpy(ShadowBase, kShadowTLSAlignment, ShadowPtr,
8361 kShadowTLSAlignment, ArgSize);
8362 if (MS.TrackOrigins)
8363 IRB.CreateMemCpy(OriginBase, kShadowTLSAlignment, OriginPtr,
8364 kShadowTLSAlignment, ArgSize);
8365 } else {
8366 ArgKind AK = classifyArgument(A);
8367 if (AK == AK_GeneralPurpose && GpOffset >= AMD64GpEndOffset)
8368 AK = AK_Memory;
8369 if (AK == AK_FloatingPoint && FpOffset >= AMD64FpEndOffset)
8370 AK = AK_Memory;
8371 Value *ShadowBase, *OriginBase = nullptr;
8372 switch (AK) {
8373 case AK_GeneralPurpose:
8374 ShadowBase = getShadowPtrForVAArgument(IRB, GpOffset);
8375 if (MS.TrackOrigins)
8376 OriginBase = getOriginPtrForVAArgument(IRB, GpOffset);
8377 GpOffset += 8;
8378 assert(GpOffset <= kParamTLSSize);
8379 break;
8380 case AK_FloatingPoint:
8381 ShadowBase = getShadowPtrForVAArgument(IRB, FpOffset);
8382 if (MS.TrackOrigins)
8383 OriginBase = getOriginPtrForVAArgument(IRB, FpOffset);
8384 FpOffset += 16;
8385 assert(FpOffset <= kParamTLSSize);
8386 break;
8387 case AK_Memory:
8388 if (IsFixed)
8389 continue;
8390 uint64_t ArgSize = DL.getTypeAllocSize(A->getType());
8391 uint64_t AlignedSize = alignTo(ArgSize, 8);
8392 unsigned BaseOffset = OverflowOffset;
8393 ShadowBase = getShadowPtrForVAArgument(IRB, OverflowOffset);
8394 if (MS.TrackOrigins) {
8395 OriginBase = getOriginPtrForVAArgument(IRB, OverflowOffset);
8396 }
8397 OverflowOffset += AlignedSize;
8398 if (OverflowOffset > kParamTLSSize) {
8399 // We have no space to copy shadow there.
8400 CleanUnusedTLS(IRB, ShadowBase, BaseOffset);
8401 continue;
8402 }
8403 }
8404 // Take fixed arguments into account for GpOffset and FpOffset,
8405 // but don't actually store shadows for them.
8406 // TODO(glider): don't call get*PtrForVAArgument() for them.
8407 if (IsFixed)
8408 continue;
8409 Value *Shadow = MSV.getShadow(A);
8410 IRB.CreateAlignedStore(Shadow, ShadowBase, kShadowTLSAlignment);
8411 if (MS.TrackOrigins) {
8412 Value *Origin = MSV.getOrigin(A);
8413 TypeSize StoreSize = DL.getTypeStoreSize(Shadow->getType());
8414 MSV.paintOrigin(IRB, Origin, OriginBase, StoreSize,
8416 }
8417 }
8418 }
8419 Constant *OverflowSize =
8420 ConstantInt::get(IRB.getInt64Ty(), OverflowOffset - AMD64FpEndOffset);
8421 IRB.CreateStore(OverflowSize, MS.VAArgOverflowSizeTLS);
8422 }
8423
8424 void finalizeInstrumentation() override {
8425 assert(!VAArgOverflowSize && !VAArgTLSCopy &&
8426 "finalizeInstrumentation called twice");
8427 if (!VAStartInstrumentationList.empty()) {
8428 // If there is a va_start in this function, make a backup copy of
8429 // va_arg_tls somewhere in the function entry block.
8430 IRBuilder<> IRB(MSV.FnPrologueEnd);
8431 VAArgOverflowSize =
8432 IRB.CreateLoad(IRB.getInt64Ty(), MS.VAArgOverflowSizeTLS);
8433 Value *CopySize = IRB.CreateAdd(
8434 ConstantInt::get(MS.IntptrTy, AMD64FpEndOffset), VAArgOverflowSize);
8435 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
8436 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
8437 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
8438 CopySize, kShadowTLSAlignment, false);
8439
8440 Value *SrcSize = IRB.CreateBinaryIntrinsic(
8441 Intrinsic::umin, CopySize,
8442 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
8443 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
8444 kShadowTLSAlignment, SrcSize);
8445 if (MS.TrackOrigins) {
8446 VAArgTLSOriginCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
8447 VAArgTLSOriginCopy->setAlignment(kShadowTLSAlignment);
8448 IRB.CreateMemCpy(VAArgTLSOriginCopy, kShadowTLSAlignment,
8449 MS.VAArgOriginTLS, kShadowTLSAlignment, SrcSize);
8450 }
8451 }
8452
8453 // Instrument va_start.
8454 // Copy va_list shadow from the backup copy of the TLS contents.
8455 for (CallInst *OrigInst : VAStartInstrumentationList) {
8456 NextNodeIRBuilder IRB(OrigInst);
8457 Value *VAListTag = OrigInst->getArgOperand(0);
8458
8459 Value *RegSaveAreaPtrPtr =
8460 IRB.CreatePtrAdd(VAListTag, ConstantInt::get(MS.IntptrTy, 16));
8461 Value *RegSaveAreaPtr = IRB.CreateLoad(MS.PtrTy, RegSaveAreaPtrPtr);
8462 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
8463 const Align Alignment = Align(16);
8464 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
8465 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(),
8466 Alignment, /*isStore*/ true);
8467 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy, Alignment,
8468 AMD64FpEndOffset);
8469 if (MS.TrackOrigins)
8470 IRB.CreateMemCpy(RegSaveAreaOriginPtr, Alignment, VAArgTLSOriginCopy,
8471 Alignment, AMD64FpEndOffset);
8472 Value *OverflowArgAreaPtrPtr =
8473 IRB.CreatePtrAdd(VAListTag, ConstantInt::get(MS.IntptrTy, 8));
8474 Value *OverflowArgAreaPtr =
8475 IRB.CreateLoad(MS.PtrTy, OverflowArgAreaPtrPtr);
8476 Value *OverflowArgAreaShadowPtr, *OverflowArgAreaOriginPtr;
8477 std::tie(OverflowArgAreaShadowPtr, OverflowArgAreaOriginPtr) =
8478 MSV.getShadowOriginPtr(OverflowArgAreaPtr, IRB, IRB.getInt8Ty(),
8479 Alignment, /*isStore*/ true);
8480 Value *SrcPtr = IRB.CreateConstGEP1_32(IRB.getInt8Ty(), VAArgTLSCopy,
8481 AMD64FpEndOffset);
8482 IRB.CreateMemCpy(OverflowArgAreaShadowPtr, Alignment, SrcPtr, Alignment,
8483 VAArgOverflowSize);
8484 if (MS.TrackOrigins) {
8485 SrcPtr = IRB.CreateConstGEP1_32(IRB.getInt8Ty(), VAArgTLSOriginCopy,
8486 AMD64FpEndOffset);
8487 IRB.CreateMemCpy(OverflowArgAreaOriginPtr, Alignment, SrcPtr, Alignment,
8488 VAArgOverflowSize);
8489 }
8490 }
8491 }
8492};
8493
8494/// AArch64-specific implementation of VarArgHelper.
8495struct VarArgAArch64Helper : public VarArgHelperBase {
8496 static const unsigned kAArch64GrArgSize = 64;
8497 static const unsigned kAArch64VrArgSize = 128;
8498
8499 static const unsigned AArch64GrBegOffset = 0;
8500 static const unsigned AArch64GrEndOffset = kAArch64GrArgSize;
8501 // Make VR space aligned to 16 bytes.
8502 static const unsigned AArch64VrBegOffset = AArch64GrEndOffset;
8503 static const unsigned AArch64VrEndOffset =
8504 AArch64VrBegOffset + kAArch64VrArgSize;
8505 static const unsigned AArch64VAEndOffset = AArch64VrEndOffset;
8506
8507 AllocaInst *VAArgTLSCopy = nullptr;
8508 Value *VAArgOverflowSize = nullptr;
8509
8510 enum ArgKind { AK_GeneralPurpose, AK_FloatingPoint, AK_Memory };
8511
8512 VarArgAArch64Helper(Function &F, MemorySanitizer &MS,
8513 MemorySanitizerVisitor &MSV)
8514 : VarArgHelperBase(F, MS, MSV, /*VAListTagSize=*/32) {}
8515
8516 // A very rough approximation of aarch64 argument classification rules.
8517 std::pair<ArgKind, uint64_t> classifyArgument(Type *T) {
8518 if (T->isIntOrPtrTy() && T->getPrimitiveSizeInBits() <= 64)
8519 return {AK_GeneralPurpose, 1};
8520 if (T->isFloatingPointTy() && T->getPrimitiveSizeInBits() <= 128)
8521 return {AK_FloatingPoint, 1};
8522
8523 if (T->isArrayTy()) {
8524 auto R = classifyArgument(T->getArrayElementType());
8525 R.second *= T->getScalarType()->getArrayNumElements();
8526 return R;
8527 }
8528
8529 if (const FixedVectorType *FV = dyn_cast<FixedVectorType>(T)) {
8530 auto R = classifyArgument(FV->getScalarType());
8531 R.second *= FV->getNumElements();
8532 return R;
8533 }
8534
8535 LLVM_DEBUG(errs() << "Unknown vararg type: " << *T << "\n");
8536 return {AK_Memory, 0};
8537 }
8538
8539 // The instrumentation stores the argument shadow in a non ABI-specific
8540 // format because it does not know which argument is named (since Clang,
8541 // like x86_64 case, lowers the va_args in the frontend and this pass only
8542 // sees the low level code that deals with va_list internals).
8543 // The first seven GR registers are saved in the first 56 bytes of the
8544 // va_arg tls arra, followed by the first 8 FP/SIMD registers, and then
8545 // the remaining arguments.
8546 // Using constant offset within the va_arg TLS array allows fast copy
8547 // in the finalize instrumentation.
8548 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
8549 unsigned GrOffset = AArch64GrBegOffset;
8550 unsigned VrOffset = AArch64VrBegOffset;
8551 unsigned OverflowOffset = AArch64VAEndOffset;
8552
8553 const DataLayout &DL = F.getDataLayout();
8554 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
8555 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
8556 auto [AK, RegNum] = classifyArgument(A->getType());
8557 if (AK == AK_GeneralPurpose &&
8558 (GrOffset + RegNum * 8) > AArch64GrEndOffset)
8559 AK = AK_Memory;
8560 if (AK == AK_FloatingPoint &&
8561 (VrOffset + RegNum * 16) > AArch64VrEndOffset)
8562 AK = AK_Memory;
8563 Value *Base;
8564 switch (AK) {
8565 case AK_GeneralPurpose:
8566 Base = getShadowPtrForVAArgument(IRB, GrOffset);
8567 GrOffset += 8 * RegNum;
8568 break;
8569 case AK_FloatingPoint:
8570 Base = getShadowPtrForVAArgument(IRB, VrOffset);
8571 VrOffset += 16 * RegNum;
8572 break;
8573 case AK_Memory:
8574 // Don't count fixed arguments in the overflow area - va_start will
8575 // skip right over them.
8576 if (IsFixed)
8577 continue;
8578 uint64_t ArgSize = DL.getTypeAllocSize(A->getType());
8579 uint64_t AlignedSize = alignTo(ArgSize, 8);
8580 unsigned BaseOffset = OverflowOffset;
8581 Base = getShadowPtrForVAArgument(IRB, BaseOffset);
8582 OverflowOffset += AlignedSize;
8583 if (OverflowOffset > kParamTLSSize) {
8584 // We have no space to copy shadow there.
8585 CleanUnusedTLS(IRB, Base, BaseOffset);
8586 continue;
8587 }
8588 break;
8589 }
8590 // Count Gp/Vr fixed arguments to their respective offsets, but don't
8591 // bother to actually store a shadow.
8592 if (IsFixed)
8593 continue;
8594 IRB.CreateAlignedStore(MSV.getShadow(A), Base, kShadowTLSAlignment);
8595 }
8596 Constant *OverflowSize =
8597 ConstantInt::get(IRB.getInt64Ty(), OverflowOffset - AArch64VAEndOffset);
8598 IRB.CreateStore(OverflowSize, MS.VAArgOverflowSizeTLS);
8599 }
8600
8601 // Retrieve a va_list field of 'void*' size.
8602 Value *getVAField64(IRBuilder<> &IRB, Value *VAListTag, int offset) {
8603 Value *SaveAreaPtrPtr =
8604 IRB.CreatePtrAdd(VAListTag, ConstantInt::get(MS.IntptrTy, offset));
8605 return IRB.CreateLoad(Type::getInt64Ty(*MS.C), SaveAreaPtrPtr);
8606 }
8607
8608 // Retrieve a va_list field of 'int' size.
8609 Value *getVAField32(IRBuilder<> &IRB, Value *VAListTag, int offset) {
8610 Value *SaveAreaPtr =
8611 IRB.CreatePtrAdd(VAListTag, ConstantInt::get(MS.IntptrTy, offset));
8612 Value *SaveArea32 = IRB.CreateLoad(IRB.getInt32Ty(), SaveAreaPtr);
8613 return IRB.CreateSExt(SaveArea32, MS.IntptrTy);
8614 }
8615
8616 void finalizeInstrumentation() override {
8617 assert(!VAArgOverflowSize && !VAArgTLSCopy &&
8618 "finalizeInstrumentation called twice");
8619 if (!VAStartInstrumentationList.empty()) {
8620 // If there is a va_start in this function, make a backup copy of
8621 // va_arg_tls somewhere in the function entry block.
8622 IRBuilder<> IRB(MSV.FnPrologueEnd);
8623 VAArgOverflowSize =
8624 IRB.CreateLoad(IRB.getInt64Ty(), MS.VAArgOverflowSizeTLS);
8625 Value *CopySize = IRB.CreateAdd(
8626 ConstantInt::get(MS.IntptrTy, AArch64VAEndOffset), VAArgOverflowSize);
8627 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
8628 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
8629 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
8630 CopySize, kShadowTLSAlignment, false);
8631
8632 Value *SrcSize = IRB.CreateBinaryIntrinsic(
8633 Intrinsic::umin, CopySize,
8634 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
8635 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
8636 kShadowTLSAlignment, SrcSize);
8637 }
8638
8639 Value *GrArgSize = ConstantInt::get(MS.IntptrTy, kAArch64GrArgSize);
8640 Value *VrArgSize = ConstantInt::get(MS.IntptrTy, kAArch64VrArgSize);
8641
8642 // Instrument va_start, copy va_list shadow from the backup copy of
8643 // the TLS contents.
8644 for (CallInst *OrigInst : VAStartInstrumentationList) {
8645 NextNodeIRBuilder IRB(OrigInst);
8646
8647 Value *VAListTag = OrigInst->getArgOperand(0);
8648
8649 // The variadic ABI for AArch64 creates two areas to save the incoming
8650 // argument registers (one for 64-bit general register xn-x7 and another
8651 // for 128-bit FP/SIMD vn-v7).
8652 // We need then to propagate the shadow arguments on both regions
8653 // 'va::__gr_top + va::__gr_offs' and 'va::__vr_top + va::__vr_offs'.
8654 // The remaining arguments are saved on shadow for 'va::stack'.
8655 // One caveat is it requires only to propagate the non-named arguments,
8656 // however on the call site instrumentation 'all' the arguments are
8657 // saved. So to copy the shadow values from the va_arg TLS array
8658 // we need to adjust the offset for both GR and VR fields based on
8659 // the __{gr,vr}_offs value (since they are stores based on incoming
8660 // named arguments).
8661 Type *RegSaveAreaPtrTy = IRB.getPtrTy();
8662
8663 // Read the stack pointer from the va_list.
8664 Value *StackSaveAreaPtr =
8665 IRB.CreateIntToPtr(getVAField64(IRB, VAListTag, 0), RegSaveAreaPtrTy);
8666
8667 // Read both the __gr_top and __gr_off and add them up.
8668 Value *GrTopSaveAreaPtr = getVAField64(IRB, VAListTag, 8);
8669 Value *GrOffSaveArea = getVAField32(IRB, VAListTag, 24);
8670
8671 Value *GrRegSaveAreaPtr = IRB.CreateIntToPtr(
8672 IRB.CreateAdd(GrTopSaveAreaPtr, GrOffSaveArea), RegSaveAreaPtrTy);
8673
8674 // Read both the __vr_top and __vr_off and add them up.
8675 Value *VrTopSaveAreaPtr = getVAField64(IRB, VAListTag, 16);
8676 Value *VrOffSaveArea = getVAField32(IRB, VAListTag, 28);
8677
8678 Value *VrRegSaveAreaPtr = IRB.CreateIntToPtr(
8679 IRB.CreateAdd(VrTopSaveAreaPtr, VrOffSaveArea), RegSaveAreaPtrTy);
8680
8681 // It does not know how many named arguments is being used and, on the
8682 // callsite all the arguments were saved. Since __gr_off is defined as
8683 // '0 - ((8 - named_gr) * 8)', the idea is to just propagate the variadic
8684 // argument by ignoring the bytes of shadow from named arguments.
8685 Value *GrRegSaveAreaShadowPtrOff =
8686 IRB.CreateAdd(GrArgSize, GrOffSaveArea);
8687
8688 Value *GrRegSaveAreaShadowPtr =
8689 MSV.getShadowOriginPtr(GrRegSaveAreaPtr, IRB, IRB.getInt8Ty(),
8690 Align(8), /*isStore*/ true)
8691 .first;
8692
8693 Value *GrSrcPtr =
8694 IRB.CreateInBoundsPtrAdd(VAArgTLSCopy, GrRegSaveAreaShadowPtrOff);
8695 Value *GrCopySize = IRB.CreateSub(GrArgSize, GrRegSaveAreaShadowPtrOff);
8696
8697 IRB.CreateMemCpy(GrRegSaveAreaShadowPtr, Align(8), GrSrcPtr, Align(8),
8698 GrCopySize);
8699
8700 // Again, but for FP/SIMD values.
8701 Value *VrRegSaveAreaShadowPtrOff =
8702 IRB.CreateAdd(VrArgSize, VrOffSaveArea);
8703
8704 Value *VrRegSaveAreaShadowPtr =
8705 MSV.getShadowOriginPtr(VrRegSaveAreaPtr, IRB, IRB.getInt8Ty(),
8706 Align(8), /*isStore*/ true)
8707 .first;
8708
8709 Value *VrSrcPtr = IRB.CreateInBoundsPtrAdd(
8710 IRB.CreateInBoundsPtrAdd(VAArgTLSCopy,
8711 IRB.getInt32(AArch64VrBegOffset)),
8712 VrRegSaveAreaShadowPtrOff);
8713 Value *VrCopySize = IRB.CreateSub(VrArgSize, VrRegSaveAreaShadowPtrOff);
8714
8715 IRB.CreateMemCpy(VrRegSaveAreaShadowPtr, Align(8), VrSrcPtr, Align(8),
8716 VrCopySize);
8717
8718 // And finally for remaining arguments.
8719 Value *StackSaveAreaShadowPtr =
8720 MSV.getShadowOriginPtr(StackSaveAreaPtr, IRB, IRB.getInt8Ty(),
8721 Align(16), /*isStore*/ true)
8722 .first;
8723
8724 Value *StackSrcPtr = IRB.CreateInBoundsPtrAdd(
8725 VAArgTLSCopy, IRB.getInt32(AArch64VAEndOffset));
8726
8727 IRB.CreateMemCpy(StackSaveAreaShadowPtr, Align(16), StackSrcPtr,
8728 Align(16), VAArgOverflowSize);
8729 }
8730 }
8731};
8732
8733/// PowerPC64-specific implementation of VarArgHelper.
8734struct VarArgPowerPC64Helper : public VarArgHelperBase {
8735 AllocaInst *VAArgTLSCopy = nullptr;
8736 Value *VAArgSize = nullptr;
8737
8738 VarArgPowerPC64Helper(Function &F, MemorySanitizer &MS,
8739 MemorySanitizerVisitor &MSV)
8740 : VarArgHelperBase(F, MS, MSV, /*VAListTagSize=*/8) {}
8741
8742 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
8743 // For PowerPC, we need to deal with alignment of stack arguments -
8744 // they are mostly aligned to 8 bytes, but vectors and i128 arrays
8745 // are aligned to 16 bytes, byvals can be aligned to 8 or 16 bytes,
8746 // For that reason, we compute current offset from stack pointer (which is
8747 // always properly aligned), and offset for the first vararg, then subtract
8748 // them.
8749 unsigned VAArgBase;
8750 Triple TargetTriple(F.getParent()->getTargetTriple());
8751 // Parameter save area starts at 48 bytes from frame pointer for ABIv1,
8752 // and 32 bytes for ABIv2. This is usually determined by target
8753 // endianness, but in theory could be overridden by function attribute.
8754 if (TargetTriple.isPPC64ELFv2ABI())
8755 VAArgBase = 32;
8756 else
8757 VAArgBase = 48;
8758 unsigned VAArgOffset = VAArgBase;
8759 const DataLayout &DL = F.getDataLayout();
8760 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
8761 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
8762 bool IsByVal = CB.isByValArgument(ArgNo);
8763 if (IsByVal) {
8764 assert(A->getType()->isPointerTy());
8765 Type *RealTy = CB.getParamByValType(ArgNo);
8766 uint64_t ArgSize = DL.getTypeAllocSize(RealTy);
8767 Align ArgAlign = CB.getParamAlign(ArgNo).value_or(Align(8));
8768 if (ArgAlign < 8)
8769 ArgAlign = Align(8);
8770 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
8771 if (!IsFixed) {
8772 Value *Base =
8773 getShadowPtrForVAArgument(IRB, VAArgOffset - VAArgBase, ArgSize);
8774 if (Base) {
8775 Value *AShadowPtr, *AOriginPtr;
8776 std::tie(AShadowPtr, AOriginPtr) =
8777 MSV.getShadowOriginPtr(A, IRB, IRB.getInt8Ty(),
8778 kShadowTLSAlignment, /*isStore*/ false);
8779
8780 IRB.CreateMemCpy(Base, kShadowTLSAlignment, AShadowPtr,
8781 kShadowTLSAlignment, ArgSize);
8782 }
8783 }
8784 VAArgOffset += alignTo(ArgSize, Align(8));
8785 } else {
8786 Value *Base;
8787 uint64_t ArgSize = DL.getTypeAllocSize(A->getType());
8788 Align ArgAlign = Align(8);
8789 if (A->getType()->isArrayTy()) {
8790 // Arrays are aligned to element size, except for long double
8791 // arrays, which are aligned to 8 bytes.
8792 Type *ElementTy = A->getType()->getArrayElementType();
8793 if (!ElementTy->isPPC_FP128Ty())
8794 ArgAlign = Align(DL.getTypeAllocSize(ElementTy));
8795 } else if (A->getType()->isVectorTy()) {
8796 // Vectors are naturally aligned.
8797 ArgAlign = Align(ArgSize);
8798 }
8799 if (ArgAlign < 8)
8800 ArgAlign = Align(8);
8801 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
8802 if (DL.isBigEndian()) {
8803 // Adjusting the shadow for argument with size < 8 to match the
8804 // placement of bits in big endian system
8805 if (ArgSize < 8)
8806 VAArgOffset += (8 - ArgSize);
8807 }
8808 if (!IsFixed) {
8809 Base =
8810 getShadowPtrForVAArgument(IRB, VAArgOffset - VAArgBase, ArgSize);
8811 if (Base)
8812 IRB.CreateAlignedStore(MSV.getShadow(A), Base, kShadowTLSAlignment);
8813 }
8814 VAArgOffset += ArgSize;
8815 VAArgOffset = alignTo(VAArgOffset, Align(8));
8816 }
8817 if (IsFixed)
8818 VAArgBase = VAArgOffset;
8819 }
8820
8821 Constant *TotalVAArgSize =
8822 ConstantInt::get(MS.IntptrTy, VAArgOffset - VAArgBase);
8823 // Here using VAArgOverflowSizeTLS as VAArgSizeTLS to avoid creation of
8824 // a new class member i.e. it is the total size of all VarArgs.
8825 IRB.CreateStore(TotalVAArgSize, MS.VAArgOverflowSizeTLS);
8826 }
8827
8828 void finalizeInstrumentation() override {
8829 assert(!VAArgSize && !VAArgTLSCopy &&
8830 "finalizeInstrumentation called twice");
8831 IRBuilder<> IRB(MSV.FnPrologueEnd);
8832 VAArgSize = IRB.CreateLoad(IRB.getInt64Ty(), MS.VAArgOverflowSizeTLS);
8833 Value *CopySize = VAArgSize;
8834
8835 if (!VAStartInstrumentationList.empty()) {
8836 // If there is a va_start in this function, make a backup copy of
8837 // va_arg_tls somewhere in the function entry block.
8838
8839 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
8840 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
8841 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
8842 CopySize, kShadowTLSAlignment, false);
8843
8844 Value *SrcSize = IRB.CreateBinaryIntrinsic(
8845 Intrinsic::umin, CopySize,
8846 ConstantInt::get(IRB.getInt64Ty(), kParamTLSSize));
8847 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
8848 kShadowTLSAlignment, SrcSize);
8849 }
8850
8851 // Instrument va_start.
8852 // Copy va_list shadow from the backup copy of the TLS contents.
8853 for (CallInst *OrigInst : VAStartInstrumentationList) {
8854 NextNodeIRBuilder IRB(OrigInst);
8855 Value *VAListTag = OrigInst->getArgOperand(0);
8856 Value *RegSaveAreaPtrPtr = IRB.CreatePtrToInt(VAListTag, MS.IntptrTy);
8857
8858 RegSaveAreaPtrPtr = IRB.CreateIntToPtr(RegSaveAreaPtrPtr, MS.PtrTy);
8859
8860 Value *RegSaveAreaPtr = IRB.CreateLoad(MS.PtrTy, RegSaveAreaPtrPtr);
8861 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
8862 const DataLayout &DL = F.getDataLayout();
8863 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
8864 const Align Alignment = Align(IntptrSize);
8865 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
8866 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(),
8867 Alignment, /*isStore*/ true);
8868 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy, Alignment,
8869 CopySize);
8870 }
8871 }
8872};
8873
8874/// PowerPC32-specific implementation of VarArgHelper.
8875struct VarArgPowerPC32Helper : public VarArgHelperBase {
8876 AllocaInst *VAArgTLSCopy = nullptr;
8877 Value *VAArgSize = nullptr;
8878
8879 VarArgPowerPC32Helper(Function &F, MemorySanitizer &MS,
8880 MemorySanitizerVisitor &MSV)
8881 : VarArgHelperBase(F, MS, MSV, /*VAListTagSize=*/12) {}
8882
8883 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
8884 unsigned VAArgBase;
8885 // Parameter save area is 8 bytes from frame pointer in PPC32
8886 VAArgBase = 8;
8887 unsigned VAArgOffset = VAArgBase;
8888 const DataLayout &DL = F.getDataLayout();
8889 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
8890 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
8891 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
8892 bool IsByVal = CB.isByValArgument(ArgNo);
8893 if (IsByVal) {
8894 assert(A->getType()->isPointerTy());
8895 Type *RealTy = CB.getParamByValType(ArgNo);
8896 uint64_t ArgSize = DL.getTypeAllocSize(RealTy);
8897 Align ArgAlign = CB.getParamAlign(ArgNo).value_or(Align(IntptrSize));
8898 if (ArgAlign < IntptrSize)
8899 ArgAlign = Align(IntptrSize);
8900 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
8901 if (!IsFixed) {
8902 Value *Base =
8903 getShadowPtrForVAArgument(IRB, VAArgOffset - VAArgBase, ArgSize);
8904 if (Base) {
8905 Value *AShadowPtr, *AOriginPtr;
8906 std::tie(AShadowPtr, AOriginPtr) =
8907 MSV.getShadowOriginPtr(A, IRB, IRB.getInt8Ty(),
8908 kShadowTLSAlignment, /*isStore*/ false);
8909
8910 IRB.CreateMemCpy(Base, kShadowTLSAlignment, AShadowPtr,
8911 kShadowTLSAlignment, ArgSize);
8912 }
8913 }
8914 VAArgOffset += alignTo(ArgSize, Align(IntptrSize));
8915 } else {
8916 Value *Base;
8917 Type *ArgTy = A->getType();
8918
8919 // On PPC 32 floating point variable arguments are stored in separate
8920 // area: fp_save_area = reg_save_area + 4*8. We do not copy shaodow for
8921 // them as they will be found when checking call arguments.
8922 if (!ArgTy->isFloatingPointTy()) {
8923 uint64_t ArgSize = DL.getTypeAllocSize(ArgTy);
8924 Align ArgAlign = Align(IntptrSize);
8925 if (ArgTy->isArrayTy()) {
8926 // Arrays are aligned to element size, except for long double
8927 // arrays, which are aligned to 8 bytes.
8928 Type *ElementTy = ArgTy->getArrayElementType();
8929 if (!ElementTy->isPPC_FP128Ty())
8930 ArgAlign = Align(DL.getTypeAllocSize(ElementTy));
8931 } else if (ArgTy->isVectorTy()) {
8932 // Vectors are naturally aligned.
8933 ArgAlign = Align(ArgSize);
8934 }
8935 if (ArgAlign < IntptrSize)
8936 ArgAlign = Align(IntptrSize);
8937 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
8938 if (DL.isBigEndian()) {
8939 // Adjusting the shadow for argument with size < IntptrSize to match
8940 // the placement of bits in big endian system
8941 if (ArgSize < IntptrSize)
8942 VAArgOffset += (IntptrSize - ArgSize);
8943 }
8944 if (!IsFixed) {
8945 Base = getShadowPtrForVAArgument(IRB, VAArgOffset - VAArgBase,
8946 ArgSize);
8947 if (Base)
8948 IRB.CreateAlignedStore(MSV.getShadow(A), Base,
8950 }
8951 VAArgOffset += ArgSize;
8952 VAArgOffset = alignTo(VAArgOffset, Align(IntptrSize));
8953 }
8954 }
8955 }
8956
8957 Constant *TotalVAArgSize =
8958 ConstantInt::get(MS.IntptrTy, VAArgOffset - VAArgBase);
8959 // Here using VAArgOverflowSizeTLS as VAArgSizeTLS to avoid creation of
8960 // a new class member i.e. it is the total size of all VarArgs.
8961 IRB.CreateStore(TotalVAArgSize, MS.VAArgOverflowSizeTLS);
8962 }
8963
8964 void finalizeInstrumentation() override {
8965 assert(!VAArgSize && !VAArgTLSCopy &&
8966 "finalizeInstrumentation called twice");
8967 IRBuilder<> IRB(MSV.FnPrologueEnd);
8968 VAArgSize = IRB.CreateLoad(MS.IntptrTy, MS.VAArgOverflowSizeTLS);
8969 Value *CopySize = VAArgSize;
8970
8971 if (!VAStartInstrumentationList.empty()) {
8972 // If there is a va_start in this function, make a backup copy of
8973 // va_arg_tls somewhere in the function entry block.
8974
8975 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
8976 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
8977 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
8978 CopySize, kShadowTLSAlignment, false);
8979
8980 Value *SrcSize = IRB.CreateBinaryIntrinsic(
8981 Intrinsic::umin, CopySize,
8982 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
8983 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
8984 kShadowTLSAlignment, SrcSize);
8985 }
8986
8987 // Instrument va_start.
8988 // Copy va_list shadow from the backup copy of the TLS contents.
8989 for (CallInst *OrigInst : VAStartInstrumentationList) {
8990 NextNodeIRBuilder IRB(OrigInst);
8991 Value *VAListTag = OrigInst->getArgOperand(0);
8992 Value *RegSaveAreaPtrPtr = IRB.CreatePtrToInt(VAListTag, MS.IntptrTy);
8993 Value *RegSaveAreaSize = CopySize;
8994
8995 // In PPC32 va_list_tag is a struct
8996 RegSaveAreaPtrPtr =
8997 IRB.CreateAdd(RegSaveAreaPtrPtr, ConstantInt::get(MS.IntptrTy, 8));
8998
8999 // On PPC 32 reg_save_area can only hold 32 bytes of data
9000 RegSaveAreaSize = IRB.CreateBinaryIntrinsic(
9001 Intrinsic::umin, CopySize, ConstantInt::get(MS.IntptrTy, 32));
9002
9003 RegSaveAreaPtrPtr = IRB.CreateIntToPtr(RegSaveAreaPtrPtr, MS.PtrTy);
9004 Value *RegSaveAreaPtr = IRB.CreateLoad(MS.PtrTy, RegSaveAreaPtrPtr);
9005
9006 const DataLayout &DL = F.getDataLayout();
9007 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
9008 const Align Alignment = Align(IntptrSize);
9009
9010 { // Copy reg save area
9011 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
9012 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
9013 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(),
9014 Alignment, /*isStore*/ true);
9015 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy,
9016 Alignment, RegSaveAreaSize);
9017
9018 RegSaveAreaShadowPtr =
9019 IRB.CreatePtrToInt(RegSaveAreaShadowPtr, MS.IntptrTy);
9020 Value *FPSaveArea = IRB.CreateAdd(RegSaveAreaShadowPtr,
9021 ConstantInt::get(MS.IntptrTy, 32));
9022 FPSaveArea = IRB.CreateIntToPtr(FPSaveArea, MS.PtrTy);
9023 // We fill fp shadow with zeroes as uninitialized fp args should have
9024 // been found during call base check
9025 IRB.CreateMemSet(FPSaveArea, ConstantInt::getNullValue(IRB.getInt8Ty()),
9026 ConstantInt::get(MS.IntptrTy, 32), Alignment);
9027 }
9028
9029 { // Copy overflow area
9030 // RegSaveAreaSize is min(CopySize, 32) -> no overflow can occur
9031 Value *OverflowAreaSize = IRB.CreateSub(CopySize, RegSaveAreaSize);
9032
9033 Value *OverflowAreaPtrPtr = IRB.CreatePtrToInt(VAListTag, MS.IntptrTy);
9034 OverflowAreaPtrPtr =
9035 IRB.CreateAdd(OverflowAreaPtrPtr, ConstantInt::get(MS.IntptrTy, 4));
9036 OverflowAreaPtrPtr = IRB.CreateIntToPtr(OverflowAreaPtrPtr, MS.PtrTy);
9037
9038 Value *OverflowAreaPtr = IRB.CreateLoad(MS.PtrTy, OverflowAreaPtrPtr);
9039
9040 Value *OverflowAreaShadowPtr, *OverflowAreaOriginPtr;
9041 std::tie(OverflowAreaShadowPtr, OverflowAreaOriginPtr) =
9042 MSV.getShadowOriginPtr(OverflowAreaPtr, IRB, IRB.getInt8Ty(),
9043 Alignment, /*isStore*/ true);
9044
9045 Value *OverflowVAArgTLSCopyPtr =
9046 IRB.CreatePtrToInt(VAArgTLSCopy, MS.IntptrTy);
9047 OverflowVAArgTLSCopyPtr =
9048 IRB.CreateAdd(OverflowVAArgTLSCopyPtr, RegSaveAreaSize);
9049
9050 OverflowVAArgTLSCopyPtr =
9051 IRB.CreateIntToPtr(OverflowVAArgTLSCopyPtr, MS.PtrTy);
9052 IRB.CreateMemCpy(OverflowAreaShadowPtr, Alignment,
9053 OverflowVAArgTLSCopyPtr, Alignment, OverflowAreaSize);
9054 }
9055 }
9056 }
9057};
9058
9059/// SystemZ-specific implementation of VarArgHelper.
9060struct VarArgSystemZHelper : public VarArgHelperBase {
9061 static const unsigned SystemZGpOffset = 16;
9062 static const unsigned SystemZGpEndOffset = 56;
9063 static const unsigned SystemZFpOffset = 128;
9064 static const unsigned SystemZFpEndOffset = 160;
9065 static const unsigned SystemZMaxVrArgs = 8;
9066 static const unsigned SystemZRegSaveAreaSize = 160;
9067 static const unsigned SystemZOverflowOffset = 160;
9068 static const unsigned SystemZVAListTagSize = 32;
9069 static const unsigned SystemZOverflowArgAreaPtrOffset = 16;
9070 static const unsigned SystemZRegSaveAreaPtrOffset = 24;
9071
9072 bool IsSoftFloatABI;
9073 AllocaInst *VAArgTLSCopy = nullptr;
9074 AllocaInst *VAArgTLSOriginCopy = nullptr;
9075 Value *VAArgOverflowSize = nullptr;
9076
9077 enum class ArgKind {
9078 GeneralPurpose,
9079 FloatingPoint,
9080 Vector,
9081 Memory,
9082 Indirect,
9083 };
9084
9085 enum class ShadowExtension { None, Zero, Sign };
9086
9087 VarArgSystemZHelper(Function &F, MemorySanitizer &MS,
9088 MemorySanitizerVisitor &MSV)
9089 : VarArgHelperBase(F, MS, MSV, SystemZVAListTagSize),
9090 IsSoftFloatABI(F.getFnAttribute("use-soft-float").getValueAsBool()) {}
9091
9092 ArgKind classifyArgument(Type *T) {
9093 // T is a SystemZABIInfo::classifyArgumentType() output, and there are
9094 // only a few possibilities of what it can be. In particular, enums, single
9095 // element structs and large types have already been taken care of.
9096
9097 // Some i128 and fp128 arguments are converted to pointers only in the
9098 // back end.
9099 if (T->isIntegerTy(128) || T->isFP128Ty())
9100 return ArgKind::Indirect;
9101 if (T->isFloatingPointTy())
9102 return IsSoftFloatABI ? ArgKind::GeneralPurpose : ArgKind::FloatingPoint;
9103 if (T->isIntegerTy() || T->isPointerTy())
9104 return ArgKind::GeneralPurpose;
9105 if (T->isVectorTy())
9106 return ArgKind::Vector;
9107 return ArgKind::Memory;
9108 }
9109
9110 ShadowExtension getShadowExtension(const CallBase &CB, unsigned ArgNo) {
9111 // ABI says: "One of the simple integer types no more than 64 bits wide.
9112 // ... If such an argument is shorter than 64 bits, replace it by a full
9113 // 64-bit integer representing the same number, using sign or zero
9114 // extension". Shadow for an integer argument has the same type as the
9115 // argument itself, so it can be sign or zero extended as well.
9116 bool ZExt = CB.paramHasAttr(ArgNo, Attribute::ZExt);
9117 bool SExt = CB.paramHasAttr(ArgNo, Attribute::SExt);
9118 if (ZExt) {
9119 assert(!SExt);
9120 return ShadowExtension::Zero;
9121 }
9122 if (SExt) {
9123 assert(!ZExt);
9124 return ShadowExtension::Sign;
9125 }
9126 return ShadowExtension::None;
9127 }
9128
9129 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
9130 unsigned GpOffset = SystemZGpOffset;
9131 unsigned FpOffset = SystemZFpOffset;
9132 unsigned VrIndex = 0;
9133 unsigned OverflowOffset = SystemZOverflowOffset;
9134 const DataLayout &DL = F.getDataLayout();
9135 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
9136 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
9137 // SystemZABIInfo does not produce ByVal parameters.
9138 assert(!CB.isByValArgument(ArgNo));
9139 Type *T = A->getType();
9140 ArgKind AK = classifyArgument(T);
9141 if (AK == ArgKind::Indirect) {
9142 T = MS.PtrTy;
9143 AK = ArgKind::GeneralPurpose;
9144 }
9145 if (AK == ArgKind::GeneralPurpose && GpOffset >= SystemZGpEndOffset)
9146 AK = ArgKind::Memory;
9147 if (AK == ArgKind::FloatingPoint && FpOffset >= SystemZFpEndOffset)
9148 AK = ArgKind::Memory;
9149 if (AK == ArgKind::Vector && (VrIndex >= SystemZMaxVrArgs || !IsFixed))
9150 AK = ArgKind::Memory;
9151 Value *ShadowBase = nullptr;
9152 Value *OriginBase = nullptr;
9153 ShadowExtension SE = ShadowExtension::None;
9154 switch (AK) {
9155 case ArgKind::GeneralPurpose: {
9156 // Always keep track of GpOffset, but store shadow only for varargs.
9157 uint64_t ArgSize = 8;
9158 if (GpOffset + ArgSize <= kParamTLSSize) {
9159 if (!IsFixed) {
9160 SE = getShadowExtension(CB, ArgNo);
9161 uint64_t GapSize = 0;
9162 if (SE == ShadowExtension::None) {
9163 uint64_t ArgAllocSize = DL.getTypeAllocSize(T);
9164 assert(ArgAllocSize <= ArgSize);
9165 GapSize = ArgSize - ArgAllocSize;
9166 }
9167 ShadowBase = getShadowAddrForVAArgument(IRB, GpOffset + GapSize);
9168 if (MS.TrackOrigins)
9169 OriginBase = getOriginPtrForVAArgument(IRB, GpOffset + GapSize);
9170 }
9171 GpOffset += ArgSize;
9172 } else {
9173 GpOffset = kParamTLSSize;
9174 }
9175 break;
9176 }
9177 case ArgKind::FloatingPoint: {
9178 // Always keep track of FpOffset, but store shadow only for varargs.
9179 uint64_t ArgSize = 8;
9180 if (FpOffset + ArgSize <= kParamTLSSize) {
9181 if (!IsFixed) {
9182 // PoP says: "A short floating-point datum requires only the
9183 // left-most 32 bit positions of a floating-point register".
9184 // Therefore, in contrast to AK_GeneralPurpose and AK_Memory,
9185 // don't extend shadow and don't mind the gap.
9186 ShadowBase = getShadowAddrForVAArgument(IRB, FpOffset);
9187 if (MS.TrackOrigins)
9188 OriginBase = getOriginPtrForVAArgument(IRB, FpOffset);
9189 }
9190 FpOffset += ArgSize;
9191 } else {
9192 FpOffset = kParamTLSSize;
9193 }
9194 break;
9195 }
9196 case ArgKind::Vector: {
9197 // Keep track of VrIndex. No need to store shadow, since vector varargs
9198 // go through AK_Memory.
9199 assert(IsFixed);
9200 VrIndex++;
9201 break;
9202 }
9203 case ArgKind::Memory: {
9204 // Keep track of OverflowOffset and store shadow only for varargs.
9205 // Ignore fixed args, since we need to copy only the vararg portion of
9206 // the overflow area shadow.
9207 if (!IsFixed) {
9208 uint64_t ArgAllocSize = DL.getTypeAllocSize(T);
9209 uint64_t ArgSize = alignTo(ArgAllocSize, 8);
9210 if (OverflowOffset + ArgSize <= kParamTLSSize) {
9211 SE = getShadowExtension(CB, ArgNo);
9212 uint64_t GapSize =
9213 SE == ShadowExtension::None ? ArgSize - ArgAllocSize : 0;
9214 ShadowBase =
9215 getShadowAddrForVAArgument(IRB, OverflowOffset + GapSize);
9216 if (MS.TrackOrigins)
9217 OriginBase =
9218 getOriginPtrForVAArgument(IRB, OverflowOffset + GapSize);
9219 OverflowOffset += ArgSize;
9220 } else {
9221 OverflowOffset = kParamTLSSize;
9222 }
9223 }
9224 break;
9225 }
9226 case ArgKind::Indirect:
9227 llvm_unreachable("Indirect must be converted to GeneralPurpose");
9228 }
9229 if (ShadowBase == nullptr)
9230 continue;
9231 Value *Shadow = MSV.getShadow(A);
9232 if (SE != ShadowExtension::None)
9233 Shadow = MSV.CreateShadowCast(IRB, Shadow, IRB.getInt64Ty(),
9234 /*Signed*/ SE == ShadowExtension::Sign);
9235 ShadowBase = IRB.CreateIntToPtr(ShadowBase, MS.PtrTy, "_msarg_va_s");
9236 IRB.CreateStore(Shadow, ShadowBase);
9237 if (MS.TrackOrigins) {
9238 Value *Origin = MSV.getOrigin(A);
9239 TypeSize StoreSize = DL.getTypeStoreSize(Shadow->getType());
9240 MSV.paintOrigin(IRB, Origin, OriginBase, StoreSize,
9242 }
9243 }
9244 Constant *OverflowSize = ConstantInt::get(
9245 IRB.getInt64Ty(), OverflowOffset - SystemZOverflowOffset);
9246 IRB.CreateStore(OverflowSize, MS.VAArgOverflowSizeTLS);
9247 }
9248
9249 void copyRegSaveArea(IRBuilder<> &IRB, Value *VAListTag) {
9250 Value *RegSaveAreaPtrPtr = IRB.CreateIntToPtr(
9251 IRB.CreateAdd(
9252 IRB.CreatePtrToInt(VAListTag, MS.IntptrTy),
9253 ConstantInt::get(MS.IntptrTy, SystemZRegSaveAreaPtrOffset)),
9254 MS.PtrTy);
9255 Value *RegSaveAreaPtr = IRB.CreateLoad(MS.PtrTy, RegSaveAreaPtrPtr);
9256 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
9257 const Align Alignment = Align(8);
9258 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
9259 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(), Alignment,
9260 /*isStore*/ true);
9261 // TODO(iii): copy only fragments filled by visitCallBase()
9262 // TODO(iii): support packed-stack && !use-soft-float
9263 // For use-soft-float functions, it is enough to copy just the GPRs.
9264 unsigned RegSaveAreaSize =
9265 IsSoftFloatABI ? SystemZGpEndOffset : SystemZRegSaveAreaSize;
9266 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy, Alignment,
9267 RegSaveAreaSize);
9268 if (MS.TrackOrigins)
9269 IRB.CreateMemCpy(RegSaveAreaOriginPtr, Alignment, VAArgTLSOriginCopy,
9270 Alignment, RegSaveAreaSize);
9271 }
9272
9273 // FIXME: This implementation limits OverflowOffset to kParamTLSSize, so we
9274 // don't know real overflow size and can't clear shadow beyond kParamTLSSize.
9275 void copyOverflowArea(IRBuilder<> &IRB, Value *VAListTag) {
9276 Value *OverflowArgAreaPtrPtr = IRB.CreateIntToPtr(
9277 IRB.CreateAdd(
9278 IRB.CreatePtrToInt(VAListTag, MS.IntptrTy),
9279 ConstantInt::get(MS.IntptrTy, SystemZOverflowArgAreaPtrOffset)),
9280 MS.PtrTy);
9281 Value *OverflowArgAreaPtr = IRB.CreateLoad(MS.PtrTy, OverflowArgAreaPtrPtr);
9282 Value *OverflowArgAreaShadowPtr, *OverflowArgAreaOriginPtr;
9283 const Align Alignment = Align(8);
9284 std::tie(OverflowArgAreaShadowPtr, OverflowArgAreaOriginPtr) =
9285 MSV.getShadowOriginPtr(OverflowArgAreaPtr, IRB, IRB.getInt8Ty(),
9286 Alignment, /*isStore*/ true);
9287 Value *SrcPtr = IRB.CreateConstGEP1_32(IRB.getInt8Ty(), VAArgTLSCopy,
9288 SystemZOverflowOffset);
9289 IRB.CreateMemCpy(OverflowArgAreaShadowPtr, Alignment, SrcPtr, Alignment,
9290 VAArgOverflowSize);
9291 if (MS.TrackOrigins) {
9292 SrcPtr = IRB.CreateConstGEP1_32(IRB.getInt8Ty(), VAArgTLSOriginCopy,
9293 SystemZOverflowOffset);
9294 IRB.CreateMemCpy(OverflowArgAreaOriginPtr, Alignment, SrcPtr, Alignment,
9295 VAArgOverflowSize);
9296 }
9297 }
9298
9299 void finalizeInstrumentation() override {
9300 assert(!VAArgOverflowSize && !VAArgTLSCopy &&
9301 "finalizeInstrumentation called twice");
9302 if (!VAStartInstrumentationList.empty()) {
9303 // If there is a va_start in this function, make a backup copy of
9304 // va_arg_tls somewhere in the function entry block.
9305 IRBuilder<> IRB(MSV.FnPrologueEnd);
9306 VAArgOverflowSize =
9307 IRB.CreateLoad(IRB.getInt64Ty(), MS.VAArgOverflowSizeTLS);
9308 Value *CopySize =
9309 IRB.CreateAdd(ConstantInt::get(MS.IntptrTy, SystemZOverflowOffset),
9310 VAArgOverflowSize);
9311 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
9312 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
9313 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
9314 CopySize, kShadowTLSAlignment, false);
9315
9316 Value *SrcSize = IRB.CreateBinaryIntrinsic(
9317 Intrinsic::umin, CopySize,
9318 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
9319 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
9320 kShadowTLSAlignment, SrcSize);
9321 if (MS.TrackOrigins) {
9322 VAArgTLSOriginCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
9323 VAArgTLSOriginCopy->setAlignment(kShadowTLSAlignment);
9324 IRB.CreateMemCpy(VAArgTLSOriginCopy, kShadowTLSAlignment,
9325 MS.VAArgOriginTLS, kShadowTLSAlignment, SrcSize);
9326 }
9327 }
9328
9329 // Instrument va_start.
9330 // Copy va_list shadow from the backup copy of the TLS contents.
9331 for (CallInst *OrigInst : VAStartInstrumentationList) {
9332 NextNodeIRBuilder IRB(OrigInst);
9333 Value *VAListTag = OrigInst->getArgOperand(0);
9334 copyRegSaveArea(IRB, VAListTag);
9335 copyOverflowArea(IRB, VAListTag);
9336 }
9337 }
9338};
9339
9340/// i386-specific implementation of VarArgHelper.
9341struct VarArgI386Helper : public VarArgHelperBase {
9342 AllocaInst *VAArgTLSCopy = nullptr;
9343 Value *VAArgSize = nullptr;
9344
9345 VarArgI386Helper(Function &F, MemorySanitizer &MS,
9346 MemorySanitizerVisitor &MSV)
9347 : VarArgHelperBase(F, MS, MSV, /*VAListTagSize=*/4) {}
9348
9349 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
9350 const DataLayout &DL = F.getDataLayout();
9351 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
9352 unsigned VAArgOffset = 0;
9353 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
9354 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
9355 bool IsByVal = CB.isByValArgument(ArgNo);
9356 if (IsByVal) {
9357 assert(A->getType()->isPointerTy());
9358 Type *RealTy = CB.getParamByValType(ArgNo);
9359 uint64_t ArgSize = DL.getTypeAllocSize(RealTy);
9360 Align ArgAlign = CB.getParamAlign(ArgNo).value_or(Align(IntptrSize));
9361 if (ArgAlign < IntptrSize)
9362 ArgAlign = Align(IntptrSize);
9363 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
9364 if (!IsFixed) {
9365 Value *Base = getShadowPtrForVAArgument(IRB, VAArgOffset, ArgSize);
9366 if (Base) {
9367 Value *AShadowPtr, *AOriginPtr;
9368 std::tie(AShadowPtr, AOriginPtr) =
9369 MSV.getShadowOriginPtr(A, IRB, IRB.getInt8Ty(),
9370 kShadowTLSAlignment, /*isStore*/ false);
9371
9372 IRB.CreateMemCpy(Base, kShadowTLSAlignment, AShadowPtr,
9373 kShadowTLSAlignment, ArgSize);
9374 }
9375 VAArgOffset += alignTo(ArgSize, Align(IntptrSize));
9376 }
9377 } else {
9378 Value *Base;
9379 uint64_t ArgSize = DL.getTypeAllocSize(A->getType());
9380 Align ArgAlign = Align(IntptrSize);
9381 VAArgOffset = alignTo(VAArgOffset, ArgAlign);
9382 if (DL.isBigEndian()) {
9383 // Adjusting the shadow for argument with size < IntptrSize to match
9384 // the placement of bits in big endian system
9385 if (ArgSize < IntptrSize)
9386 VAArgOffset += (IntptrSize - ArgSize);
9387 }
9388 if (!IsFixed) {
9389 Base = getShadowPtrForVAArgument(IRB, VAArgOffset, ArgSize);
9390 if (Base)
9391 IRB.CreateAlignedStore(MSV.getShadow(A), Base, kShadowTLSAlignment);
9392 VAArgOffset += ArgSize;
9393 VAArgOffset = alignTo(VAArgOffset, Align(IntptrSize));
9394 }
9395 }
9396 }
9397
9398 Constant *TotalVAArgSize = ConstantInt::get(MS.IntptrTy, VAArgOffset);
9399 // Here using VAArgOverflowSizeTLS as VAArgSizeTLS to avoid creation of
9400 // a new class member i.e. it is the total size of all VarArgs.
9401 IRB.CreateStore(TotalVAArgSize, MS.VAArgOverflowSizeTLS);
9402 }
9403
9404 void finalizeInstrumentation() override {
9405 assert(!VAArgSize && !VAArgTLSCopy &&
9406 "finalizeInstrumentation called twice");
9407 IRBuilder<> IRB(MSV.FnPrologueEnd);
9408 VAArgSize = IRB.CreateLoad(MS.IntptrTy, MS.VAArgOverflowSizeTLS);
9409 Value *CopySize = VAArgSize;
9410
9411 if (!VAStartInstrumentationList.empty()) {
9412 // If there is a va_start in this function, make a backup copy of
9413 // va_arg_tls somewhere in the function entry block.
9414 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
9415 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
9416 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
9417 CopySize, kShadowTLSAlignment, false);
9418
9419 Value *SrcSize = IRB.CreateBinaryIntrinsic(
9420 Intrinsic::umin, CopySize,
9421 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
9422 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
9423 kShadowTLSAlignment, SrcSize);
9424 }
9425
9426 // Instrument va_start.
9427 // Copy va_list shadow from the backup copy of the TLS contents.
9428 for (CallInst *OrigInst : VAStartInstrumentationList) {
9429 NextNodeIRBuilder IRB(OrigInst);
9430 Value *VAListTag = OrigInst->getArgOperand(0);
9431 Type *RegSaveAreaPtrTy = PointerType::getUnqual(*MS.C);
9432 Value *RegSaveAreaPtrPtr =
9433 IRB.CreateIntToPtr(IRB.CreatePtrToInt(VAListTag, MS.IntptrTy),
9434 PointerType::get(*MS.C, 0));
9435 Value *RegSaveAreaPtr =
9436 IRB.CreateLoad(RegSaveAreaPtrTy, RegSaveAreaPtrPtr);
9437 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
9438 const DataLayout &DL = F.getDataLayout();
9439 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
9440 const Align Alignment = Align(IntptrSize);
9441 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
9442 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(),
9443 Alignment, /*isStore*/ true);
9444 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy, Alignment,
9445 CopySize);
9446 }
9447 }
9448};
9449
9450/// Implementation of VarArgHelper that is used for ARM32, MIPS, RISCV,
9451/// LoongArch64.
9452struct VarArgGenericHelper : public VarArgHelperBase {
9453 AllocaInst *VAArgTLSCopy = nullptr;
9454 Value *VAArgSize = nullptr;
9455
9456 VarArgGenericHelper(Function &F, MemorySanitizer &MS,
9457 MemorySanitizerVisitor &MSV, const unsigned VAListTagSize)
9458 : VarArgHelperBase(F, MS, MSV, VAListTagSize) {}
9459
9460 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {
9461 unsigned VAArgOffset = 0;
9462 const DataLayout &DL = F.getDataLayout();
9463 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
9464 for (const auto &[ArgNo, A] : llvm::enumerate(CB.args())) {
9465 bool IsFixed = ArgNo < CB.getFunctionType()->getNumParams();
9466 if (IsFixed)
9467 continue;
9468 uint64_t ArgSize = DL.getTypeAllocSize(A->getType());
9469 if (DL.isBigEndian()) {
9470 // Adjusting the shadow for argument with size < IntptrSize to match the
9471 // placement of bits in big endian system
9472 if (ArgSize < IntptrSize)
9473 VAArgOffset += (IntptrSize - ArgSize);
9474 }
9475 Value *Base = getShadowPtrForVAArgument(IRB, VAArgOffset, ArgSize);
9476 VAArgOffset += ArgSize;
9477 VAArgOffset = alignTo(VAArgOffset, IntptrSize);
9478 if (!Base)
9479 continue;
9480 IRB.CreateAlignedStore(MSV.getShadow(A), Base, kShadowTLSAlignment);
9481 }
9482
9483 Constant *TotalVAArgSize = ConstantInt::get(MS.IntptrTy, VAArgOffset);
9484 // Here using VAArgOverflowSizeTLS as VAArgSizeTLS to avoid creation of
9485 // a new class member i.e. it is the total size of all VarArgs.
9486 IRB.CreateStore(TotalVAArgSize, MS.VAArgOverflowSizeTLS);
9487 }
9488
9489 void finalizeInstrumentation() override {
9490 assert(!VAArgSize && !VAArgTLSCopy &&
9491 "finalizeInstrumentation called twice");
9492 IRBuilder<> IRB(MSV.FnPrologueEnd);
9493 VAArgSize = IRB.CreateLoad(MS.IntptrTy, MS.VAArgOverflowSizeTLS);
9494 Value *CopySize = VAArgSize;
9495
9496 if (!VAStartInstrumentationList.empty()) {
9497 // If there is a va_start in this function, make a backup copy of
9498 // va_arg_tls somewhere in the function entry block.
9499 VAArgTLSCopy = IRB.CreateAlloca(Type::getInt8Ty(*MS.C), CopySize);
9500 VAArgTLSCopy->setAlignment(kShadowTLSAlignment);
9501 IRB.CreateMemSet(VAArgTLSCopy, Constant::getNullValue(IRB.getInt8Ty()),
9502 CopySize, kShadowTLSAlignment, false);
9503
9504 Value *SrcSize = IRB.CreateBinaryIntrinsic(
9505 Intrinsic::umin, CopySize,
9506 ConstantInt::get(MS.IntptrTy, kParamTLSSize));
9507 IRB.CreateMemCpy(VAArgTLSCopy, kShadowTLSAlignment, MS.VAArgTLS,
9508 kShadowTLSAlignment, SrcSize);
9509 }
9510
9511 // Instrument va_start.
9512 // Copy va_list shadow from the backup copy of the TLS contents.
9513 for (CallInst *OrigInst : VAStartInstrumentationList) {
9514 NextNodeIRBuilder IRB(OrigInst);
9515 Value *VAListTag = OrigInst->getArgOperand(0);
9516 Type *RegSaveAreaPtrTy = PointerType::getUnqual(*MS.C);
9517 Value *RegSaveAreaPtrPtr =
9518 IRB.CreateIntToPtr(IRB.CreatePtrToInt(VAListTag, MS.IntptrTy),
9519 PointerType::get(*MS.C, 0));
9520 Value *RegSaveAreaPtr =
9521 IRB.CreateLoad(RegSaveAreaPtrTy, RegSaveAreaPtrPtr);
9522 Value *RegSaveAreaShadowPtr, *RegSaveAreaOriginPtr;
9523 const DataLayout &DL = F.getDataLayout();
9524 unsigned IntptrSize = DL.getTypeStoreSize(MS.IntptrTy);
9525 const Align Alignment = Align(IntptrSize);
9526 std::tie(RegSaveAreaShadowPtr, RegSaveAreaOriginPtr) =
9527 MSV.getShadowOriginPtr(RegSaveAreaPtr, IRB, IRB.getInt8Ty(),
9528 Alignment, /*isStore*/ true);
9529 IRB.CreateMemCpy(RegSaveAreaShadowPtr, Alignment, VAArgTLSCopy, Alignment,
9530 CopySize);
9531 }
9532 }
9533};
9534
9535// ARM32, Loongarch64, MIPS and RISCV share the same calling conventions
9536// regarding VAArgs.
9537using VarArgARM32Helper = VarArgGenericHelper;
9538using VarArgRISCVHelper = VarArgGenericHelper;
9539using VarArgMIPSHelper = VarArgGenericHelper;
9540using VarArgLoongArch64Helper = VarArgGenericHelper;
9541using VarArgHexagonHelper = VarArgGenericHelper;
9542
9543/// A no-op implementation of VarArgHelper.
9544struct VarArgNoOpHelper : public VarArgHelper {
9545 VarArgNoOpHelper(Function &F, MemorySanitizer &MS,
9546 MemorySanitizerVisitor &MSV) {}
9547
9548 void visitCallBase(CallBase &CB, IRBuilder<> &IRB) override {}
9549
9550 void visitVAStartInst(VAStartInst &I) override {}
9551
9552 void visitVACopyInst(VACopyInst &I) override {}
9553
9554 void finalizeInstrumentation() override {}
9555};
9556
9557} // end anonymous namespace
9558
9559static VarArgHelper *CreateVarArgHelper(Function &Func, MemorySanitizer &Msan,
9560 MemorySanitizerVisitor &Visitor) {
9561 // VarArg handling is only implemented on AMD64. False positives are possible
9562 // on other platforms.
9563 Triple TargetTriple(Func.getParent()->getTargetTriple());
9564
9565 if (TargetTriple.getArch() == Triple::x86)
9566 return new VarArgI386Helper(Func, Msan, Visitor);
9567
9568 if (TargetTriple.getArch() == Triple::x86_64)
9569 return new VarArgAMD64Helper(Func, Msan, Visitor);
9570
9571 if (TargetTriple.isARM())
9572 return new VarArgARM32Helper(Func, Msan, Visitor, /*VAListTagSize=*/4);
9573
9574 if (TargetTriple.isAArch64())
9575 return new VarArgAArch64Helper(Func, Msan, Visitor);
9576
9577 if (TargetTriple.isSystemZ())
9578 return new VarArgSystemZHelper(Func, Msan, Visitor);
9579
9580 // On PowerPC32 VAListTag is a struct
9581 // {char, char, i16 padding, char *, char *}
9582 if (TargetTriple.isPPC32())
9583 return new VarArgPowerPC32Helper(Func, Msan, Visitor);
9584
9585 if (TargetTriple.isPPC64())
9586 return new VarArgPowerPC64Helper(Func, Msan, Visitor);
9587
9588 if (TargetTriple.isRISCV32())
9589 return new VarArgRISCVHelper(Func, Msan, Visitor, /*VAListTagSize=*/4);
9590
9591 if (TargetTriple.isRISCV64())
9592 return new VarArgRISCVHelper(Func, Msan, Visitor, /*VAListTagSize=*/8);
9593
9594 if (TargetTriple.isMIPS32())
9595 return new VarArgMIPSHelper(Func, Msan, Visitor, /*VAListTagSize=*/4);
9596
9597 if (TargetTriple.isMIPS64())
9598 return new VarArgMIPSHelper(Func, Msan, Visitor, /*VAListTagSize=*/8);
9599
9600 if (TargetTriple.isLoongArch64())
9601 return new VarArgLoongArch64Helper(Func, Msan, Visitor,
9602 /*VAListTagSize=*/8);
9603
9604 if (TargetTriple.getArch() == Triple::hexagon)
9605 return new VarArgHexagonHelper(Func, Msan, Visitor, /*VAListTagSize=*/12);
9606
9607 return new VarArgNoOpHelper(Func, Msan, Visitor);
9608}
9609
9610bool MemorySanitizer::sanitizeFunction(Function &F, TargetLibraryInfo &TLI) {
9611 if (!CompileKernel && F.getName() == kMsanModuleCtorName)
9612 return false;
9613
9614 if (F.hasFnAttribute(Attribute::DisableSanitizerInstrumentation))
9615 return false;
9616
9617 MemorySanitizerVisitor Visitor(F, *this, TLI);
9618
9619 // Clear out memory attributes.
9621 B.addAttribute(Attribute::Memory).addAttribute(Attribute::Speculatable);
9622 F.removeFnAttrs(B);
9623
9624 return Visitor.runOnFunction();
9625}
#define Success
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
unsigned Imm
unsigned uint64_t
constexpr LLT S1
AMDGPU Uniform Intrinsic Combine
This file implements a class to represent arbitrary precision integral constant values and operations...
static bool isStore(int Opcode)
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
static cl::opt< ITMode > IT(cl::desc("IT block support"), cl::Hidden, cl::init(DefaultIT), cl::values(clEnumValN(DefaultIT, "arm-default-it", "Generate any type of IT block"), clEnumValN(RestrictedIT, "arm-restrict-it", "Disallow complex IT blocks")))
static const size_t kNumberOfAccessSizes
VarLocInsertPt getNextNode(const DbgRecord *DVR)
Atomic ordering constants.
This file contains the simple types necessary to represent the attributes associated with functions a...
#define X(NUM, ENUM, NAME)
Definition ELF.h:857
static GCRegistry::Add< ShadowStackGC > C("shadow-stack", "Very portable GC for uncooperative code generators")
static GCRegistry::Add< ErlangGC > A("erlang", "erlang-compatible garbage collector")
static GCRegistry::Add< StatepointGC > D("statepoint-example", "an example strategy for statepoint")
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
This file contains the declarations for the subclasses of Constant, which represent the different fla...
static bool insertModuleCtor(Module &M)
Definition CopyProf.cpp:77
const MemoryMapParams Linux_LoongArch64_MemoryMapParams
const MemoryMapParams Linux_X86_64_MemoryMapParams
static AtomicOrdering addReleaseOrdering(AtomicOrdering AO)
const MemoryMapParams Linux_S390X_MemoryMapParams
static AtomicOrdering addAcquireOrdering(AtomicOrdering AO)
const MemoryMapParams Linux_AArch64_MemoryMapParams
static bool isAMustTailRetVal(Value *RetVal)
This file provides an implementation of debug counters.
#define DEBUG_COUNTER(VARNAME, COUNTERNAME, DESC)
This file defines the DenseMap class.
This file builds on the ADT/GraphTraits.h file to build generic depth first graph iterator.
static bool runOnFunction(Function &F, bool PostInlining)
This is the interface for a simple mod/ref and alias analysis over globals.
static size_t TypeSizeToSizeIndex(uint32_t TypeSize)
#define op(i)
Hexagon Common GEP
#define _
Module.h This file contains the declarations for the Module class.
static LVOptions Options
Definition LVOptions.cpp:25
static bool isZero(Value *V, const DataLayout &DL, DominatorTree *DT, AssumptionCache *AC)
Definition Lint.cpp:540
#define F(x, y, z)
Definition MD5.cpp:54
#define I(x, y, z)
Definition MD5.cpp:57
static const PlatformMemoryMapParams Linux_S390_MemoryMapParams
static const Align kMinOriginAlignment
static const PlatformMemoryMapParams Linux_X86_MemoryMapParams
static const PlatformMemoryMapParams Linux_LoongArch_MemoryMapParams
static const MemoryMapParams NetBSD_X86_64_MemoryMapParams
static const PlatformMemoryMapParams Linux_MIPS_MemoryMapParams
static const unsigned kOriginSize
static const Align kShadowTLSAlignment
static const PlatformMemoryMapParams Linux_ARM_MemoryMapParams
static Constant * getOrInsertGlobal(Module &M, StringRef Name, Type *Ty)
static const PlatformMemoryMapParams Linux_Hexagon_MemoryMapParams_P
static const MemoryMapParams Linux_I386_MemoryMapParams
const char kMsanInitName[]
OddOrEvenLanes
@ kOddLanes
@ kEvenLanes
@ kBothLanes
static const MemoryMapParams FreeBSD_X86_64_MemoryMapParams
static GlobalVariable * createPrivateConstGlobalForString(Module &M, StringRef Str)
Create a non-const global initialized with the given string.
static const PlatformMemoryMapParams Linux_PowerPC_MemoryMapParams
static const size_t kNumberOfAccessSizes
static VarArgHelper * CreateVarArgHelper(Function &Func, MemorySanitizer &Msan, MemorySanitizerVisitor &Visitor)
static const MemoryMapParams Linux_MIPS64_MemoryMapParams
static const MemoryMapParams Linux_PowerPC64_MemoryMapParams
static const MemoryMapParams Linux_Hexagon_MemoryMapParams
static const PlatformMemoryMapParams FreeBSD_X86_MemoryMapParams
static const PlatformMemoryMapParams FreeBSD_ARM_MemoryMapParams
static const unsigned kParamTLSSize
static const PlatformMemoryMapParams NetBSD_X86_MemoryMapParams
static const unsigned kRetvalTLSSize
static const MemoryMapParams FreeBSD_AArch64_MemoryMapParams
const char kMsanModuleCtorName[]
static const MemoryMapParams FreeBSD_I386_MemoryMapParams
#define T
uint64_t IntrinsicInst * II
FunctionAnalysisManager FAM
if(PassOpts->AAPipeline)
const SmallVectorImpl< MachineOperand > & Cond
static void visit(BasicBlock &Start, std::function< bool(BasicBlock *)> op)
static const char * name
This file implements a set that has insertion order iteration characteristics.
This file defines the SmallPtrSet class.
This file defines the SmallVector class.
This file contains some functions that are useful when dealing with strings.
#define LLVM_DEBUG(...)
Definition Debug.h:119
static SymbolRef::Type getType(const Symbol *Sym)
Definition TapiFile.cpp:39
Value * RHS
Value * LHS
static APInt getSignedMinValue(unsigned numBits)
Gets minimum signed value of APInt for a specific bit width.
Definition APInt.h:215
void setAlignment(Align Align)
PassT::Result & getResult(IRUnitT &IR, ExtraArgTs... ExtraArgs)
Get the result of an analysis pass for a given IR unit.
const T & front() const
Get the first element.
Definition ArrayRef.h:144
static LLVM_ABI ArrayType * get(Type *ElementType, uint64_t NumElements)
This static method is the primary way to construct an ArrayType.
This class stores enough information to efficiently remove some attributes from an existing AttrBuild...
AttributeMask & addAttribute(Attribute::AttrKind Val)
Add an attribute to the mask.
iterator end()
Definition BasicBlock.h:459
LLVM_ABI const_iterator getFirstInsertionPt() const
Returns an iterator to the first instruction in this block that is suitable for inserting a non-PHI i...
LLVM_ABI const BasicBlock * getSinglePredecessor() const
Return the predecessor of this block if it has a single predecessor block.
InstListType::iterator iterator
Instruction iterators...
Definition BasicBlock.h:170
bool isInlineAsm() const
Check if this call is an inline asm statement.
Function * getCalledFunction() const
Returns the function called, or null if this is an indirect function invocation or the function signa...
bool hasRetAttr(Attribute::AttrKind Kind) const
Determine whether the return value has the given attribute.
LLVM_ABI bool paramHasAttr(unsigned ArgNo, Attribute::AttrKind Kind) const
Determine whether the argument or parameter has the given attribute.
bool isByValArgument(unsigned ArgNo) const
Determine whether this argument is passed by value.
void removeFnAttrs(const AttributeMask &AttrsToRemove)
Removes the attributes from the function.
void setCannotMerge()
MaybeAlign getParamAlign(unsigned ArgNo) const
Extract the alignment for a call or parameter (0=unknown).
Type * getParamByValType(unsigned ArgNo) const
Extract the byval type for a call or parameter.
Value * getCalledOperand() const
Type * getParamElementType(unsigned ArgNo) const
Extract the elementtype type for a parameter.
Value * getArgOperand(unsigned i) const
void setArgOperand(unsigned i, Value *v)
FunctionType * getFunctionType() const
iterator_range< User::op_iterator > args()
Iteration adapter for range-for loops.
void addParamAttr(unsigned ArgNo, Attribute::AttrKind Kind)
Adds the attribute to the indicated argument.
Predicate
This enumeration lists the possible predicates for CmpInst subclasses.
Definition InstrTypes.h:740
@ ICMP_SLT
signed less than
Definition InstrTypes.h:769
@ ICMP_SLE
signed less or equal
Definition InstrTypes.h:770
@ ICMP_SGT
signed greater than
Definition InstrTypes.h:767
@ ICMP_SGE
signed greater or equal
Definition InstrTypes.h:768
static LLVM_ABI Constant * get(ArrayType *T, ArrayRef< Constant * > V)
static LLVM_ABI Constant * getString(LLVMContext &Context, StringRef Initializer, bool AddNull=true, bool ByteString=false)
This method constructs a CDS and initializes it with a text string.
static LLVM_ABI Constant * get(LLVMContext &Context, ArrayRef< uint8_t > Elts)
get() constructors - Return a constant with vector type with an element count and element type matchi...
static ConstantInt * getSigned(IntegerType *Ty, int64_t V, bool ImplicitTrunc=false)
Return a ConstantInt with the specified value for the specified type.
Definition Constants.h:135
static LLVM_ABI ConstantInt * getBool(LLVMContext &Context, bool V)
static LLVM_ABI Constant * get(StructType *T, ArrayRef< Constant * > V)
static LLVM_ABI Constant * getSplat(ElementCount EC, Constant *Elt)
Return a ConstantVector with the specified constant in each element.
static LLVM_ABI Constant * get(ArrayRef< Constant * > V)
This is an important base class in LLVM.
Definition Constant.h:43
bool isNullValue() const
Return true if this is the value that would be returned by getNullValue.
Definition Constant.h:64
static LLVM_ABI Constant * getAllOnesValue(Type *Ty)
LLVM_ABI bool isAllOnesValue() const
Return true if this is the value that would be returned by getAllOnesValue.
Definition Constants.cpp:68
static LLVM_ABI Constant * getNullValue(Type *Ty)
Constructor to create a '0' constant of arbitrary type.
LLVM_ABI Constant * getAggregateElement(unsigned Elt) const
For aggregates (struct/array/vector) return the constant that corresponds to the specified element if...
static bool shouldExecute(CounterInfo &Counter)
bool empty() const
Definition DenseMap.h:717
unsigned getNumElements() const
static LLVM_ABI FixedVectorType * get(Type *ElementType, unsigned NumElts)
Definition Type.cpp:843
static FixedVectorType * getHalfElementsVectorType(FixedVectorType *VTy)
A handy container for a FunctionType+Callee-pointer pair, which can be passed around as a single enti...
unsigned getNumParams() const
Return the number of fixed parameters this function type requires.
LLVM_ABI void setComdat(Comdat *C)
Definition Globals.cpp:287
@ PrivateLinkage
Like Internal, but omit from symbol table.
Definition GlobalValue.h:61
@ ExternalLinkage
Externally visible function.
Definition GlobalValue.h:53
Analysis pass providing a never-invalidated alias analysis result.
ConstantInt * getInt1(bool V)
Get a constant value representing either true or false.
Definition IRBuilder.h:449
LLVM_ABI CallInst * CreateIntrinsicWithoutFolding(Intrinsic::ID ID, ArrayRef< Type * > OverloadTypes, ArrayRef< Value * > Args, FMFSource FMFSource={}, const Twine &Name="", ArrayRef< OperandBundleDef > OpBundles={})
Create a call to intrinsic ID with Args, mangled using OverloadTypes.
LLVM_ABI Value * CreateAndReduce(Value *Src)
Create a vector int AND reduction intrinsic of the source vector.
Value * CreateInsertElement(Type *VecTy, Value *NewElt, Value *Idx, const Twine &Name="")
Definition IRBuilder.h:2678
Value * CreateConstGEP1_32(Type *Ty, Value *Ptr, unsigned Idx0, const Twine &Name="")
Definition IRBuilder.h:2033
AllocaInst * CreateAlloca(Type *Ty, unsigned AddrSpace, Value *ArraySize=nullptr, const Twine &Name="")
Definition IRBuilder.h:1889
IntegerType * getInt1Ty()
Fetch the type representing a single bit.
Definition IRBuilder.h:516
LLVM_ABI CallInst * CreateMaskedCompressStore(Value *Val, Value *Ptr, MaybeAlign Align, Value *Mask=nullptr)
Create a call to Masked Compress Store intrinsic.
Value * CreateInsertValue(Value *Agg, Value *Val, ArrayRef< unsigned > Idxs, const Twine &Name="")
Definition IRBuilder.h:2732
LLVM_ABI Value * CreateAllocationSize(Type *DestTy, AllocaInst *AI)
Get allocation size of an alloca as a runtime Value* (handles both static and dynamic allocas and vsc...
Value * CreateExtractElement(Value *Vec, Value *Idx, const Twine &Name="")
Definition IRBuilder.h:2666
IntegerType * getIntNTy(unsigned N)
Fetch the type representing an N-bit integer.
Definition IRBuilder.h:544
LoadInst * CreateAlignedLoad(Type *Ty, Value *Ptr, MaybeAlign Align, const char *Name)
Definition IRBuilder.h:1943
CallInst * CreateMemCpy(Value *Dst, MaybeAlign DstAlign, Value *Src, MaybeAlign SrcAlign, uint64_t Size, bool isVolatile=false, const AAMDNodes &AAInfo=AAMDNodes())
Create and insert a memcpy between the specified pointers.
Definition IRBuilder.h:660
Value * CreatePointerCast(Value *V, Type *DestTy, const Twine &Name="")
Definition IRBuilder.h:2306
Value * CreateExtractValue(Value *Agg, ArrayRef< unsigned > Idxs, const Twine &Name="")
Definition IRBuilder.h:2725
LLVM_ABI CallInst * CreateMaskedLoad(Type *Ty, Value *Ptr, Align Alignment, Value *Mask, Value *PassThru=nullptr, const Twine &Name="")
Create a call to Masked Load intrinsic.
LLVM_ABI Value * CreateSelect(Value *C, Value *True, Value *False, const Twine &Name="", Instruction *MDFrom=nullptr)
BasicBlock::iterator GetInsertPoint() const
Definition IRBuilder.h:180
Value * CreateSExt(Value *V, Type *DestTy, const Twine &Name="")
Definition IRBuilder.h:2142
Value * CreateIntToPtr(Value *V, Type *DestTy, const Twine &Name="")
Definition IRBuilder.h:2247
Value * CreateLShr(Value *LHS, Value *RHS, const Twine &Name="", bool isExact=false)
Definition IRBuilder.h:1537
IntegerType * getInt32Ty()
Fetch the type representing a 32-bit integer.
Definition IRBuilder.h:531
ConstantInt * getInt8(uint8_t C)
Get a constant 8-bit value.
Definition IRBuilder.h:464
Value * CreatePtrAdd(Value *Ptr, Value *Offset, const Twine &Name="", GEPNoWrapFlags NW=GEPNoWrapFlags::none())
Definition IRBuilder.h:2101
IntegerType * getInt64Ty()
Fetch the type representing a 64-bit integer.
Definition IRBuilder.h:536
Value * CreateUDiv(Value *LHS, Value *RHS, const Twine &Name="", bool isExact=false)
Definition IRBuilder.h:1478
Value * CreateICmpNE(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:2395
Value * CreateGEP(Type *Ty, Value *Ptr, ArrayRef< Value * > IdxList, const Twine &Name="", GEPNoWrapFlags NW=GEPNoWrapFlags::none())
Definition IRBuilder.h:2020
Value * CreateNeg(Value *V, const Twine &Name="", bool HasNSW=false)
Definition IRBuilder.h:1835
LLVM_ABI Value * CreateBinaryIntrinsic(Intrinsic::ID ID, Value *LHS, Value *RHS, FMFSource FMFSource={}, const Twine &Name="")
Create a call to intrinsic ID with 2 operands which is mangled on the first type.
LLVM_ABI Value * CreateOrReduce(Value *Src)
Create a vector int OR reduction intrinsic of the source vector.
ConstantInt * getInt32(uint32_t C)
Get a constant 32-bit value.
Definition IRBuilder.h:474
PHINode * CreatePHI(Type *Ty, unsigned NumReservedValues, const Twine &Name="")
Definition IRBuilder.h:2556
Value * CreateNot(Value *V, const Twine &Name="")
Definition IRBuilder.h:1859
Value * CreateICmpEQ(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:2391
LLVM_ABI DebugLoc getCurrentDebugLocation() const
Get location information used by debugging information.
Definition IRBuilder.cpp:65
Value * CreateSub(Value *LHS, Value *RHS, const Twine &Name="", bool HasNUW=false, bool HasNSW=false)
Definition IRBuilder.h:1444
Value * CreateBitCast(Value *V, Type *DestTy, const Twine &Name="")
Definition IRBuilder.h:2252
ConstantInt * getIntN(unsigned N, uint64_t C)
Get a constant N-bit value, zero extended from a 64-bit value.
Definition IRBuilder.h:484
LoadInst * CreateLoad(Type *Ty, Value *Ptr, const char *Name)
Provided to resolve 'CreateLoad(Ty, Ptr, "...")' correctly, instead of converting the string to 'bool...
Definition IRBuilder.h:1916
Value * CreateShl(Value *LHS, Value *RHS, const Twine &Name="", bool HasNUW=false, bool HasNSW=false)
Definition IRBuilder.h:1516
CallInst * CreateMemSet(Value *Ptr, Value *Val, uint64_t Size, MaybeAlign Align, bool isVolatile=false, const AAMDNodes &AAInfo=AAMDNodes())
Create and insert a memset to the specified pointer and the specified value.
Definition IRBuilder.h:605
Value * CreateZExt(Value *V, Type *DestTy, const Twine &Name="", bool IsNonNeg=false)
Definition IRBuilder.h:2130
Value * CreateShuffleVector(Value *V1, Value *V2, Value *Mask, const Twine &Name="")
Definition IRBuilder.h:2700
LLVMContext & getContext() const
Definition IRBuilder.h:181
Value * CreateAnd(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:1575
LLVM_ABI Value * CreateIntrinsic(Intrinsic::ID ID, ArrayRef< Type * > OverloadTypes, ArrayRef< Value * > Args, FMFSource FMFSource={}, const Twine &Name="", ArrayRef< OperandBundleDef > OpBundles={}, function_ref< void(CallInst *)> SetFn=[](CallInst *) {})
Variant to create a possibly constant-folded intrinsic.
StoreInst * CreateStore(Value *Val, Value *Ptr, bool isVolatile=false)
Definition IRBuilder.h:1934
LLVM_ABI CallInst * CreateMaskedStore(Value *Val, Value *Ptr, Align Alignment, Value *Mask)
Create a call to Masked Store intrinsic.
Value * CreateAdd(Value *LHS, Value *RHS, const Twine &Name="", bool HasNUW=false, bool HasNSW=false)
Definition IRBuilder.h:1427
Value * CreatePtrToInt(Value *V, Type *DestTy, const Twine &Name="")
Definition IRBuilder.h:2242
Value * CreateIsNotNull(Value *Arg, const Twine &Name="")
Return a boolean value testing if Arg != 0.
Definition IRBuilder.h:2772
CallInst * CreateCall(FunctionType *FTy, Value *Callee, ArrayRef< Value * > Args={}, const Twine &Name="", MDNode *FPMathTag=nullptr)
Definition IRBuilder.h:2570
Value * CreateTrunc(Value *V, Type *DestTy, const Twine &Name="", bool IsNUW=false, bool IsNSW=false)
Definition IRBuilder.h:2116
PointerType * getPtrTy(unsigned AddrSpace=0)
Fetch the type representing a pointer.
Definition IRBuilder.h:574
Value * CreateBinOp(Instruction::BinaryOps Opc, Value *LHS, Value *RHS, const Twine &Name="", MDNode *FPMathTag=nullptr)
Definition IRBuilder.h:1736
Value * CreateICmpSLT(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:2423
LLVM_ABI Value * CreateTypeSize(Type *Ty, TypeSize Size)
Create an expression which evaluates to the number of units in Size at runtime.
Value * CreateICmpUGE(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:2403
Value * CreateIntCast(Value *V, Type *DestTy, bool isSigned, const Twine &Name="")
Definition IRBuilder.h:2332
Value * CreateIsNull(Value *Arg, const Twine &Name="")
Return a boolean value testing if Arg == 0.
Definition IRBuilder.h:2767
void SetInsertPoint(BasicBlock *TheBB)
This specifies that created instructions should be appended to the end of the specified block.
Definition IRBuilder.h:199
Type * getVoidTy()
Fetch the type representing void.
Definition IRBuilder.h:569
StoreInst * CreateAlignedStore(Value *Val, Value *Ptr, MaybeAlign Align, bool isVolatile=false)
Definition IRBuilder.h:1962
LLVM_ABI CallInst * CreateMaskedExpandLoad(Type *Ty, Value *Ptr, MaybeAlign Align, Value *Mask=nullptr, Value *PassThru=nullptr, const Twine &Name="")
Create a call to Masked Expand Load intrinsic.
Value * CreateInBoundsPtrAdd(Value *Ptr, Value *Offset, const Twine &Name="")
Definition IRBuilder.h:2106
Value * CreateAShr(Value *LHS, Value *RHS, const Twine &Name="", bool isExact=false)
Definition IRBuilder.h:1556
Value * CreateXor(Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:1627
Value * CreateICmp(CmpInst::Predicate P, Value *LHS, Value *RHS, const Twine &Name="")
Definition IRBuilder.h:2501
Value * CreateOr(Value *LHS, Value *RHS, const Twine &Name="", bool IsDisjoint=false)
Definition IRBuilder.h:1597
IntegerType * getInt8Ty()
Fetch the type representing an 8-bit integer.
Definition IRBuilder.h:521
Value * CreateMul(Value *LHS, Value *RHS, const Twine &Name="", bool HasNUW=false, bool HasNSW=false)
Definition IRBuilder.h:1461
LLVM_ABI CallInst * CreateMaskedScatter(Value *Val, Value *Ptrs, Align Alignment, Value *Mask=nullptr)
Create a call to Masked Scatter intrinsic.
LLVM_ABI CallInst * CreateMaskedGather(Type *Ty, Value *Ptrs, Align Alignment, Value *Mask=nullptr, Value *PassThru=nullptr, const Twine &Name="")
Create a call to Masked Gather intrinsic.
Value * CreateFCmpULT(Value *LHS, Value *RHS, const Twine &Name="", MDNode *FPMathTag=nullptr)
Definition IRBuilder.h:2486
This provides a uniform API for creating instructions and inserting them into a basic block: either a...
Definition IRBuilder.h:2918
std::vector< ConstraintInfo > ConstraintInfoVector
Definition InlineAsm.h:123
void visit(Iterator Start, Iterator End)
Definition InstVisitor.h:87
const DebugLoc & getDebugLoc() const
Return the debug location for this node as a DebugLoc.
LLVM_ABI InstListType::iterator eraseFromParent()
This method unlinks 'this' from the containing basic block and deletes it.
MDNode * getMetadata(unsigned KindID) const
Get the metadata of given kind attached to this Instruction.
LLVM_ABI bool comesBefore(const Instruction *Other) const
Given an instruction Other in the same basic block as this instruction, return true if this instructi...
static LLVM_ABI IntegerType * get(LLVMContext &C, unsigned NumBits)
This static method is the primary way of constructing an IntegerType.
Definition Type.cpp:338
LLVM_ABI MDNode * createUnlikelyBranchWeights()
Return metadata containing two branch weights, with significant bias towards false destination.
Definition MDBuilder.cpp:48
A Module instance is used to store all the information related to an LLVM module.
Definition Module.h:68
void addIncoming(Value *V, BasicBlock *BB)
Add an incoming value to the end of the PHI list.
static LLVM_ABI PoisonValue * get(Type *T)
Static factory methods - Return an 'poison' object of the specified type.
A set of analyses that are preserved following a run of a transformation pass.
Definition Analysis.h:112
static PreservedAnalyses none()
Convenience factory function for the empty preserved set.
Definition Analysis.h:115
static PreservedAnalyses all()
Construct a special preserved set that preserves all passes.
Definition Analysis.h:118
PreservedAnalyses & abandon()
Mark an analysis as abandoned.
Definition Analysis.h:171
bool remove(const value_type &X)
Remove an item from the set vector.
Definition SetVector.h:187
bool insert(const value_type &X)
Insert a new element into the SetVector.
Definition SetVector.h:157
void append(ItTy in_start, ItTy in_end)
Add the specified range to the end of the SmallVector.
void push_back(const T &Elt)
Represent a constant reference to a string, i.e.
Definition StringRef.h:56
static LLVM_ABI StructType * get(LLVMContext &Context, ArrayRef< Type * > Elements, bool isPacked=false)
This static method is the primary way to create a literal StructType.
Definition Type.cpp:467
unsigned getNumElements() const
Random access to the elements.
Type * getElementType(unsigned N) const
Analysis pass providing the TargetLibraryInfo.
Provides information about what library functions are available for the current target.
AttributeList getAttrList(LLVMContext *C, ArrayRef< unsigned > ArgNos, bool Signed, bool Ret=false, AttributeList AL=AttributeList()) const
LibFunc getLibFunc(StringRef funcName) const
Searches for a particular function name.
Triple - Helper class for working with autoconf configuration names.
Definition Triple.h:48
bool isMIPS64() const
Tests whether the target is MIPS 64-bit (little and big endian).
Definition Triple.h:1133
@ loongarch64
Definition Triple.h:66
bool isRISCV32() const
Tests whether the target is 32-bit RISC-V.
Definition Triple.h:1174
bool isPPC32() const
Tests whether the target is 32-bit PowerPC (little and big endian).
Definition Triple.h:1147
ArchType getArch() const
Get the parsed architecture type of this triple.
Definition Triple.h:515
bool isRISCV64() const
Tests whether the target is 64-bit RISC-V.
Definition Triple.h:1179
bool isLoongArch64() const
Tests whether the target is 64-bit LoongArch.
Definition Triple.h:1122
bool isMIPS32() const
Tests whether the target is MIPS 32-bit (little and big endian).
Definition Triple.h:1128
bool isARM() const
Tests whether the target is ARM (little and big endian).
Definition Triple.h:1006
bool isPPC64() const
Tests whether the target is 64-bit PowerPC (little and big endian).
Definition Triple.h:1152
bool isAArch64() const
Tests whether the target is AArch64 (little and big endian).
Definition Triple.h:1099
bool isSystemZ() const
Tests whether the target is SystemZ.
Definition Triple.h:1198
The instances of the Type class are immutable: once they are created, they are never changed.
Definition Type.h:46
LLVM_ABI unsigned getIntegerBitWidth() const
bool isVectorTy() const
True if this is an instance of VectorType.
Definition Type.h:283
bool isArrayTy() const
True if this is an instance of ArrayType.
Definition Type.h:274
bool isIntOrIntVectorTy() const
Return true if this is an integer type or a vector of integer types.
Definition Type.h:258
bool isPointerTy() const
True if this is an instance of PointerType.
Definition Type.h:277
Type * getArrayElementType() const
Definition Type.h:420
bool isPPC_FP128Ty() const
Return true if this is powerpc long double.
Definition Type.h:167
bool isSized() const
Return true if it makes sense to take the size of this type.
Definition Type.h:321
static LLVM_ABI Type * getVoidTy(LLVMContext &C)
Definition Type.cpp:272
Type * getScalarType() const
If this is a vector type, return the element type, otherwise return 'this'.
Definition Type.h:363
LLVM_ABI TypeSize getPrimitiveSizeInBits() const LLVM_READONLY
Return the basic size of this type if it is a primitive type.
Definition Type.cpp:187
LLVM_ABI unsigned getScalarSizeInBits() const LLVM_READONLY
If this is a vector type, return the getPrimitiveSizeInBits value for the element type.
Definition Type.cpp:222
bool isFloatingPointTy() const
Return true if this is one of the floating-point types.
Definition Type.h:186
LLVM_ABI bool isScalableTy() const
Return true if this is a type whose size is a known multiple of vscale.
Definition Type.cpp:61
bool isIntOrPtrTy() const
Return true if this is an integer type or a pointer type.
Definition Type.h:265
bool isIntegerTy() const
True if this is an instance of IntegerType.
Definition Type.h:252
bool isFPOrFPVectorTy() const
Return true if this is a FP type or a vector of FP.
Definition Type.h:222
bool isVoidTy() const
Return true if this is 'void'.
Definition Type.h:141
Value * getOperand(unsigned i) const
Definition User.h:207
unsigned getNumOperands() const
Definition User.h:229
size_type count(const KeyT &Val) const
Return 1 if the specified key is in the map, 0 otherwise.
Definition ValueMap.h:156
Type * getType() const
All values are typed, get the type of this value.
Definition Value.h:257
LLVM_ABI void setName(const Twine &Name)
Change the name of the value.
Definition Value.cpp:394
LLVM_ABI StringRef getName() const
Return a constant reference to the value's name.
Definition Value.cpp:319
ElementCount getElementCount() const
Return an ElementCount instance to represent the (possibly scalable) number of elements in the vector...
Type * getElementType() const
constexpr ScalarTy getFixedValue() const
Definition TypeSize.h:200
constexpr bool isScalable() const
Returns whether the quantity is scaled by a runtime quantity (vscale).
Definition TypeSize.h:168
An efficient, type-erasing, non-owning reference to a callable.
const ParentTy * getParent() const
Definition ilist_node.h:34
self_iterator getIterator()
Definition ilist_node.h:123
This class implements an extremely fast bulk output stream that can only output to a stream.
Definition raw_ostream.h:53
CallInst * Call
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
constexpr char Align[]
Key for Kernel::Arg::Metadata::mAlign.
constexpr std::underlying_type_t< E > Mask()
Get a bitmask with 1s in all places up to the high-order bit of E's largest value.
@ BasicBlock
Various leaf nodes.
Definition ISDOpcodes.h:83
LLVM_ABI StringRef getBaseName(ID id)
Return the LLVM name for an intrinsic, without encoded types for overloading, such as "llvm....
Function * Kernel
Summary of a kernel (=entry point for target offloading).
Definition OpenMPOpt.h:21
NodeAddr< FuncNode * > Func
Definition RDFGraph.h:393
friend class Instruction
Iterator for Instructions in a `BasicBlock.
Definition BasicBlock.h:73
This is an optimization pass for GlobalISel generic memory operations.
unsigned Log2_32_Ceil(uint32_t Value)
Return the ceil log base 2 of the specified value, 32 if the value is zero.
Definition MathExtras.h:339
@ Offset
Definition DWP.cpp:577
auto size(R &&Range, std::enable_if_t< std::is_base_of< std::random_access_iterator_tag, typename std::iterator_traits< decltype(Range.begin())>::iterator_category >::value, void > *=nullptr)
Get the size of a range.
Definition STLExtras.h:1685
RelativeUniformCounterPtr Values
Definition InstrProf.h:91
auto enumerate(FirstRange &&First, RestRanges &&...Rest)
Given two or more input ranges, returns a new range whose values are tuples (A, B,...
Definition STLExtras.h:2570
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
@ Done
Definition Threading.h:60
bool isAligned(Align Lhs, uint64_t SizeInBytes)
Checks that SizeInBytes is a multiple of the alignment.
Definition Alignment.h:139
@ Store
The extracted value is stored (ExtractElement only).
LLVM_ABI std::pair< Instruction *, Value * > SplitBlockAndInsertSimpleForLoop(Value *End, BasicBlock::iterator SplitBefore)
Insert a for (int i = 0; i < End; i++) loop structure (with the exception that End is assumed > 0,...
InnerAnalysisManagerProxy< FunctionAnalysisManager, Module > FunctionAnalysisManagerModuleProxy
Provide the FunctionAnalysisManager to Module proxy.
constexpr bool isPowerOf2_64(uint64_t Value)
Return true if the argument is a power of two > 0 (64 bit edition.)
Definition MathExtras.h:285
unsigned Log2_64(uint64_t Value)
Return the floor log base 2 of the specified value, -1 if the value is zero.
Definition MathExtras.h:332
RelativeUniformCounterPtr ValuesPtrExpr VTableAddr Value
Definition InstrProf.h:143
LLVM_ABI bool removeUnreachableBlocks(Function &F, DomTreeUpdater *DTU=nullptr, MemorySSAUpdater *MSSAU=nullptr, bool FoldInstsToUnreachable=true)
Remove all blocks that can not be reached from the function's entry.
Definition Local.cpp:2916
auto dyn_cast_or_null(const Y &Val)
Definition Casting.h:753
LLVM_ABI std::pair< Function *, FunctionCallee > getOrCreateSanitizerCtorAndInitFunctions(Module &M, StringRef CtorName, StringRef InitName, ArrayRef< Type * > InitArgTypes, ArrayRef< Value * > InitArgs, function_ref< void(Function *, FunctionCallee)> FunctionsCreatedCallback, StringRef VersionCheckName=StringRef(), bool Weak=false)
Creates sanitizer constructor function lazily.
LLVM_ABI raw_ostream & dbgs()
dbgs() - This returns a reference to a raw_ostream for debugging messages.
Definition Debug.cpp:209
IRBuilder(LLVMContext &, FolderTy, InserterTy) -> IRBuilder< FolderTy, InserterTy >
LLVM_ABI void report_fatal_error(Error Err, bool gen_crash_diag=true)
Definition Error.cpp:163
constexpr uint64_t alignTo(uint64_t Size, Align A)
Returns a multiple of A needed to store Size bytes.
Definition Alignment.h:149
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
bool isa(const From &Val)
isa<X> - Return true if the parameter to the template is an instance of one of the template type argu...
Definition Casting.h:547
LLVM_ABI bool isKnownNonZero(const Value *V, const SimplifyQuery &Q, unsigned Depth=0)
Return true if the given value is known to be non-zero when defined.
LLVM_ABI raw_fd_ostream & errs()
This returns a reference to a raw_ostream for standard error.
AtomicOrdering
Atomic ordering for LLVM's memory model.
@ First
Helpers to iterate all locations in the MemoryEffectsBase class.
Definition ModRef.h:74
@ Or
Bitwise or logical OR of integers.
@ And
Bitwise or logical AND of integers.
@ Add
Sum of integers.
IntPtrTy
Definition InstrProf.h:82
DWARFExpression::Operation Op
RoundingMode
Rounding mode.
ArrayRef(const T &OneElt) -> ArrayRef< T >
constexpr unsigned BitWidth
LLVM_ABI void appendToGlobalCtors(Module &M, Function *F, int Priority, Constant *Data=nullptr)
Append F to the list of global ctors of module M with the given Priority.
decltype(auto) cast(const From &Val)
cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:559
constexpr bool valueOr(BoolOrDefault X, bool Default)
iterator_range< df_iterator< T > > depth_first(const T &G)
LLVM_ABI Instruction * SplitBlockAndInsertIfThen(Value *Cond, BasicBlock::iterator SplitBefore, bool Unreachable, MDNode *BranchWeights=nullptr, DomTreeUpdater *DTU=nullptr, LoopInfo *LI=nullptr, BasicBlock *ThenBlock=nullptr)
Split the containing block at the specified instruction - everything before SplitBefore stays in the ...
LLVM_ABI void maybeMarkSanitizerLibraryCallNoBuiltin(CallInst *CI, const TargetLibraryInfo *TLI)
Given a CallInst, check if it calls a string function known to CodeGen, and mark it with NoBuiltin if...
Definition Local.cpp:3902
LLVM_ABI bool checkIfAlreadyInstrumented(Module &M, StringRef Flag)
Check if module has flag attached, if not add the flag.
std::string itostr(int64_t X)
AnalysisManager< Module > ModuleAnalysisManager
Convenience typedef for the Module analysis manager.
Definition MIRParser.h:39
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:82
LLVM_ABI void printPipeline(raw_ostream &OS, function_ref< StringRef(StringRef)> MapClassName2PassName)
LLVM_ABI PreservedAnalyses run(Module &M, ModuleAnalysisManager &AM)