LLVM 24.0.0git
AArch64FrameLowering.cpp
Go to the documentation of this file.
1//===- AArch64FrameLowering.cpp - AArch64 Frame Lowering -------*- C++ -*-====//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8//
9// This file contains the AArch64 implementation of TargetFrameLowering class.
10//
11// On AArch64, stack frames are structured as follows:
12//
13// The stack grows downward.
14//
15// All of the individual frame areas on the frame below are optional, i.e. it's
16// possible to create a function so that the particular area isn't present
17// in the frame.
18//
19// At function entry, the "frame" looks as follows:
20//
21// | | Higher address
22// |-----------------------------------|
23// | |
24// | arguments passed on the stack |
25// | |
26// |-----------------------------------| <- sp
27// | | Lower address
28//
29//
30// After the prologue has run, the frame has the following general structure.
31// Note that this doesn't depict the case where a red-zone is used. Also,
32// technically the last frame area (VLAs) doesn't get created until in the
33// main function body, after the prologue is run. However, it's depicted here
34// for completeness.
35//
36// | | Higher address
37// |-----------------------------------|
38// | |
39// | arguments passed on the stack |
40// | |
41// |-----------------------------------|
42// | |
43// | (Win64 only) varargs from reg |
44// | |
45// |-----------------------------------|
46// | |
47// | (Win64 only) callee-saved SVE reg |
48// | |
49// |-----------------------------------|
50// | |
51// | callee-saved gpr registers | <--.
52// | | | On Darwin platforms these
53// |- - - - - - - - - - - - - - - - - -| | callee saves are swapped,
54// | prev_lr | | (frame record first)
55// | prev_fp | <--'
56// | async context if needed |
57// | (a.k.a. "frame record") |
58// |-----------------------------------| <- fp(=x29)
59// Default SVE stack layout Split SVE objects
60// (aarch64-split-sve-objects=false) (aarch64-split-sve-objects=true)
61// |-----------------------------------| |-----------------------------------|
62// | <hazard padding> | | callee-saved PPR registers |
63// |-----------------------------------| |-----------------------------------|
64// | | | PPR stack objects |
65// | callee-saved fp/simd/SVE regs | |-----------------------------------|
66// | | | <hazard padding> |
67// |-----------------------------------| |-----------------------------------|
68// | | | callee-saved ZPR/FPR registers |
69// | SVE stack objects | |-----------------------------------|
70// | | | ZPR stack objects |
71// |-----------------------------------| |-----------------------------------|
72// ^ NB: FPR CSRs are promoted to ZPRs
73// |-----------------------------------|
74// |.empty.space.to.make.part.below....|
75// |.aligned.in.case.it.needs.more.than| (size of this area is unknown at
76// |.the.standard.16-byte.alignment....| compile time; if present)
77// |-----------------------------------|
78// | local variables of fixed size |
79// | including spill slots |
80// | <FPR> |
81// | <hazard padding> |
82// | <GPR> |
83// |-----------------------------------| <- bp(not defined by ABI,
84// |.variable-sized.local.variables....| LLVM chooses X19)
85// |.(VLAs)............................| (size of this area is unknown at
86// |...................................| compile time)
87// |-----------------------------------| <- sp
88// | | Lower address
89//
90//
91// To access the data in a frame, at-compile time, a constant offset must be
92// computable from one of the pointers (fp, bp, sp) to access it. The size
93// of the areas with a dotted background cannot be computed at compile-time
94// if they are present, making it required to have all three of fp, bp and
95// sp to be set up to be able to access all contents in the frame areas,
96// assuming all of the frame areas are non-empty.
97//
98// For most functions, some of the frame areas are empty. For those functions,
99// it may not be necessary to set up fp or bp:
100// * A base pointer is definitely needed when there are both VLAs and local
101// variables with more-than-default alignment requirements.
102// * A frame pointer is definitely needed when there are local variables with
103// more-than-default alignment requirements.
104//
105// For Darwin platforms the frame-record (fp, lr) is stored at the top of the
106// callee-saved area, since the unwind encoding does not allow for encoding
107// this dynamically and existing tools depend on this layout. For other
108// platforms, the frame-record is stored at the bottom of the (gpr) callee-saved
109// area to allow SVE stack objects (allocated directly below the callee-saves,
110// if available) to be accessed directly from the framepointer.
111// The SVE spill/fill instructions have VL-scaled addressing modes such
112// as:
113// ldr z8, [fp, #-7 mul vl]
114// For SVE the size of the vector length (VL) is not known at compile-time, so
115// '#-7 mul vl' is an offset that can only be evaluated at runtime. With this
116// layout, we don't need to add an unscaled offset to the framepointer before
117// accessing the SVE object in the frame.
118//
119// In some cases when a base pointer is not strictly needed, it is generated
120// anyway when offsets from the frame pointer to access local variables become
121// so large that the offset can't be encoded in the immediate fields of loads
122// or stores.
123//
124// Outgoing function arguments must be at the bottom of the stack frame when
125// calling another function. If we do not have variable-sized stack objects, we
126// can allocate a "reserved call frame" area at the bottom of the local
127// variable area, large enough for all outgoing calls. If we do have VLAs, then
128// the stack pointer must be decremented and incremented around each call to
129// make space for the arguments below the VLAs.
130//
131// FIXME: also explain the redzone concept.
132//
133// About stack hazards: Under some SME contexts, a coprocessor with its own
134// separate cache can used for FP operations. This can create hazards if the CPU
135// and the SME unit try to access the same area of memory, including if the
136// access is to an area of the stack. To try to alleviate this we attempt to
137// introduce extra padding into the stack frame between FP and GPR accesses,
138// controlled by the aarch64-stack-hazard-size option. Without changing the
139// layout of the stack frame in the diagram above, a stack object of size
140// aarch64-stack-hazard-size is added between GPR and FPR CSRs. Another is added
141// to the stack objects section, and stack objects are sorted so that FPR >
142// Hazard padding slot > GPRs (where possible). Unfortunately some things are
143// not handled well (VLA area, arguments on the stack, objects with both GPR and
144// FPR accesses), but if those are controlled by the user then the entire stack
145// frame becomes GPR at the start/end with FPR in the middle, surrounded by
146// Hazard padding.
147//
148// An example of the prologue:
149//
150// .globl __foo
151// .align 2
152// __foo:
153// Ltmp0:
154// .cfi_startproc
155// .cfi_personality 155, ___gxx_personality_v0
156// Leh_func_begin:
157// .cfi_lsda 16, Lexception33
158//
159// stp xa,bx, [sp, -#offset]!
160// ...
161// stp x28, x27, [sp, #offset-32]
162// stp fp, lr, [sp, #offset-16]
163// add fp, sp, #offset - 16
164// sub sp, sp, #1360
165//
166// The Stack:
167// +-------------------------------------------+
168// 10000 | ........ | ........ | ........ | ........ |
169// 10004 | ........ | ........ | ........ | ........ |
170// +-------------------------------------------+
171// 10008 | ........ | ........ | ........ | ........ |
172// 1000c | ........ | ........ | ........ | ........ |
173// +===========================================+
174// 10010 | X28 Register |
175// 10014 | X28 Register |
176// +-------------------------------------------+
177// 10018 | X27 Register |
178// 1001c | X27 Register |
179// +===========================================+
180// 10020 | Frame Pointer |
181// 10024 | Frame Pointer |
182// +-------------------------------------------+
183// 10028 | Link Register |
184// 1002c | Link Register |
185// +===========================================+
186// 10030 | ........ | ........ | ........ | ........ |
187// 10034 | ........ | ........ | ........ | ........ |
188// +-------------------------------------------+
189// 10038 | ........ | ........ | ........ | ........ |
190// 1003c | ........ | ........ | ........ | ........ |
191// +-------------------------------------------+
192//
193// [sp] = 10030 :: >>initial value<<
194// sp = 10020 :: stp fp, lr, [sp, #-16]!
195// fp = sp == 10020 :: mov fp, sp
196// [sp] == 10020 :: stp x28, x27, [sp, #-16]!
197// sp == 10010 :: >>final value<<
198//
199// The frame pointer (w29) points to address 10020. If we use an offset of
200// '16' from 'w29', we get the CFI offsets of -8 for w30, -16 for w29, -24
201// for w27, and -32 for w28:
202//
203// Ltmp1:
204// .cfi_def_cfa w29, 16
205// Ltmp2:
206// .cfi_offset w30, -8
207// Ltmp3:
208// .cfi_offset w29, -16
209// Ltmp4:
210// .cfi_offset w27, -24
211// Ltmp5:
212// .cfi_offset w28, -32
213//
214//===----------------------------------------------------------------------===//
215
216#include "AArch64FrameLowering.h"
217#include "AArch64InstrInfo.h"
220#include "AArch64RegisterInfo.h"
221#include "AArch64SMEAttributes.h"
222#include "AArch64Subtarget.h"
225#include "llvm/ADT/ScopeExit.h"
226#include "llvm/ADT/SmallVector.h"
244#include "llvm/IR/Attributes.h"
245#include "llvm/IR/CallingConv.h"
246#include "llvm/IR/DataLayout.h"
247#include "llvm/IR/DebugLoc.h"
248#include "llvm/IR/Function.h"
249#include "llvm/IR/Module.h"
250#include "llvm/MC/MCAsmInfo.h"
251#include "llvm/MC/MCDwarf.h"
252#include "llvm/Support/Debug.h"
258#include <cassert>
259#include <cstdint>
260#include <iterator>
261#include <optional>
262#include <vector>
263
264using namespace llvm;
265
266#define DEBUG_TYPE "frame-info"
267
268int64_t
270 MachineBasicBlock &MBB) const {
271 MachineBasicBlock::iterator MBBI = MBB.getLastNonDebugInstr();
273 bool IsTailCallReturn = (MBB.end() != MBBI)
275 : false;
276
277 int64_t ArgumentPopSize = 0;
278 if (IsTailCallReturn) {
279 MachineOperand &StackAdjust = MBBI->getOperand(1);
280
281 // For a tail-call in a callee-pops-arguments environment, some or all of
282 // the stack may actually be in use for the call's arguments, this is
283 // calculated during LowerCall and consumed here...
284 ArgumentPopSize = StackAdjust.getImm();
285 } else {
286 // ... otherwise the amount to pop is *all* of the argument space,
287 // conveniently stored in the MachineFunctionInfo by
288 // LowerFormalArguments. This will, of course, be zero for the C calling
289 // convention.
290 ArgumentPopSize = AFI->getArgumentStackToRestore();
291 }
292
293 return ArgumentPopSize;
294}
295
297 MachineFunction &MF);
298
299enum class AssignObjectOffsets { No, Yes };
300/// Process all the SVE stack objects and the SVE stack size and offsets for
301/// each object. If AssignOffsets is "Yes", the offsets get assigned (and SVE
302/// stack sizes set). Returns the size of the SVE stack.
304 AssignObjectOffsets AssignOffsets);
305
306static unsigned getStackHazardSize(const MachineFunction &MF) {
307 return MF.getSubtarget<AArch64Subtarget>().getStreamingHazardSize();
308}
309
315
318 // With split SVE objects, the hazard padding is added to the PPR region,
319 // which places it between the [GPR, PPR] area and the [ZPR, FPR] area. This
320 // avoids hazards between both GPRs and FPRs and ZPRs and PPRs.
323 : 0,
324 AFI->getStackSizePPR());
325}
326
327// Conservatively, returns true if the function is likely to have SVE vectors
328// on the stack. This function is safe to be called before callee-saves or
329// object offsets have been determined.
331 const MachineFunction &MF) {
332 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
333 if (AFI->isSVECC())
334 return true;
335
336 if (AFI->hasCalculatedStackSizeSVE())
337 return bool(AFL.getSVEStackSize(MF));
338
339 const MachineFrameInfo &MFI = MF.getFrameInfo();
340 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd(); FI++) {
341 if (MFI.hasScalableStackID(FI))
342 return true;
343 }
344
345 return false;
346}
347
348static bool isTargetWindows(const MachineFunction &MF) {
349 // TODO: Should this include targets like UEFI (which use Windows CFI)?
350 // Note: Currently, there is not AArch64 support for UEFI. The value returned
351 // here must align with the predicate used for returning the list of callee
352 // saved regs in AArch64RegisterInfo::getCalleeSavedRegs(), so that we use
353 // invalidateWindowsRegisterPairing() where appropriate.
355}
356
358 const MachineFunction &MF) const {
359 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
360 return isTargetWindows(MF) && AFI->getSVECalleeSavedStackSize();
361}
362
363/// Returns true if a homogeneous prolog or epilog code can be emitted
364/// for the size optimization. If possible, a frame helper call is injected.
365/// When Exit block is given, this check is for epilog.
366bool AArch64FrameLowering::homogeneousPrologEpilog(
367 MachineFunction &MF, MachineBasicBlock *Exit) const {
368 if (!MF.getFunction().hasMinSize())
369 return false;
370 const AArch64Options &CLOpts =
371 MF.getSubtarget<AArch64Subtarget>().getCLOpts();
372 if (!CLOpts.homogeneous_prolog_epilog)
373 return false;
374 if (CLOpts.redzone)
375 return false;
376
377 // TODO: Window is supported yet.
378 if (isTargetWindows(MF))
379 return false;
380
381 // TODO: SVE is not supported yet.
382 if (isLikelyToHaveSVEStack(*this, MF))
383 return false;
384
385 // Bail on stack adjustment needed on return for simplicity.
386 const MachineFrameInfo &MFI = MF.getFrameInfo();
387 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
388 if (MFI.hasVarSizedObjects() || RegInfo->hasStackRealignment(MF))
389 return false;
390 if (Exit && getArgumentStackToRestore(MF, *Exit))
391 return false;
392
393 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
394 if (AFI->hasSwiftAsyncContext() || AFI->hasStreamingModeChanges())
395 return false;
396
397 // If there are an odd number of GPRs before LR and FP in the CSRs list,
398 // they will not be paired into one RegPairInfo, which is incompatible with
399 // the assumption made by the homogeneous prolog epilog pass.
400 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
401 unsigned NumGPRs = 0;
402 for (unsigned I = 0; CSRegs[I]; ++I) {
403 Register Reg = CSRegs[I];
404 if (Reg == AArch64::LR) {
405 assert(CSRegs[I + 1] == AArch64::FP);
406 if (NumGPRs % 2 != 0)
407 return false;
408 break;
409 }
410 if (AArch64::GPR64RegClass.contains(Reg))
411 ++NumGPRs;
412 }
413
414 return true;
415}
416
417/// Returns true if CSRs should be paired.
418bool AArch64FrameLowering::producePairRegisters(MachineFunction &MF) const {
419 return produceCompactUnwindFrame(*this, MF) || homogeneousPrologEpilog(MF);
420}
421
422/// This is the biggest offset to the stack pointer we can encode in aarch64
423/// instructions (without using a separate calculation and a temp register).
424/// Note that the exception here are vector stores/loads which cannot encode any
425/// displacements (see estimateRSStackSizeLimit(), isAArch64FrameOffsetLegal()).
426static const unsigned DefaultSafeSPDisplacement = 255;
427
428/// Look at each instruction that references stack frames and return the stack
429/// size limit beyond which some of these instructions will require a scratch
430/// register during their expansion later.
432 // FIXME: For now, just conservatively guesstimate based on unscaled indexing
433 // range. We'll end up allocating an unnecessary spill slot a lot, but
434 // realistically that's not a big deal at this stage of the game.
435 for (MachineBasicBlock &MBB : MF) {
436 for (MachineInstr &MI : MBB) {
437 if (MI.isDebugInstr() || MI.isPseudo() ||
438 MI.getOpcode() == AArch64::ADDXri ||
439 MI.getOpcode() == AArch64::ADDSXri)
440 continue;
441
442 for (const MachineOperand &MO : MI.operands()) {
443 if (!MO.isFI())
444 continue;
445
447 if (isAArch64FrameOffsetLegal(MI, Offset, nullptr, nullptr, nullptr) ==
449 return 0;
450 }
451 }
452 }
454}
455
460
461unsigned
462AArch64FrameLowering::getFixedObjectSize(const MachineFunction &MF,
463 const AArch64FunctionInfo *AFI,
464 bool IsWin64, bool IsFunclet) const {
465 assert(AFI->getTailCallReservedStack() % 16 == 0 &&
466 "Tail call reserved stack must be aligned to 16 bytes");
467 if (!IsWin64 || IsFunclet) {
468 return AFI->getTailCallReservedStack();
469 } else {
470 if (AFI->getTailCallReservedStack() != 0 &&
471 !MF.getFunction().getAttributes().hasAttrSomewhere(
472 Attribute::SwiftAsync))
473 report_fatal_error("cannot generate ABI-changing tail call for Win64");
474 unsigned FixedObjectSize = AFI->getTailCallReservedStack();
475
476 // Var args are stored here in the primary function.
477 FixedObjectSize += AFI->getVarArgsGPRSize();
478
479 if (MF.hasEHFunclets()) {
480 // Catch objects are stored here in the primary function.
481 const MachineFrameInfo &MFI = MF.getFrameInfo();
482 const WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
483 SmallSetVector<int, 8> CatchObjFrameIndices;
484 for (const WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
485 for (const WinEHHandlerType &H : TBME.HandlerArray) {
486 int FrameIndex = H.CatchObj.FrameIndex;
487 if ((FrameIndex != INT_MAX) &&
488 CatchObjFrameIndices.insert(FrameIndex)) {
489 FixedObjectSize = alignTo(FixedObjectSize,
490 MFI.getObjectAlign(FrameIndex).value()) +
491 MFI.getObjectSize(FrameIndex);
492 }
493 }
494 }
495 // To support EH funclets we allocate an UnwindHelp object
496 FixedObjectSize += 8;
497 }
498 return alignTo(FixedObjectSize, 16);
499 }
500}
501
503 if (!MF.getSubtarget<AArch64Subtarget>().getCLOpts().redzone)
504 return false;
505
506 // Don't use the red zone if the function explicitly asks us not to.
507 // This is typically used for kernel code.
508 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
509 const unsigned RedZoneSize =
511 if (!RedZoneSize)
512 return false;
513
514 const MachineFrameInfo &MFI = MF.getFrameInfo();
516 uint64_t NumBytes = AFI->getLocalStackSize();
517
518 // If neither NEON or SVE are available, a COPY from one Q-reg to
519 // another requires a spill -> reload sequence. We can do that
520 // using a pre-decrementing store/post-decrementing load, but
521 // if we do so, we can't use the Red Zone.
522 bool LowerQRegCopyThroughMem = Subtarget.hasFPARMv8() &&
523 !Subtarget.isNeonAvailable() &&
524 !Subtarget.hasSVE();
525
526 return !(MFI.hasCalls() || hasFP(MF) || NumBytes > RedZoneSize ||
527 AFI->hasSVEStackSize() || LowerQRegCopyThroughMem);
528}
529
530/// hasFPImpl - Return true if the specified function should have a dedicated
531/// frame pointer register.
533 const MachineFrameInfo &MFI = MF.getFrameInfo();
534 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
536
537 // Win64 EH requires a frame pointer if funclets are present, as the locals
538 // are accessed off the frame pointer in both the parent function and the
539 // funclets.
540 if (MF.hasEHFunclets())
541 return true;
542
543 // When the stack guard is mixed with the frame pointer, a dedicated FP is
544 // required so the guard value remains stable in the presence of dynamic
545 // stack allocations (e.g. _alloca on MSVCRT).
546 if (MFI.hasStackProtectorIndex()) {
547 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
548 if (Subtarget.getTargetLowering()->useStackGuardMixFP())
549 return true;
550 }
551
552 // Retain behavior of always omitting the FP for leaf functions when possible.
554 return true;
555 if (MFI.hasVarSizedObjects() || MFI.isFrameAddressTaken() ||
556 MFI.hasStackMap() || MFI.hasPatchPoint() ||
557 RegInfo->hasStackRealignment(MF))
558 return true;
559
560 // If we:
561 //
562 // 1. Have streaming mode changes
563 // OR:
564 // 2. Have a streaming body with SVE stack objects
565 //
566 // Then the value of VG restored when unwinding to this function may not match
567 // the value of VG used to set up the stack.
568 //
569 // This is a problem as the CFA can be described with an expression of the
570 // form: CFA = SP + NumBytes + VG * NumScalableBytes.
571 //
572 // If the value of VG used in that expression does not match the value used to
573 // set up the stack, an incorrect address for the CFA will be computed, and
574 // unwinding will fail.
575 //
576 // We work around this issue by ensuring the frame-pointer can describe the
577 // CFA in either of these cases.
578 if (AFI.needsDwarfUnwindInfo(MF) &&
581 return true;
582 // With large callframes around we may need to use FP to access the scavenging
583 // emergency spillslot.
584 //
585 // Unfortunately some calls to hasFP() like machine verifier ->
586 // getReservedReg() -> hasFP in the middle of global isel are too early
587 // to know the max call frame size. Hopefully conservatively returning "true"
588 // in those cases is fine.
589 // DefaultSafeSPDisplacement is fine as we only emergency spill GP regs.
590 if (!MFI.isMaxCallFrameSizeComputed() ||
592 return true;
593
594 return false;
595}
596
597/// Should the Frame Pointer be reserved for the current function?
599 const Triple &TT = MF.getFunction().getParent()->getTargetTriple();
600
601 // These OSes require the frame chain is valid, even if the current frame does
602 // not use a frame pointer.
603 if (TT.isOSDarwin() || TT.isOSWindows())
604 return true;
605
606 // If the function has a frame pointer, it is reserved.
607 if (hasFP(MF))
608 return true;
609
610 // Frontend has requested to preserve the frame pointer.
611 if (MF.framePointerIsReserved())
612 return true;
613
614 return false;
615}
616
617/// hasReservedCallFrame - Under normal circumstances, when a frame pointer is
618/// not required, we reserve argument space for call sites in the function
619/// immediately on entry to the current function. This eliminates the need for
620/// add/sub sp brackets around call sites. Returns true if the call frame is
621/// included as part of the stack frame.
623 const MachineFunction &MF) const {
624 // The stack probing code for the dynamically allocated outgoing arguments
625 // area assumes that the stack is probed at the top - either by the prologue
626 // code, which issues a probe if `hasVarSizedObjects` return true, or by the
627 // most recent variable-sized object allocation. Changing the condition here
628 // may need to be followed up by changes to the probe issuing logic.
629 return !MF.getFrameInfo().hasVarSizedObjects();
630}
631
635
636 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
637 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
638 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
639 [[maybe_unused]] MachineFrameInfo &MFI = MF.getFrameInfo();
640 DebugLoc DL = I->getDebugLoc();
641 unsigned Opc = I->getOpcode();
642 bool IsDestroy = Opc == TII->getCallFrameDestroyOpcode();
643 uint64_t CalleePopAmount = IsDestroy ? I->getOperand(1).getImm() : 0;
644
645 if (!hasReservedCallFrame(MF)) {
646 int64_t Amount = I->getOperand(0).getImm();
647 Amount = alignTo(Amount, getStackAlign());
648 if (!IsDestroy)
649 Amount = -Amount;
650
651 // N.b. if CalleePopAmount is valid but zero (i.e. callee would pop, but it
652 // doesn't have to pop anything), then the first operand will be zero too so
653 // this adjustment is a no-op.
654 if (CalleePopAmount == 0) {
655 // FIXME: in-function stack adjustment for calls is limited to 24-bits
656 // because there's no guaranteed temporary register available.
657 //
658 // ADD/SUB (immediate) has only LSL #0 and LSL #12 available.
659 // 1) For offset <= 12-bit, we use LSL #0
660 // 2) For 12-bit <= offset <= 24-bit, we use two instructions. One uses
661 // LSL #0, and the other uses LSL #12.
662 //
663 // Most call frames will be allocated at the start of a function so
664 // this is OK, but it is a limitation that needs dealing with.
665 assert(Amount > -0xffffff && Amount < 0xffffff && "call frame too large");
666
667 if (TLI->hasInlineStackProbe(MF) &&
669 // When stack probing is enabled, the decrement of SP may need to be
670 // probed. We only need to do this if the call site needs 1024 bytes of
671 // space or more, because a region smaller than that is allowed to be
672 // unprobed at an ABI boundary. We rely on the fact that SP has been
673 // probed exactly at this point, either by the prologue or most recent
674 // dynamic allocation.
676 "non-reserved call frame without var sized objects?");
677 Register ScratchReg =
678 MF.getRegInfo().createVirtualRegister(&AArch64::GPR64RegClass);
679 inlineStackProbeFixed(I, ScratchReg, -Amount, StackOffset::get(0, 0));
680 } else {
681 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
682 StackOffset::getFixed(Amount), TII);
683 }
684 }
685 } else if (CalleePopAmount != 0) {
686 // If the calling convention demands that the callee pops arguments from the
687 // stack, we want to add it back if we have a reserved call frame.
688 assert(CalleePopAmount < 0xffffff && "call frame too large");
689 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
690 StackOffset::getFixed(-(int64_t)CalleePopAmount), TII);
691 }
692 return MBB.erase(I);
693}
694
696 MachineBasicBlock &MBB) const {
697
698 MachineFunction &MF = *MBB.getParent();
699 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
700 const auto &TRI = *Subtarget.getRegisterInfo();
701 const auto &MFI = *MF.getInfo<AArch64FunctionInfo>();
702
703 CFIInstBuilder CFIBuilder(MBB, MBB.begin(), MachineInstr::NoFlags);
704
705 // Reset the CFA to `SP + 0`.
706 CFIBuilder.buildDefCFA(AArch64::SP, 0);
707
708 // Flip the RA sign state.
709 if (MFI.shouldSignReturnAddress(MF)) {
710 if (MFI.branchProtectionPAuthLR()) {
711 CFIBuilder.buildNegateRAStateWithPC();
712 } else if (!MF.getTarget().getTargetTriple().isOSBinFormatMachO()) {
713 CFIBuilder.buildNegateRAState();
714 }
715 }
716
717 // Shadow call stack uses X18, reset it.
718 if (MFI.needsShadowCallStackPrologueEpilogue(MF))
719 CFIBuilder.buildSameValue(AArch64::X18);
720
721 // Emit .cfi_same_value for callee-saved registers.
722 const std::vector<CalleeSavedInfo> &CSI =
724 for (const auto &Info : CSI) {
725 MCRegister Reg = Info.getReg();
726 if (!TRI.regNeedsCFI(Reg, Reg))
727 continue;
728 CFIBuilder.buildSameValue(Reg);
729 }
730}
731
733 switch (Reg.id()) {
734 default:
735 // The called routine is expected to preserve r19-r28
736 // r29 and r30 are used as frame pointer and link register resp.
737 return 0;
738
739 // GPRs
740#define CASE(n) \
741 case AArch64::W##n: \
742 case AArch64::X##n: \
743 return AArch64::X##n
744 CASE(0);
745 CASE(1);
746 CASE(2);
747 CASE(3);
748 CASE(4);
749 CASE(5);
750 CASE(6);
751 CASE(7);
752 CASE(8);
753 CASE(9);
754 CASE(10);
755 CASE(11);
756 CASE(12);
757 CASE(13);
758 CASE(14);
759 CASE(15);
760 CASE(16);
761 CASE(17);
762 CASE(18);
763#undef CASE
764
765 // FPRs
766#define CASE(n) \
767 case AArch64::B##n: \
768 case AArch64::H##n: \
769 case AArch64::S##n: \
770 case AArch64::D##n: \
771 case AArch64::Q##n: \
772 return HasSVE ? AArch64::Z##n : AArch64::Q##n
773 CASE(0);
774 CASE(1);
775 CASE(2);
776 CASE(3);
777 CASE(4);
778 CASE(5);
779 CASE(6);
780 CASE(7);
781 CASE(8);
782 CASE(9);
783 CASE(10);
784 CASE(11);
785 CASE(12);
786 CASE(13);
787 CASE(14);
788 CASE(15);
789 CASE(16);
790 CASE(17);
791 CASE(18);
792 CASE(19);
793 CASE(20);
794 CASE(21);
795 CASE(22);
796 CASE(23);
797 CASE(24);
798 CASE(25);
799 CASE(26);
800 CASE(27);
801 CASE(28);
802 CASE(29);
803 CASE(30);
804 CASE(31);
805#undef CASE
806 }
807}
808
809void AArch64FrameLowering::emitZeroCallUsedRegs(BitVector RegsToZero,
811 RegScavenger *) const {
812 // Insertion point.
814
815 // Fake a debug loc.
816 DebugLoc DL;
817 if (MBBI != MBB.end())
818 DL = MBBI->getDebugLoc();
819
820 const MachineFunction &MF = *MBB.getParent();
821 const AArch64Subtarget &STI = MF.getSubtarget<AArch64Subtarget>();
822 const AArch64RegisterInfo &TRI = *STI.getRegisterInfo();
823
824 BitVector GPRsToZero(TRI.getNumRegs());
825 BitVector FPRsToZero(TRI.getNumRegs());
826 bool HasSVE = STI.isSVEorStreamingSVEAvailable();
827 // Without an FP unit (e.g. -mgeneral-regs-only) the FP/vector registers can't
828 // hold a value and there is no instruction to clear them, so leave them out.
829 bool HasFPR = STI.hasFPARMv8();
830 for (MCRegister Reg : RegsToZero.set_bits()) {
831 if (TRI.isGeneralPurposeRegister(MF, Reg)) {
832 // For GPRs, we only care to clear out the 64-bit register.
833 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
834 GPRsToZero.set(XReg);
835 } else if (HasFPR && AArch64InstrInfo::isFpOrNEON(Reg)) {
836 // For FPRs,
837 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
838 FPRsToZero.set(XReg);
839 }
840 }
841
842 const AArch64InstrInfo &TII = *STI.getInstrInfo();
843
844 // Zero out GPRs.
845 for (MCRegister Reg : GPRsToZero.set_bits())
846 TII.buildClearRegister(Reg, MBB, MBBI, DL);
847
848 // Zero out FP/vector registers.
849 for (MCRegister Reg : FPRsToZero.set_bits())
850 TII.buildClearRegister(Reg, MBB, MBBI, DL);
851
852 if (HasSVE) {
853 for (MCRegister PReg :
854 {AArch64::P0, AArch64::P1, AArch64::P2, AArch64::P3, AArch64::P4,
855 AArch64::P5, AArch64::P6, AArch64::P7, AArch64::P8, AArch64::P9,
856 AArch64::P10, AArch64::P11, AArch64::P12, AArch64::P13, AArch64::P14,
857 AArch64::P15}) {
858 if (RegsToZero[PReg])
859 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PFALSE), PReg);
860 }
861 }
862}
863
864bool AArch64FrameLowering::windowsRequiresStackProbe(
865 const MachineFunction &MF, uint64_t StackSizeInBytes) const {
866 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
867 const AArch64FunctionInfo &MFI = *MF.getInfo<AArch64FunctionInfo>();
868 // TODO: When implementing stack protectors, take that into account
869 // for the probe threshold.
870 return Subtarget.isTargetWindows() && MFI.hasStackProbing() &&
871 StackSizeInBytes >= uint64_t(MFI.getStackProbeSize());
872}
873
875 const MachineBasicBlock &MBB) {
876 const MachineFunction *MF = MBB.getParent();
877 LiveRegs.addLiveIns(MBB);
878 // Mark callee saved registers as used so we will not choose them.
879 const MCPhysReg *CSRegs = MF->getRegInfo().getCalleeSavedRegs();
880 for (unsigned i = 0; CSRegs[i]; ++i)
881 LiveRegs.addReg(CSRegs[i]);
882}
883
885AArch64FrameLowering::findScratchNonCalleeSaveRegister(MachineBasicBlock *MBB,
886 bool HasCall) const {
888
889 // If MBB is an entry block, use X9 as the scratch register
890 // preserve_none functions may be using X9 to pass arguments,
891 // so prefer to pick an available register below.
892 if (&MF->front() == MBB &&
894 return AArch64::X9;
895
896 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
897 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
898 LivePhysRegs LiveRegs(TRI);
899 getLiveRegsForEntryMBB(LiveRegs, *MBB);
900 if (HasCall) {
901 LiveRegs.addReg(AArch64::X16);
902 LiveRegs.addReg(AArch64::X17);
903 LiveRegs.addReg(AArch64::X18);
904 }
905
906 // Prefer X9 since it was historically used for the prologue scratch reg.
907 const MachineRegisterInfo &MRI = MF->getRegInfo();
908 if (LiveRegs.available(MRI, AArch64::X9))
909 return AArch64::X9;
910
911 for (unsigned Reg : AArch64::GPR64RegClass) {
912 if (LiveRegs.available(MRI, Reg))
913 return Reg;
914 }
915 return Register();
916}
917
919 const MachineBasicBlock &MBB) const {
920 const MachineFunction *MF = MBB.getParent();
921 MachineBasicBlock *TmpMBB = const_cast<MachineBasicBlock *>(&MBB);
922 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
923 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
924 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
926
927 if (AFI->hasSwiftAsyncContext()) {
928 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
929 const MachineRegisterInfo &MRI = MF->getRegInfo();
932 // The StoreSwiftAsyncContext clobbers X16 and X17. Make sure they are
933 // available.
934 if (!LiveRegs.available(MRI, AArch64::X16) ||
935 !LiveRegs.available(MRI, AArch64::X17))
936 return false;
937 }
938
939 // Certain stack probing sequences might clobber flags, then we can't use
940 // the block as a prologue if the flags register is a live-in.
942 MBB.isLiveIn(AArch64::NZCV))
943 return false;
944
945 if (RegInfo->hasStackRealignment(*MF) || TLI->hasInlineStackProbe(*MF))
946 if (!findScratchNonCalleeSaveRegister(TmpMBB).isValid())
947 return false;
948
949 // May need a scratch register (for return value) if require making a special
950 // call
951 if (requiresSaveVG(*MF) ||
952 windowsRequiresStackProbe(*MF, std::numeric_limits<uint64_t>::max()))
953 if (!findScratchNonCalleeSaveRegister(TmpMBB, true).isValid())
954 return false;
955
956 return true;
957}
958
960 const Function &F = MF.getFunction();
961 return MF.getTarget().getMCAsmInfo().usesWindowsCFI() &&
962 F.needsUnwindTableEntry();
963}
964
965bool AArch64FrameLowering::shouldSignReturnAddressEverywhere(
966 const MachineFunction &MF) const {
967 // FIXME: With WinCFI, extra care should be taken to place SEH_PACSignLR
968 // and SEH_EpilogEnd instructions in the correct order.
970 return false;
973}
974
975// Given a load or a store instruction, generate an appropriate unwinding SEH
976// code on Windows.
978AArch64FrameLowering::insertSEH(MachineBasicBlock::iterator MBBI,
979 const AArch64InstrInfo &TII,
980 MachineInstr::MIFlag Flag) const {
981 unsigned Opc = MBBI->getOpcode();
982 MachineBasicBlock *MBB = MBBI->getParent();
983 MachineFunction &MF = *MBB->getParent();
984 DebugLoc DL = MBBI->getDebugLoc();
985 unsigned ImmIdx = MBBI->getNumOperands() - 1;
986 int Imm = MBBI->getOperand(ImmIdx).getImm();
988 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
989 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
990
991 switch (Opc) {
992 default:
993 report_fatal_error("No SEH Opcode for this instruction");
994 case AArch64::STR_ZXI:
995 case AArch64::LDR_ZXI: {
996 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
997 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveZReg))
998 .addImm(Reg0)
999 .addImm(Imm)
1000 .setMIFlag(Flag);
1001 break;
1002 }
1003 case AArch64::STR_PXI:
1004 case AArch64::LDR_PXI: {
1005 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1006 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SavePReg))
1007 .addImm(Reg0)
1008 .addImm(Imm)
1009 .setMIFlag(Flag);
1010 break;
1011 }
1012 case AArch64::LDPDpost:
1013 Imm = -Imm;
1014 [[fallthrough]];
1015 case AArch64::STPDpre: {
1016 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1017 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1018 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP_X))
1019 .addImm(Reg0)
1020 .addImm(Reg1)
1021 .addImm(Imm * 8)
1022 .setMIFlag(Flag);
1023 break;
1024 }
1025 case AArch64::LDPXpost:
1026 Imm = -Imm;
1027 [[fallthrough]];
1028 case AArch64::STPXpre: {
1029 Register Reg0 = MBBI->getOperand(1).getReg();
1030 Register Reg1 = MBBI->getOperand(2).getReg();
1031 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1032 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR_X))
1033 .addImm(Imm * 8)
1034 .setMIFlag(Flag);
1035 else
1036 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP_X))
1037 .addImm(RegInfo->getSEHRegNum(Reg0))
1038 .addImm(RegInfo->getSEHRegNum(Reg1))
1039 .addImm(Imm * 8)
1040 .setMIFlag(Flag);
1041 break;
1042 }
1043 case AArch64::LDRDpost:
1044 Imm = -Imm;
1045 [[fallthrough]];
1046 case AArch64::STRDpre: {
1047 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1048 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg_X))
1049 .addImm(Reg)
1050 .addImm(Imm)
1051 .setMIFlag(Flag);
1052 break;
1053 }
1054 case AArch64::LDRXpost:
1055 Imm = -Imm;
1056 [[fallthrough]];
1057 case AArch64::STRXpre: {
1058 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1059 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg_X))
1060 .addImm(Reg)
1061 .addImm(Imm)
1062 .setMIFlag(Flag);
1063 break;
1064 }
1065 case AArch64::STPDi:
1066 case AArch64::LDPDi: {
1067 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1068 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1069 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP))
1070 .addImm(Reg0)
1071 .addImm(Reg1)
1072 .addImm(Imm * 8)
1073 .setMIFlag(Flag);
1074 break;
1075 }
1076 case AArch64::STPXi:
1077 case AArch64::LDPXi: {
1078 Register Reg0 = MBBI->getOperand(0).getReg();
1079 Register Reg1 = MBBI->getOperand(1).getReg();
1080
1081 int SEHReg0 = RegInfo->getSEHRegNum(Reg0);
1082 int SEHReg1 = RegInfo->getSEHRegNum(Reg1);
1083
1084 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1085 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR))
1086 .addImm(Imm * 8)
1087 .setMIFlag(Flag);
1088 else if (SEHReg0 >= 19 && SEHReg1 >= 19)
1089 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP))
1090 .addImm(SEHReg0)
1091 .addImm(SEHReg1)
1092 .addImm(Imm * 8)
1093 .setMIFlag(Flag);
1094 else
1095 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegIP))
1096 .addImm(SEHReg0)
1097 .addImm(SEHReg1)
1098 .addImm(Imm * 8)
1099 .setMIFlag(Flag);
1100 break;
1101 }
1102 case AArch64::STRXui:
1103 case AArch64::LDRXui: {
1104 int Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1105 if (Reg >= 19)
1106 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg))
1107 .addImm(Reg)
1108 .addImm(Imm * 8)
1109 .setMIFlag(Flag);
1110 else
1111 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegI))
1112 .addImm(Reg)
1113 .addImm(Imm * 8)
1114 .setMIFlag(Flag);
1115 break;
1116 }
1117 case AArch64::STRDui:
1118 case AArch64::LDRDui: {
1119 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1120 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg))
1121 .addImm(Reg)
1122 .addImm(Imm * 8)
1123 .setMIFlag(Flag);
1124 break;
1125 }
1126 case AArch64::STPQi:
1127 case AArch64::LDPQi: {
1128 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1129 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1130 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQP))
1131 .addImm(Reg0)
1132 .addImm(Reg1)
1133 .addImm(Imm * 16)
1134 .setMIFlag(Flag);
1135 break;
1136 }
1137 case AArch64::LDPQpost:
1138 Imm = -Imm;
1139 [[fallthrough]];
1140 case AArch64::STPQpre: {
1141 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1142 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1143 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQPX))
1144 .addImm(Reg0)
1145 .addImm(Reg1)
1146 .addImm(Imm * 16)
1147 .setMIFlag(Flag);
1148 break;
1149 }
1150 }
1151 auto I = MBB->insertAfter(MBBI, MIB);
1152 return I;
1153}
1154
1157 if (!AFI->needsDwarfUnwindInfo(MF) || !AFI->hasStreamingModeChanges())
1158 return false;
1159 // For Darwin platforms we don't save VG for non-SVE functions, even if SME
1160 // is enabled with streaming mode changes.
1161 auto &ST = MF.getSubtarget<AArch64Subtarget>();
1162 if (ST.isTargetDarwin())
1163 return ST.hasSVE();
1164 return true;
1165}
1166
1168 MachineFunction &MF) const {
1169 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1170 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
1171
1172 auto EmitSignRA = [&](MachineBasicBlock &MBB) {
1173 DebugLoc DL; // Set debug location to unknown.
1175
1176 BuildMI(MBB, MBBI, DL, TII->get(AArch64::PAUTH_PROLOGUE))
1178 };
1179
1180 auto EmitAuthRA = [&](MachineBasicBlock &MBB) {
1181 DebugLoc DL;
1182 MachineBasicBlock::iterator MBBI = MBB.getFirstTerminator();
1183 if (MBBI != MBB.end())
1184 DL = MBBI->getDebugLoc();
1185
1186 TII->createPauthEpilogueInstr(MBB, DL);
1187 };
1188
1189 // This should be in sync with PEIImpl::calculateSaveRestoreBlocks.
1190 EmitSignRA(MF.front());
1191 for (MachineBasicBlock &MBB : MF) {
1192 if (MBB.isEHFuncletEntry())
1193 EmitSignRA(MBB);
1194 if (MBB.isReturnBlock())
1195 EmitAuthRA(MBB);
1196 }
1197}
1198
1200 MachineBasicBlock &MBB) const {
1201 AArch64PrologueEmitter PrologueEmitter(MF, MBB, *this);
1202 PrologueEmitter.emitPrologue();
1203}
1204
1206 MachineBasicBlock &MBB) const {
1207 AArch64EpilogueEmitter EpilogueEmitter(MF, MBB, *this);
1208 EpilogueEmitter.emitEpilogue();
1209}
1210
1213 MF.getInfo<AArch64FunctionInfo>()->needsDwarfUnwindInfo(MF);
1214}
1215
1217 return enableCFIFixup(MF) &&
1218 MF.getInfo<AArch64FunctionInfo>()->needsAsyncDwarfUnwindInfo(MF);
1219}
1220
1221/// getFrameIndexReference - Provide a base+offset reference to an FI slot for
1222/// debug info. It's the same as what we use for resolving the code-gen
1223/// references for now. FIXME: This can go wrong when references are
1224/// SP-relative and simple call frames aren't used.
1227 Register &FrameReg) const {
1229 MF, FI, FrameReg,
1230 /*PreferFP=*/
1231 MF.getFunction().hasFnAttribute(Attribute::SanitizeHWAddress) ||
1232 MF.getFunction().hasFnAttribute(Attribute::SanitizeMemTag),
1233 /*ForSimm=*/false);
1234}
1235
1238 int FI) const {
1239 // This function serves to provide a comparable offset from a single reference
1240 // point (the value of SP at function entry) that can be used for analysis,
1241 // e.g. the stack-frame-layout analysis pass. It is not guaranteed to be
1242 // correct for all objects in the presence of VLA-area objects or dynamic
1243 // stack re-alignment.
1244
1245 const auto &MFI = MF.getFrameInfo();
1246
1247 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1248 StackOffset ZPRStackSize = getZPRStackSize(MF);
1249 StackOffset PPRStackSize = getPPRStackSize(MF);
1250 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1251
1252 // For VLA-area objects, just emit an offset at the end of the stack frame.
1253 // Whilst not quite correct, these objects do live at the end of the frame and
1254 // so it is more useful for analysis for the offset to reflect this.
1255 if (MFI.isVariableSizedObjectIndex(FI)) {
1256 return StackOffset::getFixed(-((int64_t)MFI.getStackSize())) - SVEStackSize;
1257 }
1258
1259 // This is correct in the absence of any SVE stack objects.
1260 if (!SVEStackSize)
1261 return StackOffset::getFixed(ObjectOffset - getOffsetOfLocalArea());
1262
1263 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1264 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1265 if (MFI.hasScalableStackID(FI)) {
1266 if (FPAfterSVECalleeSaves &&
1267 -ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1268 assert(!AFI->hasSplitSVEObjects() &&
1269 "split-sve-objects not supported with FPAfterSVECalleeSaves");
1270 return StackOffset::getScalable(ObjectOffset);
1271 }
1272 StackOffset AccessOffset{};
1273 // The scalable vectors are below (lower address) the scalable predicates
1274 // with split SVE objects, so we must subtract the size of the predicates.
1275 if (AFI->hasSplitSVEObjects() &&
1276 MFI.getStackID(FI) == TargetStackID::ScalableVector)
1277 AccessOffset = -PPRStackSize;
1278 return AccessOffset +
1279 StackOffset::get(-((int64_t)AFI->getCalleeSavedStackSize()),
1280 ObjectOffset);
1281 }
1282
1283 bool IsFixed = MFI.isFixedObjectIndex(FI);
1284 bool IsCSR =
1285 !IsFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1286
1287 StackOffset ScalableOffset = {};
1288 if (!IsFixed && !IsCSR) {
1289 ScalableOffset = -SVEStackSize;
1290 } else if (FPAfterSVECalleeSaves && IsCSR) {
1291 ScalableOffset =
1293 }
1294
1295 return StackOffset::getFixed(ObjectOffset) + ScalableOffset;
1296}
1297
1303
1304StackOffset AArch64FrameLowering::getFPOffset(const MachineFunction &MF,
1305 int64_t ObjectOffset) const {
1306 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1307 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1308 const Function &F = MF.getFunction();
1309 bool IsWin64 = Subtarget.isCallingConvWin64(F.getCallingConv(), F.isVarArg());
1310 unsigned FixedObject =
1311 getFixedObjectSize(MF, AFI, IsWin64, /*IsFunclet=*/false);
1312 int64_t CalleeSaveSize = AFI->getCalleeSavedStackSize(MF.getFrameInfo());
1313 int64_t FPAdjust =
1314 CalleeSaveSize - AFI->getCalleeSaveBaseToFrameRecordOffset();
1315 return StackOffset::getFixed(ObjectOffset + FixedObject + FPAdjust);
1316}
1317
1318StackOffset AArch64FrameLowering::getStackOffset(const MachineFunction &MF,
1319 int64_t ObjectOffset) const {
1320 const auto &MFI = MF.getFrameInfo();
1321 return StackOffset::getFixed(ObjectOffset + (int64_t)MFI.getStackSize());
1322}
1323
1324// TODO: This function currently does not work for scalable vectors.
1326 int FI) const {
1327 const AArch64RegisterInfo *RegInfo =
1328 MF.getSubtarget<AArch64Subtarget>().getRegisterInfo();
1329 int ObjectOffset = MF.getFrameInfo().getObjectOffset(FI);
1330 return RegInfo->getLocalAddressRegister(MF) == AArch64::FP
1331 ? getFPOffset(MF, ObjectOffset).getFixed()
1332 : getStackOffset(MF, ObjectOffset).getFixed();
1333}
1334
1336 const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP,
1337 bool ForSimm) const {
1338 const auto &MFI = MF.getFrameInfo();
1339 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1340 bool isFixed = MFI.isFixedObjectIndex(FI);
1341 auto StackID = static_cast<TargetStackID::Value>(MFI.getStackID(FI));
1342 return resolveFrameOffsetReference(MF, ObjectOffset, isFixed, StackID,
1343 FrameReg, PreferFP, ForSimm);
1344}
1345
1347 const MachineFunction &MF, int64_t ObjectOffset, bool isFixed,
1348 TargetStackID::Value StackID, Register &FrameReg, bool PreferFP,
1349 bool ForSimm) const {
1350 const auto &MFI = MF.getFrameInfo();
1351 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1352 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1353 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1354
1355 int64_t FPOffset = getFPOffset(MF, ObjectOffset).getFixed();
1356 int64_t Offset = getStackOffset(MF, ObjectOffset).getFixed();
1357
1358 // The fixed object area sits above the callee-saved area and can grow
1359 // large enough to displace it; account for that displacement here so CSR
1360 // objects aren't misclassified as locals and addressed via the base
1361 // pointer.
1362 bool IsWin64 = Subtarget.isCallingConvWin64(MF.getFunction().getCallingConv(),
1363 MF.getFunction().isVarArg());
1364 const int64_t FixedObjectSize =
1365 getFixedObjectSize(MF, AFI, IsWin64, /*IsFunclet*/ false);
1366 bool isCSR =
1367 !isFixed && ObjectOffset >= -((int64_t)AFI->getCalleeSavedStackSize(MFI) +
1368 FixedObjectSize);
1369 bool isSVE = MFI.isScalableStackID(StackID);
1370
1371 StackOffset ZPRStackSize = getZPRStackSize(MF);
1372 StackOffset PPRStackSize = getPPRStackSize(MF);
1373 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1374
1375 // Use frame pointer to reference fixed objects. Use it for locals if
1376 // there are VLAs or a dynamically realigned SP (and thus the SP isn't
1377 // reliable as a base). Make sure useFPForScavengingIndex() does the
1378 // right thing for the emergency spill slot.
1379 bool UseFP = false;
1380 if (AFI->hasStackFrame() && !isSVE) {
1381 // We shouldn't prefer using the FP to access fixed-sized stack objects when
1382 // there are scalable (SVE) objects in between the FP and the fixed-sized
1383 // objects.
1384 PreferFP &= !SVEStackSize;
1385
1386 // Note: Keeping the following as multiple 'if' statements rather than
1387 // merging to a single expression for readability.
1388 //
1389 // Argument access should always use the FP.
1390 if (isFixed) {
1391 UseFP = hasFP(MF);
1392 } else if (isCSR && RegInfo->hasStackRealignment(MF)) {
1393 // References to the CSR area must use FP if we're re-aligning the stack
1394 // since the dynamically-sized alignment padding is between the SP/BP and
1395 // the CSR area.
1396 assert(hasFP(MF) && "Re-aligned stack must have frame pointer");
1397 UseFP = true;
1398 } else if (hasFP(MF) && !RegInfo->hasStackRealignment(MF)) {
1399 // If the FPOffset is negative and we're producing a signed immediate, we
1400 // have to keep in mind that the available offset range for negative
1401 // offsets is smaller than for positive ones. If an offset is available
1402 // via the FP and the SP, use whichever is closest.
1403 bool FPOffsetFits = !ForSimm || FPOffset >= -256;
1404 PreferFP |= Offset > -FPOffset && !SVEStackSize;
1405
1406 if (FPOffset >= 0) {
1407 // If the FPOffset is positive, that'll always be best, as the SP/BP
1408 // will be even further away.
1409 UseFP = true;
1410 } else if (MFI.hasVarSizedObjects()) {
1411 // If we have variable sized objects, we can use either FP or BP, as the
1412 // SP offset is unknown. We can use the base pointer if we have one and
1413 // FP is not preferred. If not, we're stuck with using FP.
1414 bool CanUseBP = RegInfo->hasBasePointer(MF);
1415 if (FPOffsetFits && CanUseBP) // Both are ok. Pick the best.
1416 UseFP = PreferFP;
1417 else if (!CanUseBP) // Can't use BP. Forced to use FP.
1418 UseFP = true;
1419 // else we can use BP and FP, but the offset from FP won't fit.
1420 // That will make us scavenge registers which we can probably avoid by
1421 // using BP. If it won't fit for BP either, we'll scavenge anyway.
1422 } else if (MF.hasEHFunclets() && !RegInfo->hasBasePointer(MF)) {
1423 // Funclets access the locals contained in the parent's stack frame
1424 // via the frame pointer, so we have to use the FP in the parent
1425 // function.
1426 (void) Subtarget;
1427 assert(Subtarget.isCallingConvWin64(MF.getFunction().getCallingConv(),
1428 MF.getFunction().isVarArg()) &&
1429 "Funclets should only be present on Win64");
1430 UseFP = true;
1431 } else {
1432 // We have the choice between FP and (SP or BP).
1433 if (FPOffsetFits && PreferFP) // If FP is the best fit, use it.
1434 UseFP = true;
1435 }
1436 }
1437 }
1438
1439 assert(
1440 ((isFixed || isCSR) || !RegInfo->hasStackRealignment(MF) || !UseFP) &&
1441 "In the presence of dynamic stack pointer realignment, "
1442 "non-argument/CSR objects cannot be accessed through the frame pointer");
1443
1444 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1445
1446 if (isSVE) {
1447 StackOffset FPOffset = StackOffset::get(
1448 -AFI->getCalleeSaveBaseToFrameRecordOffset(), ObjectOffset);
1449 StackOffset SPOffset =
1450 SVEStackSize +
1451 StackOffset::get(MFI.getStackSize() - AFI->getCalleeSavedStackSize(),
1452 ObjectOffset);
1453
1454 // With split SVE objects the ObjectOffset is relative to the split area
1455 // (i.e. the PPR area or ZPR area respectively).
1456 if (AFI->hasSplitSVEObjects() && StackID == TargetStackID::ScalableVector) {
1457 // If we're accessing an SVE vector with split SVE objects...
1458 // - From the FP we need to move down past the PPR area:
1459 FPOffset -= PPRStackSize;
1460 // - From the SP we only need to move up to the ZPR area:
1461 SPOffset -= PPRStackSize;
1462 // Note: `SPOffset = SVEStackSize + ...`, so `-= PPRStackSize` results in
1463 // `SPOffset = ZPRStackSize + ...`.
1464 }
1465
1466 if (FPAfterSVECalleeSaves) {
1468 if (-ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1471 }
1472 }
1473
1474 // Always use the FP for SVE spills if available and beneficial.
1475 if (hasFP(MF) && (SPOffset.getFixed() ||
1476 FPOffset.getScalable() < SPOffset.getScalable() ||
1477 RegInfo->hasStackRealignment(MF))) {
1478 FrameReg = RegInfo->getFrameRegister(MF);
1479 return FPOffset;
1480 }
1481 FrameReg = RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister()
1482 : MCRegister(AArch64::SP);
1483
1484 return SPOffset;
1485 }
1486
1487 StackOffset SVEAreaOffset = {};
1488 if (FPAfterSVECalleeSaves) {
1489 // In this stack layout, the FP is in between the callee saves and other
1490 // SVE allocations.
1491 StackOffset SVECalleeSavedStack =
1493 if (UseFP) {
1494 if (isFixed)
1495 SVEAreaOffset = SVECalleeSavedStack;
1496 else if (!isCSR)
1497 SVEAreaOffset = SVECalleeSavedStack - SVEStackSize;
1498 } else {
1499 if (isFixed)
1500 SVEAreaOffset = SVEStackSize;
1501 else if (isCSR)
1502 SVEAreaOffset = SVEStackSize - SVECalleeSavedStack;
1503 }
1504 } else {
1505 if (UseFP && !(isFixed || isCSR))
1506 SVEAreaOffset = -SVEStackSize;
1507 if (!UseFP && (isFixed || isCSR))
1508 SVEAreaOffset = SVEStackSize;
1509 }
1510
1511 if (UseFP) {
1512 FrameReg = RegInfo->getFrameRegister(MF);
1513 return StackOffset::getFixed(FPOffset) + SVEAreaOffset;
1514 }
1515
1516 // Use the base pointer if we have one.
1517 if (RegInfo->hasBasePointer(MF))
1518 FrameReg = RegInfo->getBaseRegister();
1519 else {
1520 assert(!MFI.hasVarSizedObjects() &&
1521 "Can't use SP when we have var sized objects.");
1522 FrameReg = AArch64::SP;
1523 // If we're using the red zone for this function, the SP won't actually
1524 // be adjusted, so the offsets will be negative. They're also all
1525 // within range of the signed 9-bit immediate instructions.
1526 if (canUseRedZone(MF))
1527 Offset -= AFI->getLocalStackSize();
1528 }
1529
1530 return StackOffset::getFixed(Offset) + SVEAreaOffset;
1531}
1532
1534 // Do not set a kill flag on values that are also marked as live-in. This
1535 // happens with the @llvm-returnaddress intrinsic and with arguments passed in
1536 // callee saved registers.
1537 // Omitting the kill flags is conservatively correct even if the live-in
1538 // is not used after all.
1539 bool IsLiveIn = MF.getRegInfo().isLiveIn(Reg);
1540 return getKillRegState(!IsLiveIn);
1541}
1542
1544 MachineFunction &MF) {
1545 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1548 return Subtarget.isTargetMachO() &&
1549 !(Subtarget.getTargetLowering()->supportSwiftError() &&
1550 Attrs.hasAttrSomewhere(Attribute::SwiftError)) &&
1552 !AFL.requiresSaveVG(MF) && !AFI->isSVECC();
1553}
1554
1555static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile,
1556 unsigned SpillCount, unsigned Reg1,
1557 unsigned Reg2, bool NeedsWinCFI,
1558 const TargetRegisterInfo *TRI) {
1559 // If we are generating register pairs for a Windows function that requires
1560 // EH support, then pair consecutive registers only. There are no unwind
1561 // opcodes for saves/restores of non-consecutive register pairs.
1562 // The unwind opcodes are save_regp, save_regp_x, save_fregp, save_frepg_x,
1563 // save_lrpair.
1564 // https://docs.microsoft.com/en-us/cpp/build/arm64-exception-handling
1565
1566 if (Reg2 == AArch64::FP)
1567 return true;
1568 if (!NeedsWinCFI)
1569 return false;
1570
1571 // ARM64EC introduced `save_any_regp`, which expects 16-byte alignment.
1572 // This is handled by only allowing paired spills for registers spilled at
1573 // even positions (which should be 16-byte aligned, as other GPRs/FPRs are
1574 // 8-bytes). We carve out an exception for {FP,LR}, which does not require
1575 // 16-byte alignment in the uop representation.
1576 if (TRI->getEncodingValue(Reg2) == TRI->getEncodingValue(Reg1) + 1)
1577 return SpillExtendedVolatile
1578 ? !((Reg1 == AArch64::FP && Reg2 == AArch64::LR) ||
1579 (SpillCount % 2) == 0)
1580 : false;
1581
1582 // If pairing a GPR with LR, the pair can be described by the save_lrpair
1583 // opcode. The save_lrpair opcode requires the first register to be odd.
1584 if (Reg1 >= AArch64::X19 && Reg1 <= AArch64::X27 &&
1585 (Reg1 - AArch64::X19) % 2 == 0 && Reg2 == AArch64::LR)
1586 return false;
1587 return true;
1588}
1589
1590/// Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
1591/// WindowsCFI requires that only consecutive registers can be paired.
1592/// LR and FP need to be allocated together when the frame needs to save
1593/// the frame-record. This means any other register pairing with LR is invalid.
1594static bool invalidateRegisterPairing(bool SpillExtendedVolatile,
1595 unsigned SpillCount, unsigned Reg1,
1596 unsigned Reg2, bool UsesWinAAPCS,
1597 bool NeedsWinCFI, bool NeedsFrameRecord,
1598 const TargetRegisterInfo *TRI) {
1599 if (UsesWinAAPCS)
1600 return invalidateWindowsRegisterPairing(SpillExtendedVolatile, SpillCount,
1601 Reg1, Reg2, NeedsWinCFI, TRI);
1602
1603 // If we need to store the frame record, don't pair any register
1604 // with LR other than FP.
1605 if (NeedsFrameRecord)
1606 return Reg2 == AArch64::LR;
1607
1608 return false;
1609}
1610
1611// Returns true if Offset (in bytes) is aligned to the instruction's scale and
1612// the scaled immediate is within the instruction's valid range.
1613static bool isValidMemOpOffset(const AArch64InstrInfo *TII, unsigned Opcode,
1614 int Offset) {
1615 int64_t MinOff, MaxOff;
1616 TypeSize ScaleValue(0U, false), Width(0U, false);
1617 if (!TII->getMemOpInfo(Opcode, ScaleValue, Width, MinOff, MaxOff))
1618 return false;
1619
1620 if (Offset % ScaleValue.getKnownMinValue() != 0)
1621 return false;
1622
1623 Offset /= ScaleValue.getKnownMinValue();
1624 return Offset >= MinOff && Offset <= MaxOff;
1625}
1626
1627namespace {
1628
1629struct RegPairInfo {
1630 Register Reg1;
1631 Register Reg2;
1632 int FrameIdx;
1633 int Offset;
1634 enum RegType { GPR, FPR64, FPR128, PPR, ZPR, VG } Type;
1635 const TargetRegisterClass *RC;
1636
1637 RegPairInfo() = default;
1638
1639 bool isPaired() const { return Reg2.isValid(); }
1640
1641 bool isScalable() const { return Type == PPR || Type == ZPR; }
1642};
1643
1644} // end anonymous namespace
1645
1647 for (unsigned PReg = AArch64::P8; PReg <= AArch64::P15; ++PReg) {
1648 if (SavedRegs.test(PReg)) {
1649 unsigned PNReg = PReg - AArch64::P0 + AArch64::PN0;
1650 return MCRegister(PNReg);
1651 }
1652 }
1653 return MCRegister();
1654}
1655
1656// The multivector LD/ST are available only for SME or SVE2p1 targets
1658 MachineFunction &MF) {
1659 if (Subtarget.getCLOpts().disable_multivector_spill_fill)
1660 return false;
1661
1662 SMEAttrs FuncAttrs = MF.getInfo<AArch64FunctionInfo>()->getSMEFnAttrs();
1663 bool IsLocallyStreaming =
1664 FuncAttrs.hasStreamingBody() && !FuncAttrs.hasStreamingInterface();
1665
1666 // Only when in streaming mode SME2 instructions can be safely used.
1667 // It is not safe to use SME2 instructions when in streaming compatible or
1668 // locally streaming mode.
1669 return Subtarget.hasSVE2p1() ||
1670 (Subtarget.hasSME2() &&
1671 (!IsLocallyStreaming && Subtarget.isStreaming()));
1672}
1673
1675 MachineFunction &MF,
1677 const TargetRegisterInfo *TRI,
1679 bool NeedsFrameRecord) {
1680
1681 if (CSI.empty())
1682 return;
1683
1684 const AArch64InstrInfo *TII =
1685 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
1686 bool IsWindows = isTargetWindows(MF);
1688 unsigned StackHazardSize = getStackHazardSize(MF);
1689 MachineFrameInfo &MFI = MF.getFrameInfo();
1691 unsigned Count = CSI.size();
1692 (void)CC;
1693 // MachO's compact unwind format relies on all registers being stored in
1694 // pairs.
1695 assert((!produceCompactUnwindFrame(AFL, MF) ||
1698 (Count & 1) == 0) &&
1699 "Odd number of callee-saved regs to spill!");
1700 int ByteOffset = AFI->getCalleeSavedStackSize();
1701 int StackFillDir = -1;
1702 int RegInc = 1;
1703 unsigned FirstReg = 0;
1704 if (IsWindows) {
1705 // For WinCFI, fill the stack from the bottom up.
1706 ByteOffset = 0;
1707 StackFillDir = 1;
1708 // As the CSI array is reversed to match PrologEpilogInserter, iterate
1709 // backwards, to pair up registers starting from lower numbered registers.
1710 RegInc = -1;
1711 FirstReg = Count - 1;
1712 }
1713
1714 bool FPAfterSVECalleeSaves = AFL.hasSVECalleeSavesAboveFrameRecord(MF);
1715 // Windows AAPCS has x9-x15 as volatile registers, x16-x17 as intra-procedural
1716 // scratch, x18 as platform reserved. However, clang has extended calling
1717 // convensions such as preserve_most and preserve_all which treat these as
1718 // CSR. As such, the ARM64 unwind uOPs bias registers by 19. We use ARM64EC
1719 // uOPs which have separate restrictions. We need to check for that.
1720 //
1721 // NOTE: we currently do not account for the D registers as LLVM does not
1722 // support non-ABI compliant D register spills.
1723 bool SpillExtendedVolatile =
1724 IsWindows && llvm::any_of(CSI, [](const CalleeSavedInfo &CSI) {
1725 const auto &Reg = CSI.getReg();
1726 return Reg >= AArch64::X0 && Reg <= AArch64::X18;
1727 });
1728
1729 int ZPRByteOffset = 0;
1730 int PPRByteOffset = 0;
1731 bool SplitPPRs = AFI->hasSplitSVEObjects();
1732 if (SplitPPRs) {
1733 ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
1734 PPRByteOffset = AFI->getPPRCalleeSavedStackSize();
1735 } else if (!FPAfterSVECalleeSaves) {
1736 ZPRByteOffset =
1738 // Unused: Everything goes in ZPR space.
1739 PPRByteOffset = 0;
1740 }
1741
1742 bool NeedGapToAlignStack = AFI->hasCalleeSaveStackFreeSpace();
1743 Register LastReg = 0;
1744 bool HasCSHazardPadding = AFI->hasStackHazardSlotIndex() && !SplitPPRs;
1745
1746 auto AlignOffset = [StackFillDir](int Offset, int Align) {
1747 if (StackFillDir < 0)
1748 return alignDown(Offset, Align);
1749 return alignTo(Offset, Align);
1750 };
1751
1752 // When iterating backwards, the loop condition relies on unsigned wraparound.
1753 for (unsigned i = FirstReg; i < Count; i += RegInc) {
1754 RegPairInfo RPI;
1755 RPI.Reg1 = CSI[i].getReg();
1756
1757 if (AArch64::GPR64RegClass.contains(RPI.Reg1)) {
1758 RPI.Type = RegPairInfo::GPR;
1759 RPI.RC = &AArch64::GPR64RegClass;
1760 } else if (AArch64::FPR64RegClass.contains(RPI.Reg1)) {
1761 RPI.Type = RegPairInfo::FPR64;
1762 RPI.RC = &AArch64::FPR64RegClass;
1763 } else if (AArch64::FPR128RegClass.contains(RPI.Reg1)) {
1764 RPI.Type = RegPairInfo::FPR128;
1765 RPI.RC = &AArch64::FPR128RegClass;
1766 } else if (AArch64::ZPRRegClass.contains(RPI.Reg1)) {
1767 RPI.Type = RegPairInfo::ZPR;
1768 RPI.RC = &AArch64::ZPRRegClass;
1769 } else if (AArch64::PPRRegClass.contains(RPI.Reg1)) {
1770 RPI.Type = RegPairInfo::PPR;
1771 RPI.RC = &AArch64::PPRRegClass;
1772 } else if (RPI.Reg1 == AArch64::VG) {
1773 RPI.Type = RegPairInfo::VG;
1774 RPI.RC = &AArch64::FIXED_REGSRegClass;
1775 } else {
1776 llvm_unreachable("Unsupported register class.");
1777 }
1778
1779 int &ScalableByteOffset = RPI.Type == RegPairInfo::PPR && SplitPPRs
1780 ? PPRByteOffset
1781 : ZPRByteOffset;
1782
1783 // Add the stack hazard size as we transition from GPR->FPR CSRs.
1784 if (HasCSHazardPadding &&
1785 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
1787 ByteOffset += StackFillDir * StackHazardSize;
1788 LastReg = RPI.Reg1;
1789
1790 bool NeedsWinCFI = AFL.needsWinCFI(MF);
1791 int Scale = TRI->getSpillSize(*RPI.RC);
1792 // Add the next reg to the pair if it is in the same register class.
1793 if (unsigned(i + RegInc) < Count && !HasCSHazardPadding) {
1794 MCRegister NextReg = CSI[i + RegInc].getReg();
1795 unsigned SpillCount = NeedsWinCFI ? FirstReg - i : i;
1796 int Aligned = AlignOffset(ByteOffset, Scale);
1797 int PairOffset = IsWindows ? Aligned : Aligned + StackFillDir * 2 * Scale;
1798 bool PairFitsImmRange =
1799 PairOffset / Scale >= -64 && PairOffset / Scale <= 63;
1800 switch (RPI.Type) {
1801 case RegPairInfo::GPR:
1802 if (AArch64::GPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1803 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1804 RPI.Reg1, NextReg, IsWindows,
1805 NeedsWinCFI, NeedsFrameRecord, TRI))
1806 RPI.Reg2 = NextReg;
1807 break;
1808 case RegPairInfo::FPR64:
1809 if (AArch64::FPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1810 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1811 RPI.Reg1, NextReg, IsWindows,
1812 NeedsWinCFI, NeedsFrameRecord, TRI))
1813 RPI.Reg2 = NextReg;
1814 break;
1815 case RegPairInfo::FPR128:
1816 if (AArch64::FPR128RegClass.contains(NextReg) && PairFitsImmRange)
1817 RPI.Reg2 = NextReg;
1818 break;
1819 case RegPairInfo::PPR:
1820 break;
1821 case RegPairInfo::ZPR:
1822 // Windows support is possible but the order requirement is reversed.
1823 // Also, WinCFI has no support for group ZPR loads/stores yet.
1824 if (isTargetWindows(MF) || AFI->getPredicateRegForFillSpill() == 0)
1825 break;
1826 // We expect to see pairs in decending order (e.g. [z9, z8]).
1827 // We ensure this in `orderZPRCalleeSavesForPairs()`. This is required
1828 // as (for Linux) StackFillDir is negative (so we start at higher
1829 // addresses) and we need to ensure the lower register in the pair has
1830 // the lower address to store the registers in the correct order.
1831 if (((NextReg - AArch64::Z0) % 2 == 0) && (NextReg + 1 == RPI.Reg1)) {
1832 const int NumRegs = 2;
1833 int Offset = (ScalableByteOffset + StackFillDir * NumRegs * Scale);
1834
1835 // Note: ST1B has the same offset constraints.
1836 if (isValidMemOpOffset(TII, AArch64::LD1B_2Z_IMM, Offset))
1837 RPI.Reg2 = NextReg;
1838 }
1839 break;
1840 case RegPairInfo::VG:
1841 break;
1842 }
1843 }
1844
1845 // GPRs and FPRs are saved in pairs of 64-bit regs. We expect the CSI
1846 // list to come in sorted by frame index so that we can issue the store
1847 // pair instructions directly. Assert if we see anything otherwise.
1848 //
1849 // The order of the registers in the list is controlled by
1850 // getCalleeSavedRegs(), so they will always be in-order, as well.
1851 assert((!RPI.isPaired() ||
1852 (CSI[i].getFrameIdx() + RegInc == CSI[i + RegInc].getFrameIdx())) &&
1853 "Out of order callee saved regs!");
1854
1855 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg2 != AArch64::FP ||
1856 RPI.Reg1 == AArch64::LR) &&
1857 "FrameRecord must be allocated together with LR");
1858
1859 // Windows AAPCS has FP and LR reversed.
1860 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg1 != AArch64::FP ||
1861 RPI.Reg2 == AArch64::LR) &&
1862 "FrameRecord must be allocated together with LR");
1863
1864 // MachO's compact unwind format relies on all registers being stored in
1865 // adjacent register pairs.
1866 assert((!produceCompactUnwindFrame(AFL, MF) ||
1869 (RPI.isPaired() &&
1870 ((RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP) ||
1871 RPI.Reg1 + 1 == RPI.Reg2))) &&
1872 "Callee-save registers not saved as adjacent register pair!");
1873
1874 RPI.FrameIdx = CSI[i].getFrameIdx();
1875 if (IsWindows &&
1876 RPI.isPaired()) // RPI.FrameIdx must be the lower index of the pair
1877 RPI.FrameIdx = CSI[i + RegInc].getFrameIdx();
1878
1879 // Realign the scalable offset if necessary. This is relevant when spilling
1880 // predicates on Windows.
1881 if (RPI.isScalable() && ScalableByteOffset % Scale != 0)
1882 ScalableByteOffset = AlignOffset(ScalableByteOffset, Scale);
1883
1884 // Realign the fixed offset if necessary. This is relevant when spilling Q
1885 // registers after spilling an odd amount of X registers.
1886 if (!RPI.isScalable() && ByteOffset % Scale != 0)
1887 ByteOffset = AlignOffset(ByteOffset, Scale);
1888
1889 int OffsetPre = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1890 assert(OffsetPre % Scale == 0);
1891
1892 if (RPI.isScalable())
1893 ScalableByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1894 else
1895 ByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1896
1897 // Swift's async context is directly before FP, so allocate an extra
1898 // 8 bytes for it.
1899 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1900 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1901 (IsWindows && RPI.Reg2 == AArch64::LR)))
1902 ByteOffset += StackFillDir * 8;
1903
1904 // Round up size of non-pair to pair size if we need to pad the
1905 // callee-save area to ensure 16-byte alignment.
1906 if (NeedGapToAlignStack && !IsWindows && !RPI.isScalable() &&
1907 RPI.Type != RegPairInfo::FPR128 && !RPI.isPaired() &&
1908 ByteOffset % 16 != 0) {
1909 ByteOffset += 8 * StackFillDir;
1910 assert(MFI.getObjectAlign(RPI.FrameIdx) <= Align(16));
1911 // A stack frame with a gap looks like this, bottom up:
1912 // d9, d8. x21, gap, x20, x19.
1913 // Set extra alignment on the x21 object to create the gap above it.
1914 MFI.setObjectAlignment(RPI.FrameIdx, Align(16));
1915 NeedGapToAlignStack = false;
1916 }
1917
1918 int OffsetPost = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1919 assert(OffsetPost % Scale == 0);
1920 // If filling top down (default), we want the offset after incrementing it.
1921 // If filling bottom up (WinCFI) we need the original offset.
1922 int Offset = IsWindows ? OffsetPre : OffsetPost;
1923
1924 // The FP, LR pair goes 8 bytes into our expanded 24-byte slot so that the
1925 // Swift context can directly precede FP.
1926 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1927 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1928 (IsWindows && RPI.Reg2 == AArch64::LR)))
1929 Offset += 8;
1930 RPI.Offset = Offset / Scale;
1931
1932 assert((!RPI.isPaired() ||
1933 (!RPI.isScalable() && RPI.Offset >= -64 && RPI.Offset <= 63) ||
1934 (RPI.isScalable() && RPI.Offset >= -256 && RPI.Offset <= 255)) &&
1935 "Offset out of bounds for LDP/STP immediate");
1936
1937 auto isFrameRecord = [&] {
1938 if (RPI.isPaired())
1939 return IsWindows ? RPI.Reg1 == AArch64::FP && RPI.Reg2 == AArch64::LR
1940 : RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP;
1941 // Otherwise, look for the frame record as two unpaired registers. This is
1942 // needed for -aarch64-stack-hazard-size=<val>, which disables register
1943 // pairing (as the padding may be too large for the LDP/STP offset). Note:
1944 // On Windows, this check works out as current reg == FP, next reg == LR,
1945 // and on other platforms current reg == FP, previous reg == LR. This
1946 // works out as the correct pre-increment or post-increment offsets
1947 // respectively.
1948 return i > 0 && RPI.Reg1 == AArch64::FP &&
1949 CSI[i - 1].getReg() == AArch64::LR;
1950 };
1951
1952 // Save the offset to frame record so that the FP register can point to the
1953 // innermost frame record (spilled FP and LR registers).
1954 if (NeedsFrameRecord && isFrameRecord())
1956
1957 RegPairs.push_back(RPI);
1958 if (RPI.isPaired())
1959 i += RegInc;
1960 }
1961 if (IsWindows) {
1962 // If we need an alignment gap in the stack, align the topmost stack
1963 // object. A stack frame with a gap looks like this, bottom up:
1964 // x19, d8. d9, gap.
1965 // Set extra alignment on the topmost stack object (the first element in
1966 // CSI, which goes top down), to create the gap above it.
1967 if (AFI->hasCalleeSaveStackFreeSpace())
1968 MFI.setObjectAlignment(CSI[0].getFrameIdx(), Align(16));
1969 // We iterated bottom up over the registers; flip RegPairs back to top
1970 // down order.
1971 std::reverse(RegPairs.begin(), RegPairs.end());
1972 }
1973}
1974
1978 MachineFunction &MF = *MBB.getParent();
1979 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1980 auto &TLI = *Subtarget.getTargetLowering();
1981 const AArch64InstrInfo &TII = *Subtarget.getInstrInfo();
1982 bool NeedsWinCFI = needsWinCFI(MF);
1983 DebugLoc DL;
1985
1986 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
1987
1988 MachineRegisterInfo &MRI = MF.getRegInfo();
1989 // Refresh the reserved regs in case there are any potential changes since the
1990 // last freeze.
1991 MRI.freezeReservedRegs();
1992
1993 if (homogeneousPrologEpilog(MF)) {
1994 auto MIB = BuildMI(MBB, MI, DL, TII.get(AArch64::HOM_Prolog))
1996
1997 for (auto &RPI : RegPairs) {
1998 MIB.addReg(RPI.Reg1);
1999 MIB.addReg(RPI.Reg2);
2000
2001 // Update register live in.
2002 if (!MRI.isReserved(RPI.Reg1))
2003 MBB.addLiveIn(RPI.Reg1);
2004 if (RPI.isPaired() && !MRI.isReserved(RPI.Reg2))
2005 MBB.addLiveIn(RPI.Reg2);
2006 }
2007 return true;
2008 }
2009 bool PTrueCreated = false;
2010 for (const RegPairInfo &RPI : llvm::reverse(RegPairs)) {
2011 Register Reg1 = RPI.Reg1;
2012 Register Reg2 = RPI.Reg2;
2013 unsigned StrOpc;
2014
2015 // Issue sequence of spills for cs regs. The first spill may be converted
2016 // to a pre-decrement store later by emitPrologue if the callee-save stack
2017 // area allocation can't be combined with the local stack area allocation.
2018 // For example:
2019 // stp x22, x21, [sp, #0] // addImm(+0)
2020 // stp x20, x19, [sp, #16] // addImm(+2)
2021 // stp fp, lr, [sp, #32] // addImm(+4)
2022 // Rationale: This sequence saves uop updates compared to a sequence of
2023 // pre-increment spills like stp xi,xj,[sp,#-16]!
2024 // Note: Similar rationale and sequence for restores in epilog.
2025 unsigned Size = TRI->getSpillSize(*RPI.RC);
2026 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2027 switch (RPI.Type) {
2028 case RegPairInfo::GPR:
2029 StrOpc = RPI.isPaired() ? AArch64::STPXi : AArch64::STRXui;
2030 break;
2031 case RegPairInfo::FPR64:
2032 StrOpc = RPI.isPaired() ? AArch64::STPDi : AArch64::STRDui;
2033 break;
2034 case RegPairInfo::FPR128:
2035 StrOpc = RPI.isPaired() ? AArch64::STPQi : AArch64::STRQui;
2036 break;
2037 case RegPairInfo::ZPR:
2038 StrOpc = RPI.isPaired() ? AArch64::ST1B_2Z_IMM : AArch64::STR_ZXI;
2039 break;
2040 case RegPairInfo::PPR:
2041 StrOpc = AArch64::STR_PXI;
2042 break;
2043 case RegPairInfo::VG:
2044 StrOpc = AArch64::STRXui;
2045 break;
2046 }
2047
2048 Register X0Scratch;
2049 llvm::scope_exit RestoreX0([&] {
2050 if (X0Scratch.isValid())
2051 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), AArch64::X0)
2052 .addReg(X0Scratch)
2054 });
2055
2056 if (Reg1 == AArch64::VG) {
2057 // Find an available register to store value of VG to.
2058 Reg1 = findScratchNonCalleeSaveRegister(&MBB, true);
2059 assert(Reg1.isValid());
2060 if (MF.getSubtarget<AArch64Subtarget>().hasSVE()) {
2061 BuildMI(MBB, MI, DL, TII.get(AArch64::CNTD_XPiI), Reg1)
2062 .addImm(31)
2063 .addImm(1)
2065 } else {
2067 if (any_of(MBB.liveins(),
2068 [&STI](const MachineBasicBlock::RegisterMaskPair &LiveIn) {
2069 return STI.getRegisterInfo()->isSuperOrSubRegisterEq(
2070 AArch64::X0, LiveIn.PhysReg);
2071 })) {
2072 X0Scratch = Reg1;
2073 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), X0Scratch)
2074 .addReg(AArch64::X0)
2076 }
2077
2078 RTLIB::Libcall LC = RTLIB::SMEABI_GET_CURRENT_VG;
2079 const uint32_t *RegMask =
2080 TRI->getCallPreservedMask(MF, TLI.getLibcallCallingConv(LC));
2081 BuildMI(MBB, MI, DL, TII.get(AArch64::BL))
2082 .addExternalSymbol(TLI.getLibcallName(LC))
2083 .addRegMask(RegMask)
2084 .addReg(AArch64::X0, RegState::ImplicitDefine)
2086 Reg1 = AArch64::X0;
2087 }
2088 }
2089
2090 LLVM_DEBUG({
2091 dbgs() << "CSR spill: (" << printReg(Reg1, TRI);
2092 if (RPI.isPaired())
2093 dbgs() << ", " << printReg(Reg2, TRI);
2094 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2095 if (RPI.isPaired())
2096 dbgs() << ", " << RPI.FrameIdx + 1;
2097 dbgs() << ")\n";
2098 });
2099
2100 assert((!isTargetWindows(MF) ||
2101 !(Reg1 == AArch64::LR && Reg2 == AArch64::FP)) &&
2102 "Windows unwdinding requires a consecutive (FP,LR) pair");
2103 unsigned FrameIdxReg1 = RPI.FrameIdx;
2104 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2105
2106 if (RPI.isPaired() && RPI.isScalable()) {
2107 assert(!isTargetWindows(MF) &&
2108 "Scalable register groups are not supported by Windows WinCFI");
2109 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2112 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2113 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2114 "Expects SVE2.1 or SME2 target and a predicate register");
2115#ifdef EXPENSIVE_CHECKS
2116 auto IsPPR = [](const RegPairInfo &c) {
2117 return c.Type == RegPairInfo::PPR;
2118 };
2119 auto PPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsPPR);
2120 auto IsZPR = [](const RegPairInfo &c) {
2121 return c.Type == RegPairInfo::ZPR;
2122 };
2123 auto ZPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsZPR);
2124 assert(!(PPRBegin < ZPRBegin) &&
2125 "Expected callee save predicate to be handled first");
2126#endif
2127 if (!PTrueCreated) {
2128 PTrueCreated = true;
2129 BuildMI(MBB, MI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2131 }
2132 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2133 if (!MRI.isReserved(Reg1))
2134 MBB.addLiveIn(Reg1);
2135 if (!MRI.isReserved(Reg2))
2136 MBB.addLiveIn(Reg2);
2137 assert(RPI.Reg2 + 1 == RPI.Reg1 && "Expected reversed ZPR pair");
2138 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg2 - AArch64::Z0));
2140 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2141 MachineMemOperand::MOStore, Size, Alignment));
2142 MIB.addReg(PnReg);
2143 MIB.addReg(AArch64::SP)
2144 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale],
2145 // where 2*vscale is implicit
2148 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2149 MachineMemOperand::MOStore, Size, Alignment));
2150 } else { // The code when the pair of ZReg is not present
2151 // Windows unwind codes require consecutive registers if registers are
2152 // paired. Make the switch here, so that the code below will save (x,x+1)
2153 // and not (x+1,x).
2154 if (isTargetWindows(MF) && RPI.isPaired()) {
2155 std::swap(Reg1, Reg2);
2156 std::swap(FrameIdxReg1, FrameIdxReg2);
2157 }
2158 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2159 if (!MRI.isReserved(Reg1))
2160 MBB.addLiveIn(Reg1);
2161 if (RPI.isPaired()) {
2162 if (!MRI.isReserved(Reg2))
2163 MBB.addLiveIn(Reg2);
2164 MIB.addReg(Reg2, getPrologueDeath(MF, Reg2));
2166 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2167 MachineMemOperand::MOStore, Size, Alignment));
2168 }
2169 MIB.addReg(Reg1, getPrologueDeath(MF, Reg1))
2170 .addReg(AArch64::SP)
2171 .addImm(RPI.Offset) // [sp, #offset*vscale],
2172 // where factor*vscale is implicit
2175 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2176 MachineMemOperand::MOStore, Size, Alignment));
2177 if (NeedsWinCFI)
2178 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2179 }
2180 // Update the StackIDs of the SVE stack slots.
2181 MachineFrameInfo &MFI = MF.getFrameInfo();
2182 if (RPI.Type == RegPairInfo::ZPR) {
2183 MFI.setStackID(FrameIdxReg1, TargetStackID::ScalableVector);
2184 if (RPI.isPaired())
2185 MFI.setStackID(FrameIdxReg2, TargetStackID::ScalableVector);
2186 } else if (RPI.Type == RegPairInfo::PPR) {
2188 if (RPI.isPaired())
2190 }
2191 }
2192 return true;
2193}
2194
2198 MachineFunction &MF = *MBB.getParent();
2199 const AArch64InstrInfo &TII =
2200 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
2201 DebugLoc DL;
2203 bool NeedsWinCFI = needsWinCFI(MF);
2204
2205 if (MBBI != MBB.end())
2206 DL = MBBI->getDebugLoc();
2207
2208 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
2209 if (homogeneousPrologEpilog(MF, &MBB)) {
2210 auto MIB = BuildMI(MBB, MBBI, DL, TII.get(AArch64::HOM_Epilog))
2212 for (auto &RPI : RegPairs) {
2213 MIB.addReg(RPI.Reg1, RegState::Define);
2214 MIB.addReg(RPI.Reg2, RegState::Define);
2215 }
2216 return true;
2217 }
2218
2219 // For performance reasons restore SVE register in increasing order
2220 auto IsPPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::PPR; };
2221 auto PPRBegin = llvm::find_if(RegPairs, IsPPR);
2222 auto PPREnd = std::find_if_not(PPRBegin, RegPairs.end(), IsPPR);
2223 std::reverse(PPRBegin, PPREnd);
2224 auto IsZPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::ZPR; };
2225 auto ZPRBegin = llvm::find_if(RegPairs, IsZPR);
2226 auto ZPREnd = std::find_if_not(ZPRBegin, RegPairs.end(), IsZPR);
2227 std::reverse(ZPRBegin, ZPREnd);
2228
2229 bool PTrueCreated = false;
2230 for (const RegPairInfo &RPI : RegPairs) {
2231 Register Reg1 = RPI.Reg1;
2232 Register Reg2 = RPI.Reg2;
2233
2234 // Issue sequence of restores for cs regs. The last restore may be converted
2235 // to a post-increment load later by emitEpilogue if the callee-save stack
2236 // area allocation can't be combined with the local stack area allocation.
2237 // For example:
2238 // ldp fp, lr, [sp, #32] // addImm(+4)
2239 // ldp x20, x19, [sp, #16] // addImm(+2)
2240 // ldp x22, x21, [sp, #0] // addImm(+0)
2241 // Note: see comment in spillCalleeSavedRegisters()
2242 unsigned LdrOpc;
2243 unsigned Size = TRI->getSpillSize(*RPI.RC);
2244 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2245 switch (RPI.Type) {
2246 case RegPairInfo::GPR:
2247 LdrOpc = RPI.isPaired() ? AArch64::LDPXi : AArch64::LDRXui;
2248 break;
2249 case RegPairInfo::FPR64:
2250 LdrOpc = RPI.isPaired() ? AArch64::LDPDi : AArch64::LDRDui;
2251 break;
2252 case RegPairInfo::FPR128:
2253 LdrOpc = RPI.isPaired() ? AArch64::LDPQi : AArch64::LDRQui;
2254 break;
2255 case RegPairInfo::ZPR:
2256 LdrOpc = RPI.isPaired() ? AArch64::LD1B_2Z_IMM : AArch64::LDR_ZXI;
2257 break;
2258 case RegPairInfo::PPR:
2259 LdrOpc = AArch64::LDR_PXI;
2260 break;
2261 case RegPairInfo::VG:
2262 continue;
2263 }
2264 LLVM_DEBUG({
2265 dbgs() << "CSR restore: (" << printReg(Reg1, TRI);
2266 if (RPI.isPaired())
2267 dbgs() << ", " << printReg(Reg2, TRI);
2268 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2269 if (RPI.isPaired())
2270 dbgs() << ", " << RPI.FrameIdx + 1;
2271 dbgs() << ")\n";
2272 });
2273
2274 unsigned FrameIdxReg1 = RPI.FrameIdx;
2275 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2276
2278 if (RPI.isPaired() && RPI.isScalable()) {
2279 assert(!isTargetWindows(MF) &&
2280 "Scalable register groups are not supported by Windows WinCFI");
2281 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2283 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2284 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2285 "Expects SVE2.1 or SME2 target and a predicate register");
2286#ifdef EXPENSIVE_CHECKS
2287 assert(!(PPRBegin < ZPRBegin) &&
2288 "Expected callee save predicate to be handled first");
2289#endif
2290 if (!PTrueCreated) {
2291 PTrueCreated = true;
2292 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2294 }
2295 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2296 assert(RPI.Reg2 + 1 == RPI.Reg1 && "Expected reversed ZPR pair");
2297 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg2 - AArch64::Z0),
2298 getDefRegState(true));
2300 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2301 MachineMemOperand::MOLoad, Size, Alignment));
2302 MIB.addReg(PnReg);
2303 MIB.addReg(AArch64::SP)
2304 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale]
2305 // where 2*vscale is implicit
2308 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2309 MachineMemOperand::MOLoad, Size, Alignment));
2310 } else {
2311 // Windows unwind codes require consecutive registers if registers are
2312 // paired. Make the switch here, so that the code below will save (x,x+1)
2313 // and not (x+1,x).
2314 if (isTargetWindows(MF) && RPI.isPaired()) {
2315 std::swap(Reg1, Reg2);
2316 std::swap(FrameIdxReg1, FrameIdxReg2);
2317 }
2318 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2319 if (RPI.isPaired()) {
2320 MIB.addReg(Reg2, getDefRegState(true));
2322 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2323 MachineMemOperand::MOLoad, Size, Alignment));
2324 }
2325 MIB.addReg(Reg1, getDefRegState(true));
2326 MIB.addReg(AArch64::SP)
2327 .addImm(RPI.Offset) // [sp, #offset*vscale]
2328 // where factor*vscale is implicit
2331 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2332 MachineMemOperand::MOLoad, Size, Alignment));
2333 if (NeedsWinCFI)
2334 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2335 }
2336 }
2337 return true;
2338}
2339
2340// Return the FrameID for a MMO.
2341static std::optional<int> getMMOFrameID(MachineMemOperand *MMO,
2342 const MachineFrameInfo &MFI) {
2343 auto *PSV =
2345 if (PSV)
2346 return std::optional<int>(PSV->getFrameIndex());
2347
2348 if (MMO->getValue()) {
2349 if (auto *Al = dyn_cast<AllocaInst>(getUnderlyingObject(MMO->getValue()))) {
2350 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd();
2351 FI++)
2352 if (MFI.getObjectAllocation(FI) == Al)
2353 return FI;
2354 }
2355 }
2356
2357 return std::nullopt;
2358}
2359
2360// Return the FrameID for a Load/Store instruction by looking at the first MMO.
2361static std::optional<int> getLdStFrameID(const MachineInstr &MI,
2362 const MachineFrameInfo &MFI) {
2363 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
2364 return std::nullopt;
2365
2366 return getMMOFrameID(*MI.memoperands_begin(), MFI);
2367}
2368
2369// Returns true if the LDST MachineInstr \p MI is a PPR access.
2370static bool isPPRAccess(const MachineInstr &MI) {
2371 return AArch64::PPRRegClass.contains(MI.getOperand(0).getReg());
2372}
2373
2374// Check if a Hazard slot is needed for the current function, and if so create
2375// one for it. The index is stored in AArch64FunctionInfo->StackHazardSlotIndex,
2376// which can be used to determine if any hazard padding is needed.
2377void AArch64FrameLowering::determineStackHazardSlot(
2378 MachineFunction &MF, BitVector &SavedRegs) const {
2379 unsigned StackHazardSize = getStackHazardSize(MF);
2380 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2381 if (StackHazardSize == 0 || StackHazardSize % 16 != 0 ||
2383 return;
2384
2385 // Stack hazards are only needed in streaming functions.
2386 SMEAttrs Attrs = AFI->getSMEFnAttrs();
2387 const AArch64Options &CLOpts =
2388 MF.getSubtarget<AArch64Subtarget>().getCLOpts();
2389 if (!CLOpts.stack_hazard_in_non_streaming &&
2390 Attrs.hasNonStreamingInterfaceAndBody())
2391 return;
2392
2393 MachineFrameInfo &MFI = MF.getFrameInfo();
2394
2395 // Add a hazard slot if there are any CSR FPR registers, or are any fp-only
2396 // stack objects.
2397 bool HasFPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2398 return AArch64::FPR64RegClass.contains(Reg) ||
2399 AArch64::FPR128RegClass.contains(Reg) ||
2400 AArch64::ZPRRegClass.contains(Reg);
2401 });
2402 bool HasPPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2403 return AArch64::PPRRegClass.contains(Reg);
2404 });
2405 bool HasFPRStackObjects = false;
2406 bool HasPPRStackObjects = false;
2407 if (!HasFPRCSRs || CLOpts.split_sve_objects) {
2408 enum SlotType : uint8_t {
2409 Unknown = 0,
2410 ZPRorFPR = 1 << 0,
2411 PPR = 1 << 1,
2412 GPR = 1 << 2,
2414 };
2415
2416 // Find stack slots solely used for one kind of register (ZPR, PPR, etc.),
2417 // based on the kinds of accesses used in the function.
2418 SmallVector<SlotType> SlotTypes(MFI.getObjectIndexEnd(), SlotType::Unknown);
2419 for (auto &MBB : MF) {
2420 for (auto &MI : MBB) {
2421 std::optional<int> FI = getLdStFrameID(MI, MFI);
2422 if (!FI || FI < 0 || FI > int(SlotTypes.size()))
2423 continue;
2424 if (MFI.hasScalableStackID(*FI)) {
2425 SlotTypes[*FI] |=
2426 isPPRAccess(MI) ? SlotType::PPR : SlotType::ZPRorFPR;
2427 } else {
2428 SlotTypes[*FI] |= AArch64InstrInfo::isFpOrNEON(MI)
2429 ? SlotType::ZPRorFPR
2430 : SlotType::GPR;
2431 }
2432 }
2433 }
2434
2435 for (int FI = 0; FI < int(SlotTypes.size()); ++FI) {
2436 HasFPRStackObjects |= SlotTypes[FI] == SlotType::ZPRorFPR;
2437 // For SplitSVEObjects remember that this stack slot is a predicate, this
2438 // will be needed later when determining the frame layout.
2439 if (SlotTypes[FI] == SlotType::PPR) {
2441 HasPPRStackObjects = true;
2442 }
2443 }
2444 }
2445
2446 if (HasFPRCSRs || HasFPRStackObjects) {
2447 int ID = MFI.CreateStackObject(StackHazardSize, Align(16), false);
2448 LLVM_DEBUG(dbgs() << "Created Hazard slot at " << ID << " size "
2449 << StackHazardSize << "\n");
2450 AFI->setStackHazardSlotIndex(ID);
2451 }
2452
2453 if (!AFI->hasStackHazardSlotIndex())
2454 return;
2455
2456 if (CLOpts.split_sve_objects) {
2457 CallingConv::ID CC = MF.getFunction().getCallingConv();
2458 if (AFI->isSVECC() || CC == CallingConv::AArch64_SVE_VectorCall) {
2459 AFI->setSplitSVEObjects(true);
2460 LLVM_DEBUG(dbgs() << "Using SplitSVEObjects for SVE CC function\n");
2461 return;
2462 }
2463
2464 // We only use SplitSVEObjects in non-SVE CC functions if there's a
2465 // possibility of a stack hazard between PPRs and ZPRs/FPRs.
2466 LLVM_DEBUG(dbgs() << "Determining if SplitSVEObjects should be used in "
2467 "non-SVE CC function...\n");
2468
2469 // If another calling convention is explicitly set FPRs can't be promoted to
2470 // ZPR callee-saves.
2472 LLVM_DEBUG(
2473 dbgs()
2474 << "Calling convention is not supported with SplitSVEObjects\n");
2475 return;
2476 }
2477
2478 if (!HasPPRCSRs && !HasPPRStackObjects) {
2479 LLVM_DEBUG(
2480 dbgs() << "Not using SplitSVEObjects as no PPRs are on the stack\n");
2481 return;
2482 }
2483
2484 if (!HasFPRCSRs && !HasFPRStackObjects) {
2485 LLVM_DEBUG(
2486 dbgs()
2487 << "Not using SplitSVEObjects as no FPRs or ZPRs are on the stack\n");
2488 return;
2489 }
2490
2491 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2492 MF.getSubtarget<AArch64Subtarget>();
2494 "Expected SVE to be available for PPRs");
2495
2496 const TargetRegisterInfo *TRI = MF.getSubtarget().getRegisterInfo();
2497 // With SplitSVEObjects the CS hazard padding is placed between the
2498 // PPRs and ZPRs. If there are any FPR CS there would be a hazard between
2499 // them and the CS GRPs. Avoid this by promoting all FPR CS to ZPRs.
2500 BitVector FPRZRegs(SavedRegs.size());
2501 for (size_t Reg = 0, E = SavedRegs.size(); HasFPRCSRs && Reg < E; ++Reg) {
2502 BitVector::reference RegBit = SavedRegs[Reg];
2503 if (!RegBit)
2504 continue;
2505 unsigned SubRegIdx = 0;
2506 if (AArch64::FPR64RegClass.contains(Reg))
2507 SubRegIdx = AArch64::dsub;
2508 else if (AArch64::FPR128RegClass.contains(Reg))
2509 SubRegIdx = AArch64::zsub;
2510 else
2511 continue;
2512 // Clear the bit for the FPR save.
2513 RegBit = false;
2514 // Mark that we should save the corresponding ZPR.
2515 Register ZReg =
2516 TRI->getMatchingSuperReg(Reg, SubRegIdx, &AArch64::ZPRRegClass);
2517 FPRZRegs.set(ZReg);
2518 }
2519 SavedRegs |= FPRZRegs;
2520
2521 AFI->setSplitSVEObjects(true);
2522 LLVM_DEBUG(dbgs() << "SplitSVEObjects enabled!\n");
2523 }
2524}
2525
2527 BitVector &SavedRegs,
2528 RegScavenger *RS) const {
2529 // All calls are tail calls in GHC calling conv, and functions have no
2530 // prologue/epilogue.
2532 return;
2533
2534 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2535
2537 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
2539 Register UnspilledCSGPR;
2540 Register UnspilledCSGPRPaired;
2541
2542 MachineFrameInfo &MFI = MF.getFrameInfo();
2543 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
2544
2545 MCRegister BasePointerReg =
2546 RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister() : MCRegister();
2547
2548 Register ExtraCSSpill;
2549 bool HasUnpairedGPR64 = false;
2550 bool HasPairZReg = false;
2551 BitVector UserReservedRegs = RegInfo->getUserReservedRegs(MF);
2552 BitVector ReservedRegs = RegInfo->getReservedRegs(MF);
2553
2554 // Figure out which callee-saved registers to save/restore.
2555 for (unsigned i = 0; CSRegs[i]; ++i) {
2556 const MCRegister Reg = CSRegs[i];
2557
2558 // Add the base pointer register to SavedRegs if it is callee-save.
2559 if (Reg == BasePointerReg)
2560 SavedRegs.set(Reg);
2561
2562 // Don't save manually reserved registers set through +reserve-x#i,
2563 // even for callee-saved registers, as per GCC's behavior.
2564 if (UserReservedRegs[Reg]) {
2565 SavedRegs.reset(Reg);
2566 continue;
2567 }
2568
2569 bool RegUsed = SavedRegs.test(Reg);
2570 MCRegister PairedReg;
2571 const bool RegIsGPR64 = AArch64::GPR64RegClass.contains(Reg);
2572 if (RegIsGPR64 || AArch64::FPR64RegClass.contains(Reg) ||
2573 AArch64::FPR128RegClass.contains(Reg)) {
2574 // Compensate for odd numbers of GP CSRs.
2575 // For now, all the known cases of odd number of CSRs are of GPRs.
2576 if (HasUnpairedGPR64)
2577 PairedReg = CSRegs[i % 2 == 0 ? i - 1 : i + 1];
2578 else
2579 PairedReg = CSRegs[i ^ 1];
2580 }
2581
2582 // If the function requires all the GP registers to save (SavedRegs),
2583 // and there are an odd number of GP CSRs at the same time (CSRegs),
2584 // PairedReg could be in a different register class from Reg, which would
2585 // lead to a FPR (usually D8) accidentally being marked saved.
2586 if (RegIsGPR64 && !AArch64::GPR64RegClass.contains(PairedReg)) {
2587 PairedReg = Register();
2588 HasUnpairedGPR64 = true;
2589 }
2590 assert(!PairedReg.isValid() ||
2591 AArch64::GPR64RegClass.contains(Reg, PairedReg) ||
2592 AArch64::FPR64RegClass.contains(Reg, PairedReg) ||
2593 AArch64::FPR128RegClass.contains(Reg, PairedReg));
2594
2595 if (!RegUsed) {
2596 if (AArch64::GPR64RegClass.contains(Reg) && !ReservedRegs[Reg]) {
2597 UnspilledCSGPR = Reg;
2598 UnspilledCSGPRPaired = PairedReg;
2599 }
2600 continue;
2601 }
2602
2603 // MachO's compact unwind format relies on all registers being stored in
2604 // pairs.
2605 // FIXME: the usual format is actually better if unwinding isn't needed.
2606 if (producePairRegisters(MF) && PairedReg.isValid() &&
2607 !SavedRegs.test(PairedReg)) {
2608 SavedRegs.set(PairedReg);
2609 if (AArch64::GPR64RegClass.contains(PairedReg) &&
2610 !ReservedRegs[PairedReg])
2611 ExtraCSSpill = PairedReg;
2612 }
2613 // Check if there is a pair of ZRegs, so it can select PReg for spill/fill
2614 HasPairZReg |= (AArch64::ZPRRegClass.contains(Reg, CSRegs[i ^ 1]) &&
2615 SavedRegs.test(CSRegs[i ^ 1]));
2616 }
2617
2618 if (HasPairZReg && enableMultiVectorSpillFill(Subtarget, MF)) {
2620 // Find a suitable predicate register for the multi-vector spill/fill
2621 // instructions.
2622 MCRegister PnReg = findFreePredicateReg(SavedRegs);
2623 if (PnReg.isValid())
2624 AFI->setPredicateRegForFillSpill(PnReg);
2625 // If no free callee-save has been found assign one.
2626 if (!AFI->getPredicateRegForFillSpill() &&
2627 MF.getFunction().getCallingConv() ==
2629 SavedRegs.set(AArch64::P8);
2630 AFI->setPredicateRegForFillSpill(AArch64::PN8);
2631 }
2632
2633 assert(!ReservedRegs[AFI->getPredicateRegForFillSpill()] &&
2634 "Predicate cannot be a reserved register");
2635 }
2636
2638 !Subtarget.isTargetWindows()) {
2639 // For Windows calling convention on a non-windows OS, where X18 is treated
2640 // as reserved, back up X18 when entering non-windows code (marked with the
2641 // Windows calling convention) and restore when returning regardless of
2642 // whether the individual function uses it - it might call other functions
2643 // that clobber it.
2644 SavedRegs.set(AArch64::X18);
2645 }
2646
2647 // Determine if a Hazard slot should be used and where it should go.
2648 // If SplitSVEObjects is used, the hazard padding is placed between the PPRs
2649 // and ZPRs. Otherwise, it goes in the callee save area.
2650 determineStackHazardSlot(MF, SavedRegs);
2651
2652 // Calculates the callee saved stack size.
2653 unsigned CSStackSize = 0;
2654 unsigned ZPRCSStackSize = 0;
2655 unsigned PPRCSStackSize = 0;
2657 for (unsigned Reg : SavedRegs.set_bits()) {
2658 auto *RC = TRI->getMinimalPhysRegClass(MCRegister(Reg));
2659 assert(RC && "expected register class!");
2660 auto SpillSize = TRI->getSpillSize(*RC);
2661 bool IsZPR = AArch64::ZPRRegClass.contains(Reg);
2662 bool IsPPR = !IsZPR && AArch64::PPRRegClass.contains(Reg);
2663 if (IsZPR)
2664 ZPRCSStackSize += SpillSize;
2665 else if (IsPPR)
2666 PPRCSStackSize += SpillSize;
2667 else {
2668 // A register and its super-register can both appear in SavedRegs.
2669 // Only the widest register is actually spilled, so skip such
2670 // sub-registers here to avoid double-counting the overlap.
2671 bool SavedSuper = any_of(TRI->superregs(Reg), [&](MCPhysReg SuperReg) {
2672 return SavedRegs.test(SuperReg);
2673 });
2674 if (!SavedSuper)
2675 CSStackSize += SpillSize;
2676 }
2677 }
2678
2679 // Save number of saved regs, so we can easily update CSStackSize later to
2680 // account for any additional 64-bit GPR saves. Note: After this point
2681 // only 64-bit GPRs can be added to SavedRegs.
2682 unsigned NumSavedRegs = SavedRegs.count();
2683
2684 // If we have hazard padding in the CS area add that to the size.
2686 CSStackSize += getStackHazardSize(MF);
2687
2688 // Increase the callee-saved stack size if the function has streaming mode
2689 // changes, as we will need to spill the value of the VG register.
2690 if (requiresSaveVG(MF))
2691 CSStackSize += 8;
2692
2693 // If we must call __arm_get_current_vg in the prologue preserve the LR.
2694 if (requiresSaveVG(MF) && !Subtarget.hasSVE())
2695 SavedRegs.set(AArch64::LR);
2696
2697 // The frame record needs to be created by saving the appropriate registers
2698 uint64_t EstimatedStackSize = MFI.estimateStackSize(MF);
2699 if (hasFP(MF) ||
2700 windowsRequiresStackProbe(MF, EstimatedStackSize + CSStackSize + 16)) {
2701 SavedRegs.set(AArch64::FP);
2702 SavedRegs.set(AArch64::LR);
2703 }
2704
2705 LLVM_DEBUG({
2706 dbgs() << "*** determineCalleeSaves\nSaved CSRs:";
2707 for (unsigned Reg : SavedRegs.set_bits())
2708 dbgs() << ' ' << printReg(MCRegister(Reg), RegInfo);
2709 dbgs() << "\n";
2710 });
2711
2712 // If any callee-saved registers are used, the frame cannot be eliminated.
2713 auto [ZPRLocalStackSize, PPRLocalStackSize] =
2715 uint64_t SVELocals = ZPRLocalStackSize + PPRLocalStackSize;
2716 uint64_t SVEStackSize =
2717 alignTo(ZPRCSStackSize + PPRCSStackSize + SVELocals, 16);
2718 bool CanEliminateFrame = (SavedRegs.count() == 0) && !SVEStackSize;
2719
2720 // The CSR spill slots have not been allocated yet, so estimateStackSize
2721 // won't include them.
2722 unsigned EstimatedStackSizeLimit = estimateRSStackSizeLimit(MF);
2723
2724 // We may address some of the stack above the canonical frame address, either
2725 // for our own arguments or during a call. Include that in calculating whether
2726 // we have complicated addressing concerns.
2727 int64_t CalleeStackUsed = 0;
2728 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I) {
2729 int64_t FixedOff = MFI.getObjectOffset(I);
2730 if (FixedOff > CalleeStackUsed)
2731 CalleeStackUsed = FixedOff;
2732 }
2733
2734 // Conservatively always assume BigStack when there are SVE spills.
2735 bool BigStack = SVEStackSize || (EstimatedStackSize + CSStackSize +
2736 CalleeStackUsed) > EstimatedStackSizeLimit;
2737 if (BigStack || !CanEliminateFrame || RegInfo->cannotEliminateFrame(MF))
2738 AFI->setHasStackFrame(true);
2739
2740 // Estimate if we might need to scavenge a register at some point in order
2741 // to materialize a stack offset. If so, either spill one additional
2742 // callee-saved register or reserve a special spill slot to facilitate
2743 // register scavenging. If we already spilled an extra callee-saved register
2744 // above to keep the number of spills even, we don't need to do anything else
2745 // here.
2746 if (BigStack) {
2747 if (!ExtraCSSpill.isValid() && UnspilledCSGPR.isValid()) {
2748 LLVM_DEBUG(dbgs() << "Spilling " << printReg(UnspilledCSGPR, RegInfo)
2749 << " to get a scratch register.\n");
2750 SavedRegs.set(UnspilledCSGPR);
2751 ExtraCSSpill = UnspilledCSGPR;
2752
2753 // MachO's compact unwind format relies on all registers being stored in
2754 // pairs, so if we need to spill one extra for BigStack, then we need to
2755 // store the pair.
2756 if (producePairRegisters(MF)) {
2757 if (!UnspilledCSGPRPaired.isValid()) {
2758 // Failed to make a pair for compact unwind format, revert spilling.
2759 if (produceCompactUnwindFrame(*this, MF)) {
2760 SavedRegs.reset(UnspilledCSGPR);
2761 ExtraCSSpill = Register();
2762 }
2763 } else
2764 SavedRegs.set(UnspilledCSGPRPaired);
2765 }
2766 }
2767
2768 // If we didn't find an extra callee-saved register to spill, create
2769 // an emergency spill slot.
2770 if (!ExtraCSSpill.isValid() ||
2771 MF.getRegInfo().isPhysRegUsed(ExtraCSSpill)) {
2773 const TargetRegisterClass &RC = AArch64::GPR64RegClass;
2774 unsigned Size = TRI->getSpillSize(RC);
2775 Align Alignment = TRI->getSpillAlign(RC);
2776 int FI = MFI.CreateSpillStackObject(Size, Alignment);
2777 RS->addScavengingFrameIndex(FI);
2778 LLVM_DEBUG(dbgs() << "No available CS registers, allocated fi#" << FI
2779 << " as the emergency spill slot.\n");
2780 }
2781 }
2782
2783 // Adding the size of additional 64bit GPR saves.
2784 CSStackSize += 8 * (SavedRegs.count() - NumSavedRegs);
2785
2786 // A Swift asynchronous context extends the frame record with a pointer
2787 // directly before FP.
2788 if (hasFP(MF) && AFI->hasSwiftAsyncContext())
2789 CSStackSize += 8;
2790
2791 uint64_t AlignedCSStackSize = alignTo(CSStackSize, 16);
2792 LLVM_DEBUG(dbgs() << "Estimated stack frame size: "
2793 << EstimatedStackSize + AlignedCSStackSize << " bytes.\n");
2794
2796 AFI->getCalleeSavedStackSize() == AlignedCSStackSize) &&
2797 "Should not invalidate callee saved info");
2798
2799 // Round up to register pair alignment to avoid additional SP adjustment
2800 // instructions.
2801 AFI->setCalleeSavedStackSize(AlignedCSStackSize);
2802 AFI->setCalleeSaveStackHasFreeSpace(AlignedCSStackSize != CSStackSize);
2803 AFI->setSVECalleeSavedStackSize(ZPRCSStackSize, alignTo(PPRCSStackSize, 16));
2804}
2805
2808 std::vector<CalleeSavedInfo> &CSI) {
2809 // Reorder callee-saved ZPRs to maximize pairing which requires
2810 // consecutive even/odd registers at even scaled stack offsets.
2811 // Additional requirements are checked when the register pairs are formed.
2812 assert(!isTargetWindows(MF) &&
2813 "ZPR callee-save reordering not supported on Windows");
2814
2815 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2816 if (!AFI->getPredicateRegForFillSpill())
2817 return;
2818
2820 SmallVector<size_t> ZPRPositions;
2821
2822 for (auto [Index, CS] : llvm::enumerate(CSI)) {
2823 if (AArch64::ZPRRegClass.contains(CS.getReg())) {
2824 ZPRSaves.push_back(CS);
2825 ZPRPositions.push_back(Index);
2826 }
2827 }
2828
2829 if (ZPRSaves.size() < 2)
2830 return;
2831
2832 llvm::sort(ZPRSaves, [](const auto &A, const auto &B) {
2833 return A.getReg() < B.getReg();
2834 });
2835
2838 for (size_t i = 0; i < ZPRSaves.size();) {
2839 if (i + 1 < ZPRSaves.size() &&
2840 (ZPRSaves[i].getReg() + 1 == ZPRSaves[i + 1].getReg()) &&
2841 (ZPRSaves[i].getReg() - AArch64::Z0) % 2 == 0) {
2842 Pairs.emplace_back(ZPRSaves[i], ZPRSaves[i + 1]);
2843 i += 2;
2844 } else {
2845 Singles.push_back(ZPRSaves[i++]);
2846 }
2847 }
2848
2849 // If the lowest offset is odd, select one register to spill here
2850 // so subsequent pairs begin at an even offset.
2851 int ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
2852 if (!AFI->hasSplitSVEObjects())
2853 ZPRByteOffset += AFI->getPPRCalleeSavedStackSize();
2854
2855 const int Scale = RegInfo->getSpillSize(AArch64::ZPRRegClass);
2856 const int LowestOffset =
2857 (ZPRByteOffset / Scale) - static_cast<int>(ZPRSaves.size());
2858
2859 // Prefer the highest single otherwise split the highest pair.
2860 std::optional<CalleeSavedInfo> AlignmentSingle;
2861 if (LowestOffset % 2 != 0) {
2862 if (!Singles.empty()) {
2863 AlignmentSingle = Singles.pop_back_val();
2864 } else {
2865 assert(!Pairs.empty() && "Expected a ZPR pair to split");
2866 auto [Even, Odd] = Pairs.pop_back_val();
2867 AlignmentSingle = Odd;
2868 Singles.push_back(Even);
2869 }
2870 }
2871
2872 if (Pairs.empty())
2873 return;
2874
2875 // Build ZPRs so reverse spill emission processes the leading single first,
2876 // followed by the candidate even/odd pairs and remaining singles.
2877 SmallVector<CalleeSavedInfo> ZPRSavesInCSIOrder;
2878 llvm::append_range(ZPRSavesInCSIOrder, Singles);
2879
2880 for (const auto &[Even, Odd] : Pairs) {
2881 ZPRSavesInCSIOrder.push_back(Odd);
2882 ZPRSavesInCSIOrder.push_back(Even);
2883 }
2884
2885 if (AlignmentSingle)
2886 ZPRSavesInCSIOrder.push_back(*AlignmentSingle);
2887
2888 assert(ZPRSavesInCSIOrder.size() == ZPRPositions.size() &&
2889 "Reordering should not change the number of ZPR spills");
2890 for (auto [Position, CS] : llvm::zip(ZPRPositions, ZPRSavesInCSIOrder))
2891 CSI[Position] = CS;
2892}
2893
2895 MachineFunction &MF, const TargetRegisterInfo *RegInfo,
2896 std::vector<CalleeSavedInfo> &CSI) const {
2897 bool IsWindows = isTargetWindows(MF);
2898 unsigned StackHazardSize = getStackHazardSize(MF);
2899 // To match the canonical windows frame layout, reverse the list of
2900 // callee saved registers to get them laid out by PrologEpilogInserter
2901 // in the right order. (PrologEpilogInserter allocates stack objects top
2902 // down. Windows canonical prologs store higher numbered registers at
2903 // the top, thus have the CSI array start from the highest registers.)
2904 if (IsWindows)
2905 std::reverse(CSI.begin(), CSI.end());
2906
2907 if (CSI.empty())
2908 return true; // Early exit if no callee saved registers are modified!
2909
2910 // Now that we know which registers need to be saved and restored, allocate
2911 // stack slots for them.
2912 MachineFrameInfo &MFI = MF.getFrameInfo();
2913 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2914
2915 // Insert VG into the list of CSRs, immediately before LR if saved.
2916 if (requiresSaveVG(MF)) {
2917 CalleeSavedInfo VGInfo(AArch64::VG);
2918 auto It =
2919 find_if(CSI, [](auto &Info) { return Info.getReg() == AArch64::LR; });
2920 if (It != CSI.end())
2921 CSI.insert(It, VGInfo);
2922 else
2923 CSI.push_back(VGInfo);
2924 }
2925
2926 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2927 if (!IsWindows && enableMultiVectorSpillFill(Subtarget, MF))
2928 // The Windows stack layout is not supported by this reordering function
2929 // yet.
2930 orderZPRCalleeSavesForPairs(MF, RegInfo, CSI);
2931
2932 Register LastReg = 0;
2933 int HazardSlotIndex = std::numeric_limits<int>::max();
2934 for (auto &CS : CSI) {
2935 MCRegister Reg = CS.getReg();
2936 const TargetRegisterClass *RC = RegInfo->getMinimalPhysRegClass(Reg);
2937
2938 // Create a hazard slot as we switch between GPR and FPR CSRs.
2940 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
2942 assert(HazardSlotIndex == std::numeric_limits<int>::max() &&
2943 "Unexpected register order for hazard slot");
2944 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2945 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2946 << "\n");
2947 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2948 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2949 }
2950
2951 unsigned Size = RegInfo->getSpillSize(*RC);
2952 Align Alignment(RegInfo->getSpillAlign(*RC));
2953 int FrameIdx = MFI.CreateStackObject(Size, Alignment, true);
2954 CS.setFrameIdx(FrameIdx);
2955 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2956
2957 // Grab 8 bytes below FP for the extended asynchronous frame info.
2958 if (hasFP(MF) && AFI->hasSwiftAsyncContext() && Reg == AArch64::FP) {
2959 FrameIdx = MFI.CreateStackObject(8, Alignment, true);
2960 AFI->setSwiftAsyncContextFrameIdx(FrameIdx);
2961 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2962 }
2963 LastReg = Reg;
2964 }
2965
2966 // Add hazard slot in the case where no FPR CSRs are present.
2968 HazardSlotIndex == std::numeric_limits<int>::max()) {
2969 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2970 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2971 << "\n");
2972 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2973 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2974 }
2975
2976 return true;
2977}
2978
2980 const MachineFunction &MF) const {
2982 // If the function has streaming-mode changes, don't scavenge a
2983 // spillslot in the callee-save area, as that might require an
2984 // 'addvl' in the streaming-mode-changing call-sequence when the
2985 // function doesn't use a FP.
2986 if (AFI->hasStreamingModeChanges() && !hasFP(MF))
2987 return false;
2988 // Don't allow register salvaging with hazard slots, in case it moves objects
2989 // into the wrong place.
2990 if (AFI->hasStackHazardSlotIndex())
2991 return false;
2992 return AFI->hasCalleeSaveStackFreeSpace();
2993}
2994
2995/// returns true if there are any SVE callee saves.
2997 int &Min, int &Max) {
2998 Min = std::numeric_limits<int>::max();
2999 Max = std::numeric_limits<int>::min();
3000
3001 if (!MFI.isCalleeSavedInfoValid())
3002 return false;
3003
3004 const std::vector<CalleeSavedInfo> &CSI = MFI.getCalleeSavedInfo();
3005 for (auto &CS : CSI) {
3006 if (AArch64::ZPRRegClass.contains(CS.getReg()) ||
3007 AArch64::PPRRegClass.contains(CS.getReg())) {
3008 assert((Max == std::numeric_limits<int>::min() ||
3009 Max + 1 == CS.getFrameIdx()) &&
3010 "SVE CalleeSaves are not consecutive");
3011 Min = std::min(Min, CS.getFrameIdx());
3012 Max = std::max(Max, CS.getFrameIdx());
3013 }
3014 }
3015 return Min != std::numeric_limits<int>::max();
3016}
3017
3019 AssignObjectOffsets AssignOffsets) {
3020 MachineFrameInfo &MFI = MF.getFrameInfo();
3021 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3022
3023 SVEStackSizes SVEStack{};
3024
3025 // With SplitSVEObjects we maintain separate stack offsets for predicates
3026 // (PPRs) and SVE vectors (ZPRs). When SplitSVEObjects is disabled predicates
3027 // are included in the SVE vector area.
3028 uint64_t &ZPRStackTop = SVEStack.ZPRStackSize;
3029 uint64_t &PPRStackTop =
3030 AFI->hasSplitSVEObjects() ? SVEStack.PPRStackSize : SVEStack.ZPRStackSize;
3031
3032#ifndef NDEBUG
3033 // First process all fixed stack objects.
3034 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I)
3035 assert(!MFI.hasScalableStackID(I) &&
3036 "SVE vectors should never be passed on the stack by value, only by "
3037 "reference.");
3038#endif
3039
3040 auto AllocateObject = [&](int FI) {
3042 ? ZPRStackTop
3043 : PPRStackTop;
3044
3045 // FIXME: Given that the length of SVE vectors is not necessarily a power of
3046 // two, we'd need to align every object dynamically at runtime if the
3047 // alignment is larger than 16. This is not yet supported.
3048 Align Alignment = MFI.getObjectAlign(FI);
3049 if (Alignment > Align(16))
3051 "Alignment of scalable vectors > 16 bytes is not yet supported");
3052
3053 StackTop += MFI.getObjectSize(FI);
3054 StackTop = alignTo(StackTop, Alignment);
3055
3056 assert(StackTop < (uint64_t)std::numeric_limits<int64_t>::max() &&
3057 "SVE StackTop far too large?!");
3058
3059 int64_t Offset = -int64_t(StackTop);
3060 if (AssignOffsets == AssignObjectOffsets::Yes)
3061 MFI.setObjectOffset(FI, Offset);
3062
3063 LLVM_DEBUG(dbgs() << "alloc FI(" << FI << ") at SP[" << Offset << "]\n");
3064 };
3065
3066 // Then process all callee saved slots.
3067 int MinCSFrameIndex, MaxCSFrameIndex;
3068 if (getSVECalleeSaveSlotRange(MFI, MinCSFrameIndex, MaxCSFrameIndex)) {
3069 for (int FI = MinCSFrameIndex; FI <= MaxCSFrameIndex; ++FI)
3070 AllocateObject(FI);
3071 }
3072
3073 // Ensure the CS area is 16-byte aligned.
3074 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3075 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3076
3077 // Create a buffer of SVE objects to allocate and sort it.
3078 SmallVector<int, 8> ObjectsToAllocate;
3079 // If we have a stack protector, and we've previously decided that we have SVE
3080 // objects on the stack and thus need it to go in the SVE stack area, then it
3081 // needs to go first.
3082 int StackProtectorFI = -1;
3083 if (MFI.hasStackProtectorIndex()) {
3084 StackProtectorFI = MFI.getStackProtectorIndex();
3085 if (MFI.getStackID(StackProtectorFI) == TargetStackID::ScalableVector)
3086 ObjectsToAllocate.push_back(StackProtectorFI);
3087 }
3088
3089 for (int FI = 0, E = MFI.getObjectIndexEnd(); FI != E; ++FI) {
3090 if (FI == StackProtectorFI || MFI.isDeadObjectIndex(FI) ||
3092 continue;
3093
3096 continue;
3097
3098 ObjectsToAllocate.push_back(FI);
3099 }
3100
3101 // Allocate all SVE locals and spills
3102 for (unsigned FI : ObjectsToAllocate)
3103 AllocateObject(FI);
3104
3105 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3106 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3107
3108 if (AssignOffsets == AssignObjectOffsets::Yes)
3109 AFI->setStackSizeSVE(SVEStack.ZPRStackSize, SVEStack.PPRStackSize);
3110
3111 return SVEStack;
3112}
3113
3115 MachineFunction &MF, RegScavenger *RS) const {
3117 "Upwards growing stack unsupported");
3118
3120
3121 // If this function isn't doing Win64-style C++ EH, we don't need to do
3122 // anything.
3123 if (!MF.hasEHFunclets())
3124 return;
3125
3126 MachineFrameInfo &MFI = MF.getFrameInfo();
3127 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3128
3129 // Win64 C++ EH needs to allocate space for the catch objects in the fixed
3130 // object area right next to the UnwindHelp object.
3131 WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
3132 int64_t CurrentOffset =
3134 for (WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
3135 for (WinEHHandlerType &H : TBME.HandlerArray) {
3136 int FrameIndex = H.CatchObj.FrameIndex;
3137 if ((FrameIndex != INT_MAX) && MFI.getObjectOffset(FrameIndex) == 0) {
3138 CurrentOffset =
3139 alignTo(CurrentOffset, MFI.getObjectAlign(FrameIndex).value());
3140 CurrentOffset += MFI.getObjectSize(FrameIndex);
3141 MFI.setObjectOffset(FrameIndex, -CurrentOffset);
3142 }
3143 }
3144 }
3145
3146 // Create an UnwindHelp object.
3147 // The UnwindHelp object is allocated at the start of the fixed object area
3148 int64_t UnwindHelpOffset = alignTo(CurrentOffset + 8, Align(16));
3149 assert(UnwindHelpOffset == getFixedObjectSize(MF, AFI, /*IsWin64*/ true,
3150 /*IsFunclet*/ false) &&
3151 "UnwindHelpOffset must be at the start of the fixed object area");
3152 int UnwindHelpFI = MFI.CreateFixedObject(/*Size*/ 8, -UnwindHelpOffset,
3153 /*IsImmutable=*/false);
3154 EHInfo.UnwindHelpFrameIdx = UnwindHelpFI;
3155
3156 MachineBasicBlock &MBB = MF.front();
3157 auto MBBI = MBB.begin();
3158 while (MBBI != MBB.end() && MBBI->getFlag(MachineInstr::FrameSetup))
3159 ++MBBI;
3160
3161 // We need to store -2 into the UnwindHelp object at the start of the
3162 // function.
3163 DebugLoc DL;
3164 RS->enterBasicBlockEnd(MBB);
3165 RS->backward(MBBI);
3166 Register DstReg = RS->FindUnusedReg(&AArch64::GPR64commonRegClass);
3167 assert(DstReg && "There must be a free register after frame setup");
3168 const AArch64InstrInfo &TII =
3169 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3170 BuildMI(MBB, MBBI, DL, TII.get(AArch64::MOVi64imm), DstReg).addImm(-2);
3171 BuildMI(MBB, MBBI, DL, TII.get(AArch64::STURXi))
3172 .addReg(DstReg, getKillRegState(true))
3173 .addFrameIndex(UnwindHelpFI)
3174 .addImm(0);
3175}
3176
3177namespace {
3178struct TagStoreInstr {
3180 int64_t Offset, Size;
3181 explicit TagStoreInstr(MachineInstr *MI, int64_t Offset, int64_t Size)
3182 : MI(MI), Offset(Offset), Size(Size) {}
3183};
3184
3185class TagStoreEdit {
3186 MachineFunction *MF;
3187 MachineBasicBlock *MBB;
3188 MachineRegisterInfo *MRI;
3189 // Tag store instructions that are being replaced.
3191 // Combined memref arguments of the above instructions.
3193
3194 // Replace allocation tags in [FrameReg + FrameRegOffset, FrameReg +
3195 // FrameRegOffset + Size) with the address tag of SP.
3196 Register FrameReg;
3197 StackOffset FrameRegOffset;
3198 int64_t Size;
3199 // If not std::nullopt, move FrameReg to (FrameReg + FrameRegUpdate) at the
3200 // end.
3201 std::optional<int64_t> FrameRegUpdate;
3202 // MIFlags for any FrameReg updating instructions.
3203 unsigned FrameRegUpdateFlags;
3204
3205 // Use zeroing instruction variants.
3206 bool ZeroData;
3207 DebugLoc DL;
3208
3209 void emitUnrolled(MachineBasicBlock::iterator InsertI);
3210 void emitLoop(MachineBasicBlock::iterator InsertI);
3211
3212public:
3213 TagStoreEdit(MachineBasicBlock *MBB, bool ZeroData)
3214 : MBB(MBB), ZeroData(ZeroData) {
3215 MF = MBB->getParent();
3216 MRI = &MF->getRegInfo();
3217 }
3218 // Add an instruction to be replaced. Instructions must be added in the
3219 // ascending order of Offset, and have to be adjacent.
3220 void addInstruction(TagStoreInstr I) {
3221 assert((TagStores.empty() ||
3222 TagStores.back().Offset + TagStores.back().Size == I.Offset) &&
3223 "Non-adjacent tag store instructions.");
3224 TagStores.push_back(I);
3225 }
3226 void clear() { TagStores.clear(); }
3227 // Emit equivalent code at the given location, and erase the current set of
3228 // instructions. May skip if the replacement is not profitable. May invalidate
3229 // the input iterator and replace it with a valid one.
3230 void emitCode(MachineBasicBlock::iterator &InsertI,
3231 const AArch64FrameLowering *TFI, bool TryMergeSPUpdate);
3232};
3233
3234void TagStoreEdit::emitUnrolled(MachineBasicBlock::iterator InsertI) {
3235 const AArch64InstrInfo *TII =
3236 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3237
3238 const int64_t kMinOffset = -256 * 16;
3239 const int64_t kMaxOffset = 255 * 16;
3240
3241 Register BaseReg = FrameReg;
3242 int64_t BaseRegOffsetBytes = FrameRegOffset.getFixed();
3243 if (BaseRegOffsetBytes < kMinOffset ||
3244 BaseRegOffsetBytes + (Size - Size % 32) > kMaxOffset ||
3245 // BaseReg can be FP, which is not necessarily aligned to 16-bytes. In
3246 // that case, BaseRegOffsetBytes will not be aligned to 16 bytes, which
3247 // is required for the offset of ST2G.
3248 BaseRegOffsetBytes % 16 != 0) {
3249 Register ScratchReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3250 emitFrameOffset(*MBB, InsertI, DL, ScratchReg, BaseReg,
3251 StackOffset::getFixed(BaseRegOffsetBytes), TII);
3252 BaseReg = ScratchReg;
3253 BaseRegOffsetBytes = 0;
3254 }
3255
3256 MachineInstr *LastI = nullptr;
3257 while (Size) {
3258 int64_t InstrSize = (Size > 16) ? 32 : 16;
3259 unsigned Opcode =
3260 InstrSize == 16
3261 ? (ZeroData ? AArch64::STZGi : AArch64::STGi)
3262 : (ZeroData ? AArch64::STZ2Gi : AArch64::ST2Gi);
3263 assert(BaseRegOffsetBytes % 16 == 0);
3264 MachineInstr *I = BuildMI(*MBB, InsertI, DL, TII->get(Opcode))
3265 .addReg(AArch64::SP)
3266 .addReg(BaseReg)
3267 .addImm(BaseRegOffsetBytes / 16)
3268 .setMemRefs(CombinedMemRefs);
3269 // A store to [BaseReg, #0] should go last for an opportunity to fold the
3270 // final SP adjustment in the epilogue.
3271 if (BaseRegOffsetBytes == 0)
3272 LastI = I;
3273 BaseRegOffsetBytes += InstrSize;
3274 Size -= InstrSize;
3275 }
3276
3277 if (LastI)
3278 MBB->splice(InsertI, MBB, LastI);
3279}
3280
3281void TagStoreEdit::emitLoop(MachineBasicBlock::iterator InsertI) {
3282 const AArch64InstrInfo *TII =
3283 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3284
3285 Register BaseReg = FrameRegUpdate
3286 ? FrameReg
3287 : MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3288 Register SizeReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3289
3290 emitFrameOffset(*MBB, InsertI, DL, BaseReg, FrameReg, FrameRegOffset, TII);
3291
3292 int64_t LoopSize = Size;
3293 // If the loop size is not a multiple of 32, split off one 16-byte store at
3294 // the end to fold BaseReg update into.
3295 if (FrameRegUpdate && *FrameRegUpdate)
3296 LoopSize -= LoopSize % 32;
3297 MachineInstr *LoopI = BuildMI(*MBB, InsertI, DL,
3298 TII->get(ZeroData ? AArch64::STZGloop_wback
3299 : AArch64::STGloop_wback))
3300 .addDef(SizeReg)
3301 .addDef(BaseReg)
3302 .addImm(LoopSize)
3303 .addReg(BaseReg)
3304 .setMemRefs(CombinedMemRefs);
3305 if (FrameRegUpdate)
3306 LoopI->setFlags(FrameRegUpdateFlags);
3307
3308 int64_t ExtraBaseRegUpdate =
3309 FrameRegUpdate ? (*FrameRegUpdate - FrameRegOffset.getFixed() - Size) : 0;
3310 LLVM_DEBUG(dbgs() << "TagStoreEdit::emitLoop: LoopSize=" << LoopSize
3311 << ", Size=" << Size
3312 << ", ExtraBaseRegUpdate=" << ExtraBaseRegUpdate
3313 << ", FrameRegUpdate=" << FrameRegUpdate
3314 << ", FrameRegOffset.getFixed()="
3315 << FrameRegOffset.getFixed() << "\n");
3316 if (LoopSize < Size) {
3317 assert(FrameRegUpdate);
3318 assert(Size - LoopSize == 16);
3319 // Tag 16 more bytes at BaseReg and update BaseReg.
3320 int64_t STGOffset = ExtraBaseRegUpdate + 16;
3321 assert(STGOffset % 16 == 0 && STGOffset >= -4096 && STGOffset <= 4080 &&
3322 "STG immediate out of range");
3323 BuildMI(*MBB, InsertI, DL,
3324 TII->get(ZeroData ? AArch64::STZGPostIndex : AArch64::STGPostIndex))
3325 .addDef(BaseReg)
3326 .addReg(BaseReg)
3327 .addReg(BaseReg)
3328 .addImm(STGOffset / 16)
3329 .setMemRefs(CombinedMemRefs)
3330 .setMIFlags(FrameRegUpdateFlags);
3331 } else if (ExtraBaseRegUpdate) {
3332 // Update BaseReg.
3333 int64_t AddSubOffset = std::abs(ExtraBaseRegUpdate);
3334 assert(AddSubOffset <= 4095 && "ADD/SUB immediate out of range");
3335 BuildMI(
3336 *MBB, InsertI, DL,
3337 TII->get(ExtraBaseRegUpdate > 0 ? AArch64::ADDXri : AArch64::SUBXri))
3338 .addDef(BaseReg)
3339 .addReg(BaseReg)
3340 .addImm(AddSubOffset)
3341 .addImm(0)
3342 .setMIFlags(FrameRegUpdateFlags);
3343 }
3344}
3345
3346// Check if *II is a register update that can be merged into STGloop that ends
3347// at (Reg + Size). RemainingOffset is the required adjustment to Reg after the
3348// end of the loop.
3349bool canMergeRegUpdate(MachineBasicBlock::iterator II, unsigned Reg,
3350 int64_t Size, int64_t *TotalOffset) {
3351 MachineInstr &MI = *II;
3352 if ((MI.getOpcode() == AArch64::ADDXri ||
3353 MI.getOpcode() == AArch64::SUBXri) &&
3354 MI.getOperand(0).getReg() == Reg && MI.getOperand(1).getReg() == Reg) {
3355 unsigned Shift = AArch64_AM::getShiftValue(MI.getOperand(3).getImm());
3356 int64_t Offset = MI.getOperand(2).getImm() << Shift;
3357 if (MI.getOpcode() == AArch64::SUBXri)
3358 Offset = -Offset;
3359 int64_t PostOffset = Offset - Size;
3360 // TagStoreEdit::emitLoop might emit either an ADD/SUB after the loop, or
3361 // an STGPostIndex which does the last 16 bytes of tag write. Which one is
3362 // chosen depends on the alignment of the loop size, but the difference
3363 // between the valid ranges for the two instructions is small, so we
3364 // conservatively assume that it could be either case here.
3365 //
3366 // Max offset of STGPostIndex, minus the 16 byte tag write folded into that
3367 // instruction.
3368 const int64_t kMaxOffset = 4080 - 16;
3369 // Max offset of SUBXri.
3370 const int64_t kMinOffset = -4095;
3371 if (PostOffset <= kMaxOffset && PostOffset >= kMinOffset &&
3372 PostOffset % 16 == 0) {
3373 *TotalOffset = Offset;
3374 return true;
3375 }
3376 }
3377 return false;
3378}
3379
3380void mergeMemRefs(const SmallVectorImpl<TagStoreInstr> &TSE,
3382 MemRefs.clear();
3383 for (auto &TS : TSE) {
3384 MachineInstr *MI = TS.MI;
3385 // An instruction without memory operands may access anything. Be
3386 // conservative and return an empty list.
3387 if (MI->memoperands_empty()) {
3388 MemRefs.clear();
3389 return;
3390 }
3391 MemRefs.append(MI->memoperands_begin(), MI->memoperands_end());
3392 }
3393}
3394
3395void TagStoreEdit::emitCode(MachineBasicBlock::iterator &InsertI,
3396 const AArch64FrameLowering *TFI,
3397 bool TryMergeSPUpdate) {
3398 if (TagStores.empty())
3399 return;
3400 TagStoreInstr &FirstTagStore = TagStores[0];
3401 TagStoreInstr &LastTagStore = TagStores[TagStores.size() - 1];
3402 Size = LastTagStore.Offset - FirstTagStore.Offset + LastTagStore.Size;
3403 DL = TagStores[0].MI->getDebugLoc();
3404
3405 Register Reg;
3406 FrameRegOffset = TFI->resolveFrameOffsetReference(
3407 *MF, FirstTagStore.Offset, false /*isFixed*/,
3408 TargetStackID::Default /*StackID*/, Reg,
3409 /*PreferFP=*/false, /*ForSimm=*/true);
3410 FrameReg = Reg;
3411 FrameRegUpdate = std::nullopt;
3412
3413 mergeMemRefs(TagStores, CombinedMemRefs);
3414
3415 LLVM_DEBUG({
3416 dbgs() << "Replacing adjacent STG instructions:\n";
3417 for (const auto &Instr : TagStores) {
3418 dbgs() << " " << *Instr.MI;
3419 }
3420 });
3421
3422 // Size threshold where a loop becomes shorter than a linear sequence of
3423 // tagging instructions.
3424 const int kSetTagLoopThreshold = 176;
3425 if (Size < kSetTagLoopThreshold) {
3426 if (TagStores.size() < 2)
3427 return;
3428 emitUnrolled(InsertI);
3429 } else {
3430 MachineInstr *UpdateInstr = nullptr;
3431 int64_t TotalOffset = 0;
3432 if (TryMergeSPUpdate) {
3433 // See if we can merge base register update into the STGloop.
3434 // This is done in AArch64LoadStoreOptimizer for "normal" stores,
3435 // but STGloop is way too unusual for that, and also it only
3436 // realistically happens in function epilogue. Also, STGloop is expanded
3437 // before that pass.
3438 if (InsertI != MBB->end() &&
3439 canMergeRegUpdate(InsertI, FrameReg, FrameRegOffset.getFixed() + Size,
3440 &TotalOffset)) {
3441 UpdateInstr = &*InsertI++;
3442 LLVM_DEBUG(dbgs() << "Folding SP update into loop:\n "
3443 << *UpdateInstr);
3444 }
3445 }
3446
3447 if (!UpdateInstr && TagStores.size() < 2)
3448 return;
3449
3450 if (UpdateInstr) {
3451 FrameRegUpdate = TotalOffset;
3452 FrameRegUpdateFlags = UpdateInstr->getFlags();
3453 }
3454 emitLoop(InsertI);
3455 if (UpdateInstr)
3456 UpdateInstr->eraseFromParent();
3457 }
3458
3459 for (auto &TS : TagStores)
3460 TS.MI->eraseFromParent();
3461}
3462
3463bool isMergeableStackTaggingInstruction(MachineInstr &MI, int64_t &Offset,
3464 int64_t &Size, bool &ZeroData) {
3465 MachineFunction &MF = *MI.getParent()->getParent();
3466 const MachineFrameInfo &MFI = MF.getFrameInfo();
3467
3468 unsigned Opcode = MI.getOpcode();
3469 ZeroData = (Opcode == AArch64::STZGloop || Opcode == AArch64::STZGi ||
3470 Opcode == AArch64::STZ2Gi);
3471
3472 if (Opcode == AArch64::STGloop || Opcode == AArch64::STZGloop) {
3473 if (!MI.getOperand(0).isDead() || !MI.getOperand(1).isDead())
3474 return false;
3475 if (!MI.getOperand(2).isImm() || !MI.getOperand(3).isFI())
3476 return false;
3477 Offset = MFI.getObjectOffset(MI.getOperand(3).getIndex());
3478 Size = MI.getOperand(2).getImm();
3479 return true;
3480 }
3481
3482 if (Opcode == AArch64::STGi || Opcode == AArch64::STZGi)
3483 Size = 16;
3484 else if (Opcode == AArch64::ST2Gi || Opcode == AArch64::STZ2Gi)
3485 Size = 32;
3486 else
3487 return false;
3488
3489 if (MI.getOperand(0).getReg() != AArch64::SP || !MI.getOperand(1).isFI())
3490 return false;
3491
3492 Offset = MFI.getObjectOffset(MI.getOperand(1).getIndex()) +
3493 16 * MI.getOperand(2).getImm();
3494 return true;
3495}
3496
3497static size_t countAvailableScavengerSlots(LivePhysRegs &LiveRegs,
3499 RegScavenger *RS) {
3500 auto FreeGPRs =
3501 llvm::count_if(AArch64::GPR64RegClass, [&LiveRegs, &MRI](auto Reg) {
3502 return LiveRegs.available(MRI, Reg);
3503 });
3504
3505 size_t NumEmergencySlots = 0;
3506 if (RS)
3507 NumEmergencySlots = RS->getNumScavengingFrameIndices();
3508
3509 return FreeGPRs + NumEmergencySlots;
3510}
3511
3512// Detect a run of memory tagging instructions for adjacent stack frame slots,
3513// and replace them with a shorter instruction sequence:
3514// * replace STG + STG with ST2G
3515// * replace STGloop + STGloop with STGloop
3516// This code needs to run when stack slot offsets are already known, but before
3517// FrameIndex operands in STG instructions are eliminated.
3519 const AArch64FrameLowering *TFI,
3520 RegScavenger *RS) {
3521 bool FirstZeroData;
3522 int64_t Size, Offset;
3523 MachineInstr &MI = *II;
3526 if (&MI == &MBB->instr_back())
3527 return II;
3528 if (!isMergeableStackTaggingInstruction(MI, Offset, Size, FirstZeroData))
3529 return II;
3530
3532 Instrs.emplace_back(&MI, Offset, Size);
3533
3534 constexpr int kScanLimit = 10;
3535 int Count = 0;
3537 NextI != E && Count < kScanLimit; ++NextI) {
3538 MachineInstr &MI = *NextI;
3539 bool ZeroData;
3540 int64_t Size, Offset;
3541 // Collect instructions that update memory tags with a FrameIndex operand
3542 // and (when applicable) constant size, and whose output registers are dead
3543 // (the latter is almost always the case in practice). Since these
3544 // instructions effectively have no inputs or outputs, we are free to skip
3545 // any non-aliasing instructions in between without tracking used registers.
3546 if (isMergeableStackTaggingInstruction(MI, Offset, Size, ZeroData)) {
3547 if (ZeroData != FirstZeroData)
3548 break;
3549 Instrs.emplace_back(&MI, Offset, Size);
3550 continue;
3551 }
3552
3553 // Only count non-transient, non-tagging instructions toward the scan
3554 // limit.
3555 if (!MI.isTransient())
3556 ++Count;
3557
3558 // Just in case, stop before the epilogue code starts.
3559 if (MI.getFlag(MachineInstr::FrameSetup) ||
3561 break;
3562
3563 // Reject anything that may alias the collected instructions.
3564 if (MI.mayLoadOrStore() || MI.hasUnmodeledSideEffects() || MI.isCall())
3565 break;
3566 }
3567
3568 // New code will be inserted after the last tagging instruction we've found.
3569 MachineBasicBlock::iterator InsertI = Instrs.back().MI;
3570
3571 // All the gathered stack tag instructions are merged and placed after
3572 // last tag store in the list. The check should be made if the nzcv
3573 // flag is live at the point where we are trying to insert. Otherwise
3574 // the nzcv flag might get clobbered if any stg loops are present.
3575
3576 // FIXME : This approach of bailing out from merge is conservative in
3577 // some ways like even if stg loops are not present after merge the
3578 // insert list, this liveness check is done (which is not needed).
3580 LiveRegs.addLiveOuts(*MBB);
3581 for (auto I = MBB->rbegin();; ++I) {
3582 MachineInstr &MI = *I;
3583 if (MI == InsertI)
3584 break;
3585 LiveRegs.stepBackward(*I);
3586 }
3587 InsertI++;
3588 if (LiveRegs.contains(AArch64::NZCV))
3589 return InsertI;
3590
3591 // Emitting an MTE loop requires two physical registers (BaseReg and
3592 // SizeReg). If the function is under register pressure, the register
3593 // scavenger will crash trying to allocate them. If we don't have at least
3594 // two free slots (free registers + emergency slots), bail out and fall back
3595 // to the unrolled sequence.
3596 if (countAvailableScavengerSlots(LiveRegs, MBB->getParent()->getRegInfo(),
3597 RS) < 2) {
3598 LLVM_DEBUG(
3599 dbgs() << "Failed to merge MTE stack tagging instructions into loop "
3600 << "due to high register pressure.\n");
3601 return InsertI;
3602 }
3603
3604 llvm::stable_sort(Instrs,
3605 [](const TagStoreInstr &Left, const TagStoreInstr &Right) {
3606 return Left.Offset < Right.Offset;
3607 });
3608
3609 // Make sure that we don't have any overlapping stores.
3610 int64_t CurOffset = Instrs[0].Offset;
3611 for (auto &Instr : Instrs) {
3612 if (CurOffset > Instr.Offset)
3613 return NextI;
3614 CurOffset = Instr.Offset + Instr.Size;
3615 }
3616
3617 // Find contiguous runs of tagged memory and emit shorter instruction
3618 // sequences for them when possible.
3619 TagStoreEdit TSE(MBB, FirstZeroData);
3620 std::optional<int64_t> EndOffset;
3621 for (auto &Instr : Instrs) {
3622 if (EndOffset && *EndOffset != Instr.Offset) {
3623 // Found a gap.
3624 TSE.emitCode(InsertI, TFI, /*TryMergeSPUpdate = */ false);
3625 TSE.clear();
3626 }
3627
3628 TSE.addInstruction(Instr);
3629 EndOffset = Instr.Offset + Instr.Size;
3630 }
3631
3632 const MachineFunction *MF = MBB->getParent();
3633 // Multiple FP/SP updates in a loop cannot be described by CFI instructions.
3634 TSE.emitCode(
3635 InsertI, TFI, /*TryMergeSPUpdate = */
3637
3638 return InsertI;
3639}
3640} // namespace
3641
3643 MachineFunction &MF, RegScavenger *RS = nullptr) const {
3644 bool MergeSetTag = MF.getSubtarget<AArch64Subtarget>()
3645 .getCLOpts()
3646 .stack_tagging_merge_settag;
3647 for (auto &BB : MF)
3648 for (MachineBasicBlock::iterator II = BB.begin(); II != BB.end();) {
3649 if (MergeSetTag)
3650 II = tryMergeAdjacentSTG(II, this, RS);
3651 }
3652
3653 // By the time this method is called, most of the prologue/epilogue code is
3654 // already emitted, whether its location was affected by the shrink-wrapping
3655 // optimization or not.
3656 if (!MF.getFunction().hasFnAttribute(Attribute::Naked) &&
3657 shouldSignReturnAddressEverywhere(MF))
3659}
3660
3661/// For Win64 AArch64 EH, the offset to the Unwind object is from the SP
3662/// before the update. This is easily retrieved as it is exactly the offset
3663/// that is set in processFunctionBeforeFrameFinalized.
3665 const MachineFunction &MF, int FI, Register &FrameReg,
3666 bool IgnoreSPUpdates) const {
3667 const MachineFrameInfo &MFI = MF.getFrameInfo();
3668 if (IgnoreSPUpdates) {
3669 LLVM_DEBUG(dbgs() << "Offset from the SP for " << FI << " is "
3670 << MFI.getObjectOffset(FI) << "\n");
3671 FrameReg = AArch64::SP;
3672 return StackOffset::getFixed(MFI.getObjectOffset(FI));
3673 }
3674
3675 // Go to common code if we cannot provide sp + offset.
3676 if (MFI.hasVarSizedObjects() ||
3679 return getFrameIndexReference(MF, FI, FrameReg);
3680
3681 FrameReg = AArch64::SP;
3682 return getStackOffset(MF, MFI.getObjectOffset(FI));
3683}
3684
3685/// The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve
3686/// the parent's frame pointer
3688 const MachineFunction &MF) const {
3689 return 0;
3690}
3691
3692/// Funclets only need to account for space for the callee saved registers,
3693/// as the locals are accounted for in the parent's stack frame.
3695 const MachineFunction &MF) const {
3696 // This is the size of the pushed CSRs.
3697 unsigned CSSize =
3698 MF.getInfo<AArch64FunctionInfo>()->getCalleeSavedStackSize();
3699 // This is the amount of stack a funclet needs to allocate.
3700 return alignTo(CSSize + MF.getFrameInfo().getMaxCallFrameSize(),
3701 getStackAlign());
3702}
3703
3704namespace {
3705struct FrameObject {
3706 bool IsValid = false;
3707 // Index of the object in MFI.
3708 int ObjectIndex = 0;
3709 // Group ID this object belongs to.
3710 int GroupIndex = -1;
3711 // This object should be placed first (closest to SP).
3712 bool ObjectFirst = false;
3713 // This object's group (which always contains the object with
3714 // ObjectFirst==true) should be placed first.
3715 bool GroupFirst = false;
3716
3717 // Used to distinguish between FP and GPR accesses. The values are decided so
3718 // that they sort FPR < Hazard < GPR and they can be or'd together.
3719 unsigned Accesses = 0;
3720 enum { AccessFPR = 1, AccessHazard = 2, AccessGPR = 4 };
3721};
3722
3723class GroupBuilder {
3724 SmallVector<int, 8> CurrentMembers;
3725 int NextGroupIndex = 0;
3726 std::vector<FrameObject> &Objects;
3727
3728public:
3729 GroupBuilder(std::vector<FrameObject> &Objects) : Objects(Objects) {}
3730 void AddMember(int Index) { CurrentMembers.push_back(Index); }
3731 void EndCurrentGroup() {
3732 if (CurrentMembers.size() > 1) {
3733 // Create a new group with the current member list. This might remove them
3734 // from their pre-existing groups. That's OK, dealing with overlapping
3735 // groups is too hard and unlikely to make a difference.
3736 LLVM_DEBUG(dbgs() << "group:");
3737 for (int Index : CurrentMembers) {
3738 Objects[Index].GroupIndex = NextGroupIndex;
3739 LLVM_DEBUG(dbgs() << " " << Index);
3740 }
3741 LLVM_DEBUG(dbgs() << "\n");
3742 NextGroupIndex++;
3743 }
3744 CurrentMembers.clear();
3745 }
3746};
3747
3748bool FrameObjectCompare(const FrameObject &A, const FrameObject &B) {
3749 // Objects at a lower index are closer to FP; objects at a higher index are
3750 // closer to SP.
3751 //
3752 // For consistency in our comparison, all invalid objects are placed
3753 // at the end. This also allows us to stop walking when we hit the
3754 // first invalid item after it's all sorted.
3755 //
3756 // If we want to include a stack hazard region, order FPR accesses < the
3757 // hazard object < GPRs accesses in order to create a separation between the
3758 // two. For the Accesses field 1 = FPR, 2 = Hazard Object, 4 = GPR.
3759 //
3760 // Otherwise the "first" object goes first (closest to SP), followed by the
3761 // members of the "first" group.
3762 //
3763 // The rest are sorted by the group index to keep the groups together.
3764 // Higher numbered groups are more likely to be around longer (i.e. untagged
3765 // in the function epilogue and not at some earlier point). Place them closer
3766 // to SP.
3767 //
3768 // If all else equal, sort by the object index to keep the objects in the
3769 // original order.
3770 return std::make_tuple(!A.IsValid, A.Accesses, A.ObjectFirst, A.GroupFirst,
3771 A.GroupIndex, A.ObjectIndex) <
3772 std::make_tuple(!B.IsValid, B.Accesses, B.ObjectFirst, B.GroupFirst,
3773 B.GroupIndex, B.ObjectIndex);
3774}
3775} // namespace
3776
3778 const MachineFunction &MF, SmallVectorImpl<int> &ObjectsToAllocate) const {
3780
3781 if ((!MF.getSubtarget<AArch64Subtarget>().getCLOpts().order_frame_objects &&
3782 !AFI.hasSplitSVEObjects()) ||
3783 ObjectsToAllocate.empty())
3784 return;
3785
3786 const MachineFrameInfo &MFI = MF.getFrameInfo();
3787 std::vector<FrameObject> FrameObjects(MFI.getObjectIndexEnd());
3788 for (auto &Obj : ObjectsToAllocate) {
3789 FrameObjects[Obj].IsValid = true;
3790 FrameObjects[Obj].ObjectIndex = Obj;
3791 }
3792
3793 // Identify FPR vs GPR slots for hazards, and stack slots that are tagged at
3794 // the same time.
3795 GroupBuilder GB(FrameObjects);
3796 for (auto &MBB : MF) {
3797 for (auto &MI : MBB) {
3798 if (MI.isDebugInstr())
3799 continue;
3800
3801 if (AFI.hasStackHazardSlotIndex()) {
3802 std::optional<int> FI = getLdStFrameID(MI, MFI);
3803 if (FI && *FI >= 0 && *FI < (int)FrameObjects.size()) {
3804 if (MFI.getStackID(*FI) == TargetStackID::ScalableVector ||
3806 FrameObjects[*FI].Accesses |= FrameObject::AccessFPR;
3807 else
3808 FrameObjects[*FI].Accesses |= FrameObject::AccessGPR;
3809 }
3810 }
3811
3812 int OpIndex;
3813 switch (MI.getOpcode()) {
3814 case AArch64::STGloop:
3815 case AArch64::STZGloop:
3816 OpIndex = 3;
3817 break;
3818 case AArch64::STGi:
3819 case AArch64::STZGi:
3820 case AArch64::ST2Gi:
3821 case AArch64::STZ2Gi:
3822 OpIndex = 1;
3823 break;
3824 default:
3825 OpIndex = -1;
3826 }
3827
3828 int TaggedFI = -1;
3829 if (OpIndex >= 0) {
3830 const MachineOperand &MO = MI.getOperand(OpIndex);
3831 if (MO.isFI()) {
3832 int FI = MO.getIndex();
3833 if (FI >= 0 && FI < MFI.getObjectIndexEnd() &&
3834 FrameObjects[FI].IsValid)
3835 TaggedFI = FI;
3836 }
3837 }
3838
3839 // If this is a stack tagging instruction for a slot that is not part of a
3840 // group yet, either start a new group or add it to the current one.
3841 if (TaggedFI >= 0)
3842 GB.AddMember(TaggedFI);
3843 else
3844 GB.EndCurrentGroup();
3845 }
3846 // Groups should never span multiple basic blocks.
3847 GB.EndCurrentGroup();
3848 }
3849
3850 if (AFI.hasStackHazardSlotIndex()) {
3851 FrameObjects[AFI.getStackHazardSlotIndex()].Accesses =
3852 FrameObject::AccessHazard;
3853 // If a stack object is unknown or both GPR and FPR, sort it into GPR.
3854 for (auto &Obj : FrameObjects)
3855 if (!Obj.Accesses ||
3856 Obj.Accesses == (FrameObject::AccessGPR | FrameObject::AccessFPR))
3857 Obj.Accesses = FrameObject::AccessGPR;
3858 }
3859
3860 // If the function's tagged base pointer is pinned to a stack slot, we want to
3861 // put that slot first when possible. This will likely place it at SP + 0,
3862 // and save one instruction when generating the base pointer because IRG does
3863 // not allow an immediate offset.
3864 std::optional<int> TBPI = AFI.getTaggedBasePointerIndex();
3865 if (TBPI) {
3866 FrameObjects[*TBPI].ObjectFirst = true;
3867 FrameObjects[*TBPI].GroupFirst = true;
3868 int FirstGroupIndex = FrameObjects[*TBPI].GroupIndex;
3869 if (FirstGroupIndex >= 0)
3870 for (FrameObject &Object : FrameObjects)
3871 if (Object.GroupIndex == FirstGroupIndex)
3872 Object.GroupFirst = true;
3873 }
3874
3875 llvm::stable_sort(FrameObjects, FrameObjectCompare);
3876
3877 int i = 0;
3878 for (auto &Obj : FrameObjects) {
3879 // All invalid items are sorted at the end, so it's safe to stop.
3880 if (!Obj.IsValid)
3881 break;
3882 ObjectsToAllocate[i++] = Obj.ObjectIndex;
3883 }
3884
3885 LLVM_DEBUG({
3886 dbgs() << "Final frame order:\n";
3887 for (auto &Obj : FrameObjects) {
3888 if (!Obj.IsValid)
3889 break;
3890 dbgs() << " " << Obj.ObjectIndex << ": group " << Obj.GroupIndex;
3891 if (Obj.ObjectFirst)
3892 dbgs() << ", first";
3893 if (Obj.GroupFirst)
3894 dbgs() << ", group-first";
3895 dbgs() << "\n";
3896 }
3897 });
3898}
3899
3900/// Emit a loop to decrement SP until it is equal to TargetReg, with probes at
3901/// least every ProbeSize bytes. Returns an iterator of the first instruction
3902/// after the loop. The difference between SP and TargetReg must be an exact
3903/// multiple of ProbeSize.
3905AArch64FrameLowering::inlineStackProbeLoopExactMultiple(
3906 MachineBasicBlock::iterator MBBI, int64_t ProbeSize,
3907 Register TargetReg) const {
3908 MachineBasicBlock &MBB = *MBBI->getParent();
3909 MachineFunction &MF = *MBB.getParent();
3910 const AArch64InstrInfo *TII =
3911 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3912 DebugLoc DL = MBB.findDebugLoc(MBBI);
3913
3914 MachineFunction::iterator MBBInsertPoint = std::next(MBB.getIterator());
3915 MachineBasicBlock *LoopMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3916 MF.insert(MBBInsertPoint, LoopMBB);
3917 MachineBasicBlock *ExitMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3918 MF.insert(MBBInsertPoint, ExitMBB);
3919
3920 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not encodable
3921 // in SUB).
3922 emitFrameOffset(*LoopMBB, LoopMBB->end(), DL, AArch64::SP, AArch64::SP,
3923 StackOffset::getFixed(-ProbeSize), TII,
3925 // LDR XZR, [SP]
3926 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::LDRXui))
3927 .addDef(AArch64::XZR)
3928 .addReg(AArch64::SP)
3929 .addImm(0)
3933 Align(8)))
3935 // CMP SP, TargetReg
3936 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::SUBSXrx64),
3937 AArch64::XZR)
3938 .addReg(AArch64::SP)
3939 .addReg(TargetReg)
3942 // B.CC Loop
3943 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::Bcc))
3945 .addMBB(LoopMBB)
3947
3948 LoopMBB->addSuccessor(ExitMBB);
3949 LoopMBB->addSuccessor(LoopMBB);
3950 // Synthesize the exit MBB.
3951 ExitMBB->splice(ExitMBB->end(), &MBB, MBBI, MBB.end());
3953 MBB.addSuccessor(LoopMBB);
3954 // Update liveins.
3955 fullyRecomputeLiveIns({ExitMBB, LoopMBB});
3956
3957 return ExitMBB->begin();
3958}
3959
3960void AArch64FrameLowering::inlineStackProbeFixed(
3961 MachineBasicBlock::iterator MBBI, Register ScratchReg, int64_t FrameSize,
3962 StackOffset CFAOffset) const {
3963 MachineBasicBlock *MBB = MBBI->getParent();
3964 MachineFunction &MF = *MBB->getParent();
3965 const AArch64InstrInfo *TII =
3966 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3967 AArch64FunctionInfo *AFI = MF.getInfo<AArch64FunctionInfo>();
3968 bool EmitAsyncCFI = AFI->needsAsyncDwarfUnwindInfo(MF);
3969 bool HasFP = hasFP(MF);
3970
3971 DebugLoc DL;
3972 int64_t ProbeSize = MF.getInfo<AArch64FunctionInfo>()->getStackProbeSize();
3973 int64_t NumBlocks = FrameSize / ProbeSize;
3974 int64_t ResidualSize = FrameSize % ProbeSize;
3975
3976 LLVM_DEBUG(dbgs() << "Stack probing: total " << FrameSize << " bytes, "
3977 << NumBlocks << " blocks of " << ProbeSize
3978 << " bytes, plus " << ResidualSize << " bytes\n");
3979
3980 // Decrement SP by NumBlock * ProbeSize bytes, with either unrolled or
3981 // ordinary loop.
3982 if (NumBlocks <= AArch64::StackProbeMaxLoopUnroll) {
3983 for (int i = 0; i < NumBlocks; ++i) {
3984 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not
3985 // encodable in a SUB).
3986 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3987 StackOffset::getFixed(-ProbeSize), TII,
3988 MachineInstr::FrameSetup, false, false, nullptr,
3989 EmitAsyncCFI && !HasFP, CFAOffset);
3990 CFAOffset += StackOffset::getFixed(ProbeSize);
3991 // LDR XZR, [SP]
3992 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3993 .addDef(AArch64::XZR)
3994 .addReg(AArch64::SP)
3995 .addImm(0)
3999 Align(8)))
4001 }
4002 } else if (NumBlocks != 0) {
4003 // SUB ScratchReg, SP, #FrameSize (or equivalent if FrameSize is not
4004 // encodable in ADD). ScrathReg may temporarily become the CFA register.
4005 emitFrameOffset(*MBB, MBBI, DL, ScratchReg, AArch64::SP,
4006 StackOffset::getFixed(-ProbeSize * NumBlocks), TII,
4007 MachineInstr::FrameSetup, false, false, nullptr,
4008 EmitAsyncCFI && !HasFP, CFAOffset);
4009 CFAOffset += StackOffset::getFixed(ProbeSize * NumBlocks);
4010 MBBI = inlineStackProbeLoopExactMultiple(MBBI, ProbeSize, ScratchReg);
4011 MBB = MBBI->getParent();
4012 if (EmitAsyncCFI && !HasFP) {
4013 // Set the CFA register back to SP.
4014 CFIInstBuilder(*MBB, MBBI, MachineInstr::FrameSetup)
4015 .buildDefCFARegister(AArch64::SP);
4016 }
4017 }
4018
4019 if (ResidualSize != 0) {
4020 // SUB SP, SP, #ResidualSize (or equivalent if ResidualSize is not encodable
4021 // in SUB).
4022 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
4023 StackOffset::getFixed(-ResidualSize), TII,
4024 MachineInstr::FrameSetup, false, false, nullptr,
4025 EmitAsyncCFI && !HasFP, CFAOffset);
4026 if (ResidualSize > AArch64::StackProbeMaxUnprobedStack) {
4027 // LDR XZR, [SP]
4028 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
4029 .addDef(AArch64::XZR)
4030 .addReg(AArch64::SP)
4031 .addImm(0)
4035 Align(8)))
4037 }
4038 }
4039}
4040
4041void AArch64FrameLowering::inlineStackProbe(MachineFunction &MF,
4042 MachineBasicBlock &MBB) const {
4043 // Get the instructions that need to be replaced. We emit at most two of
4044 // these. Remember them in order to avoid complications coming from the need
4045 // to traverse the block while potentially creating more blocks.
4046 SmallVector<MachineInstr *, 4> ToReplace;
4047 for (MachineInstr &MI : MBB)
4048 if (MI.getOpcode() == AArch64::PROBED_STACKALLOC ||
4049 MI.getOpcode() == AArch64::PROBED_STACKALLOC_VAR)
4050 ToReplace.push_back(&MI);
4051
4052 for (MachineInstr *MI : ToReplace) {
4053 if (MI->getOpcode() == AArch64::PROBED_STACKALLOC) {
4054 Register ScratchReg = MI->getOperand(0).getReg();
4055 int64_t FrameSize = MI->getOperand(1).getImm();
4056 StackOffset CFAOffset = StackOffset::get(MI->getOperand(2).getImm(),
4057 MI->getOperand(3).getImm());
4058 inlineStackProbeFixed(MI->getIterator(), ScratchReg, FrameSize,
4059 CFAOffset);
4060 } else {
4061 assert(MI->getOpcode() == AArch64::PROBED_STACKALLOC_VAR &&
4062 "Stack probe pseudo-instruction expected");
4063 const AArch64InstrInfo *TII =
4064 MI->getMF()->getSubtarget<AArch64Subtarget>().getInstrInfo();
4065 Register TargetReg = MI->getOperand(0).getReg();
4066 (void)TII->probedStackAlloc(MI->getIterator(), TargetReg, true);
4067 }
4068 MI->eraseFromParent();
4069 }
4070}
4071
4074 NotAccessed = 0, // Stack object not accessed by load/store instructions.
4075 GPR = 1 << 0, // A general purpose register.
4076 PPR = 1 << 1, // A predicate register.
4077 FPR = 1 << 2, // A floating point/Neon/SVE register.
4078 };
4079
4080 int Idx;
4082 int64_t Size;
4083 unsigned AccessTypes;
4084
4086
4087 bool operator<(const StackAccess &Rhs) const {
4088 return std::make_tuple(start(), Idx) <
4089 std::make_tuple(Rhs.start(), Rhs.Idx);
4090 }
4091
4092 bool isCPU() const {
4093 // Predicate register load and store instructions execute on the CPU.
4095 }
4096 bool isSME() const { return AccessTypes & AccessType::FPR; }
4097 bool isMixed() const { return isCPU() && isSME(); }
4098
4099 int64_t start() const { return Offset.getFixed() + Offset.getScalable(); }
4100 int64_t end() const { return start() + Size; }
4101
4102 std::string getTypeString() const {
4103 switch (AccessTypes) {
4104 case AccessType::FPR:
4105 return "FPR";
4106 case AccessType::PPR:
4107 return "PPR";
4108 case AccessType::GPR:
4109 return "GPR";
4111 return "NA";
4112 default:
4113 return "Mixed";
4114 }
4115 }
4116
4117 void print(raw_ostream &OS) const {
4118 OS << getTypeString() << " stack object at [SP"
4119 << (Offset.getFixed() < 0 ? "" : "+") << Offset.getFixed();
4120 if (Offset.getScalable())
4121 OS << (Offset.getScalable() < 0 ? "" : "+") << Offset.getScalable()
4122 << " * vscale";
4123 OS << "]";
4124 }
4125};
4126
4127static inline raw_ostream &operator<<(raw_ostream &OS, const StackAccess &SA) {
4128 SA.print(OS);
4129 return OS;
4130}
4131
4132void AArch64FrameLowering::emitRemarks(
4133 const MachineFunction &MF, MachineOptimizationRemarkEmitter *ORE) const {
4134
4135 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
4137 return;
4138
4139 uint64_t HazardSize = getStackHazardSize(MF);
4140 if (!HazardSize)
4141 HazardSize = MF.getSubtarget<AArch64Subtarget>()
4142 .getCLOpts()
4143 .stack_hazard_remark_size;
4144
4145 if (HazardSize == 0)
4146 return;
4147
4148 const MachineFrameInfo &MFI = MF.getFrameInfo();
4149 // Bail if function has no stack objects.
4150 if (!MFI.hasStackObjects())
4151 return;
4152
4153 std::vector<StackAccess> StackAccesses(MFI.getNumObjects());
4154
4155 size_t NumFPLdSt = 0;
4156 size_t NumNonFPLdSt = 0;
4157
4158 // Collect stack accesses via Load/Store instructions.
4159 for (const MachineBasicBlock &MBB : MF) {
4160 for (const MachineInstr &MI : MBB) {
4161 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
4162 continue;
4163 for (MachineMemOperand *MMO : MI.memoperands()) {
4164 std::optional<int> FI = getMMOFrameID(MMO, MFI);
4165 if (FI && !MFI.isDeadObjectIndex(*FI)) {
4166 int FrameIdx = *FI;
4167
4168 size_t ArrIdx = FrameIdx + MFI.getNumFixedObjects();
4169 if (StackAccesses[ArrIdx].AccessTypes == StackAccess::NotAccessed) {
4170 StackAccesses[ArrIdx].Idx = FrameIdx;
4171 StackAccesses[ArrIdx].Offset =
4172 getFrameIndexReferenceFromSP(MF, FrameIdx);
4173 StackAccesses[ArrIdx].Size = MFI.getObjectSize(FrameIdx);
4174 }
4175
4176 unsigned RegTy = StackAccess::AccessType::GPR;
4177 if (MFI.hasScalableStackID(FrameIdx))
4180 RegTy = StackAccess::FPR;
4181
4182 StackAccesses[ArrIdx].AccessTypes |= RegTy;
4183
4184 if (RegTy == StackAccess::FPR)
4185 ++NumFPLdSt;
4186 else
4187 ++NumNonFPLdSt;
4188 }
4189 }
4190 }
4191 }
4192
4193 if (NumFPLdSt == 0 || NumNonFPLdSt == 0)
4194 return;
4195
4196 llvm::sort(StackAccesses);
4197 llvm::erase_if(StackAccesses, [](const StackAccess &S) {
4199 });
4200
4203
4204 if (StackAccesses.front().isMixed())
4205 MixedObjects.push_back(&StackAccesses.front());
4206
4207 for (auto It = StackAccesses.begin(), End = std::prev(StackAccesses.end());
4208 It != End; ++It) {
4209 const auto &First = *It;
4210 const auto &Second = *(It + 1);
4211
4212 if (Second.isMixed())
4213 MixedObjects.push_back(&Second);
4214
4215 if ((First.isSME() && Second.isCPU()) ||
4216 (First.isCPU() && Second.isSME())) {
4217 uint64_t Distance = static_cast<uint64_t>(Second.start() - First.end());
4218 if (Distance < HazardSize)
4219 HazardPairs.emplace_back(&First, &Second);
4220 }
4221 }
4222
4223 auto EmitRemark = [&](llvm::StringRef Str) {
4224 ORE->emit([&]() {
4225 auto R = MachineOptimizationRemarkAnalysis(
4226 "sme", "StackHazard", MF.getFunction().getSubprogram(), &MF.front());
4227 return R << formatv("stack hazard in '{0}': ", MF.getName()).str() << Str;
4228 });
4229 };
4230
4231 for (const auto &P : HazardPairs)
4232 EmitRemark(formatv("{0} is too close to {1}", *P.first, *P.second).str());
4233
4234 for (const auto *Obj : MixedObjects)
4235 EmitRemark(
4236 formatv("{0} accessed by both GP and FP instructions", *Obj).str());
4237}
static void getLiveRegsForEntryMBB(LivePhysRegs &LiveRegs, const MachineBasicBlock &MBB)
static const unsigned DefaultSafeSPDisplacement
This is the biggest offset to the stack pointer we can encode in aarch64 instructions (without using ...
static void orderZPRCalleeSavesForPairs(MachineFunction &MF, const TargetRegisterInfo *RegInfo, std::vector< CalleeSavedInfo > &CSI)
static RegState getPrologueDeath(MachineFunction &MF, unsigned Reg)
static bool produceCompactUnwindFrame(const AArch64FrameLowering &, MachineFunction &MF)
bool enableMultiVectorSpillFill(const AArch64Subtarget &Subtarget, MachineFunction &MF)
static std::optional< int > getLdStFrameID(const MachineInstr &MI, const MachineFrameInfo &MFI)
void computeCalleeSaveRegisterPairs(const AArch64FrameLowering &AFL, MachineFunction &MF, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI, SmallVectorImpl< RegPairInfo > &RegPairs, bool NeedsFrameRecord)
static bool invalidateRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool UsesWinAAPCS, bool NeedsWinCFI, bool NeedsFrameRecord, const TargetRegisterInfo *TRI)
Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
static bool isLikelyToHaveSVEStack(const AArch64FrameLowering &AFL, const MachineFunction &MF)
static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool NeedsWinCFI, const TargetRegisterInfo *TRI)
static SVEStackSizes determineSVEStackSizes(MachineFunction &MF, AssignObjectOffsets AssignOffsets)
Process all the SVE stack objects and the SVE stack size and offsets for each object.
static bool isTargetWindows(const MachineFunction &MF)
static unsigned estimateRSStackSizeLimit(MachineFunction &MF)
Look at each instruction that references stack frames and return the stack size limit beyond which so...
static bool getSVECalleeSaveSlotRange(const MachineFrameInfo &MFI, int &Min, int &Max)
returns true if there are any SVE callee saves.
static MCRegister getRegisterOrZero(MCRegister Reg, bool HasSVE)
static unsigned getStackHazardSize(const MachineFunction &MF)
static bool isValidMemOpOffset(const AArch64InstrInfo *TII, unsigned Opcode, int Offset)
MCRegister findFreePredicateReg(BitVector &SavedRegs)
static bool isPPRAccess(const MachineInstr &MI)
static std::optional< int > getMMOFrameID(MachineMemOperand *MMO, const MachineFrameInfo &MFI)
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
This file contains the declaration of the AArch64PrologueEmitter and AArch64EpilogueEmitter classes,...
static const int kSetTagLoopThreshold
unsigned Imm
unsigned uint64_t
MachineBasicBlock & MBB
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
MachineBasicBlock MachineBasicBlock::iterator MBBI
This file contains the simple types necessary to represent the attributes associated with functions a...
#define CASE(ATTRNAME, AANAME,...)
static GCRegistry::Add< ErlangGC > A("erlang", "erlang-compatible garbage collector")
static GCRegistry::Add< CoreCLRGC > E("coreclr", "CoreCLR-compatible GC")
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
DXIL Forward Handle Accesses
const HexagonInstrInfo * TII
IRTranslator LLVM IR MI
Module.h This file contains the declarations for the Module class.
static std::string getTypeString(Type *T)
Definition LLParser.cpp:68
This file implements the LivePhysRegs utility for tracking liveness of physical registers.
#define F(x, y, z)
Definition MD5.cpp:54
#define I(x, y, z)
Definition MD5.cpp:57
#define H(x, y, z)
Definition MD5.cpp:56
Register Reg
Register const TargetRegisterInfo * TRI
Promote Memory to Register
Definition Mem2Reg.cpp:110
uint64_t IntrinsicInst * II
#define P(N)
This file declares the machine register scavenger class.
static bool isValid(const char C)
Returns true if C is a valid mangled character: <0-9a-zA-Z_>.
static bool contains(SmallPtrSetImpl< ConstantExpr * > &Cache, ConstantExpr *Expr, Constant *C)
Definition Value.cpp:484
This file defines the scope_exit class, which executes user-defined cleanup logic at scope exit.
This file defines the SmallVector class.
#define LLVM_DEBUG(...)
Definition Debug.h:119
StackOffset getSVEStackSize(const MachineFunction &MF) const
Returns the size of the entire SVE stackframe (PPRs + ZPRs).
StackOffset getZPRStackSize(const MachineFunction &MF) const
Returns the size of the entire ZPR stackframe (calleesaves + spills).
void processFunctionBeforeFrameIndicesReplaced(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameIndicesReplaced - This method is called immediately before MO_FrameIndex op...
MachineBasicBlock::iterator eliminateCallFramePseudoInstr(MachineFunction &MF, MachineBasicBlock &MBB, MachineBasicBlock::iterator I) const override
This method is called during prolog/epilog code insertion to eliminate call frame setup and destroy p...
bool canUseAsPrologue(const MachineBasicBlock &MBB) const override
Check whether or not the given MBB can be used as a prologue for the target.
bool enableStackSlotScavenging(const MachineFunction &MF) const override
Returns true if the stack slot holes in the fixed and callee-save stack area should be used when allo...
bool assignCalleeSavedSpillSlots(MachineFunction &MF, const TargetRegisterInfo *TRI, std::vector< CalleeSavedInfo > &CSI) const override
assignCalleeSavedSpillSlots - Allows target to override spill slot assignment logic.
bool spillCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
spillCalleeSavedRegisters - Issues instruction(s) to spill all callee saved registers and returns tru...
bool restoreCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, MutableArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
restoreCalleeSavedRegisters - Issues instruction(s) to restore all callee saved registers and returns...
bool enableFullCFIFixup(const MachineFunction &MF) const override
enableFullCFIFixup - Returns true if we may need to fix the unwind information such that it is accura...
StackOffset getFrameIndexReferenceFromSP(const MachineFunction &MF, int FI) const override
getFrameIndexReferenceFromSP - This method returns the offset from the stack pointer to the slot of t...
bool enableCFIFixup(const MachineFunction &MF) const override
Returns true if we may need to fix the unwind information for the function.
StackOffset getNonLocalFrameIndexReference(const MachineFunction &MF, int FI) const override
getNonLocalFrameIndexReference - This method returns the offset used to reference a frame index locat...
TargetStackID::Value getStackIDForScalableVectors() const override
Returns the StackID that scalable vectors should be associated with.
bool hasFPImpl(const MachineFunction &MF) const override
hasFPImpl - Return true if the specified function should have a dedicated frame pointer register.
void emitPrologue(MachineFunction &MF, MachineBasicBlock &MBB) const override
emitProlog/emitEpilog - These methods insert prolog and epilog code into the function.
void resetCFIToInitialState(MachineBasicBlock &MBB) const override
Emit CFI instructions that recreate the state of the unwind information upon function entry.
bool hasReservedCallFrame(const MachineFunction &MF) const override
hasReservedCallFrame - Under normal circumstances, when a frame pointer is not required,...
bool hasSVECalleeSavesAboveFrameRecord(const MachineFunction &MF) const
StackOffset resolveFrameOffsetReference(const MachineFunction &MF, int64_t ObjectOffset, bool isFixed, TargetStackID::Value StackID, Register &FrameReg, bool PreferFP, bool ForSimm) const
bool canUseRedZone(const MachineFunction &MF) const
Can this function use the red zone for local allocations.
bool needsWinCFI(const MachineFunction &MF) const
bool isFPReserved(const MachineFunction &MF) const
Should the Frame Pointer be reserved for the current function?
void processFunctionBeforeFrameFinalized(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameFinalized - This method is called immediately before the specified function...
int getSEHFrameIndexOffset(const MachineFunction &MF, int FI) const
unsigned getWinEHFuncletFrameSize(const MachineFunction &MF) const
Funclets only need to account for space for the callee saved registers, as the locals are accounted f...
void orderFrameObjects(const MachineFunction &MF, SmallVectorImpl< int > &ObjectsToAllocate) const override
Order the symbols in the local stack frame.
void emitEpilogue(MachineFunction &MF, MachineBasicBlock &MBB) const override
StackOffset getPPRStackSize(const MachineFunction &MF) const
Returns the size of the entire PPR stackframe (calleesaves + spills + hazard padding).
int64_t getArgumentStackToRestore(MachineFunction &MF, MachineBasicBlock &MBB) const
Returns how much of the incoming argument stack area (in bytes) we should clean up in an epilogue.
void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS) const override
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
StackOffset getFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg) const override
getFrameIndexReference - Provide a base+offset reference to an FI slot for debug info.
StackOffset getFrameIndexReferencePreferSP(const MachineFunction &MF, int FI, Register &FrameReg, bool IgnoreSPUpdates) const override
For Win64 AArch64 EH, the offset to the Unwind object is from the SP before the update.
StackOffset resolveFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP, bool ForSimm) const
unsigned getWinEHParentFrameOffset(const MachineFunction &MF) const override
The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve the parent's frame pointer...
bool requiresSaveVG(const MachineFunction &MF) const
void emitPacRetPlusLeafHardening(MachineFunction &MF) const
Harden the entire function with pac-ret.
AArch64FunctionInfo - This class is derived from MachineFunctionInfo and contains private AArch64-spe...
unsigned getCalleeSavedStackSize(const MachineFrameInfo &MFI) const
void setCalleeSaveBaseToFrameRecordOffset(int Offset)
SignReturnAddress getSignReturnAddressCondition() const
void setStackSizeSVE(uint64_t ZPR, uint64_t PPR)
std::optional< int > getTaggedBasePointerIndex() const
bool needsDwarfUnwindInfo(const MachineFunction &MF) const
void setSVECalleeSavedStackSize(unsigned ZPR, unsigned PPR)
bool needsAsyncDwarfUnwindInfo(const MachineFunction &MF) const
static bool isTailCallReturnInst(const MachineInstr &MI)
Returns true if MI is one of the TCRETURN* instructions.
static bool isFpOrNEON(Register Reg)
Returns whether the physical register is FP or NEON.
const AArch64RegisterInfo * getRegisterInfo() const override
bool isNeonAvailable() const
Returns true if the target has NEON and the function at runtime is known to have NEON enabled (e....
const AArch64InstrInfo * getInstrInfo() const override
const AArch64Options & getCLOpts() const
const AArch64TargetLowering * getTargetLowering() const override
bool isSVEorStreamingSVEAvailable() const
Returns true if the target has access to either the full range of SVE instructions,...
bool isStreaming() const
Returns true if the function has a streaming body.
bool hasInlineStackProbe(const MachineFunction &MF) const override
True if stack clash protection is enabled for this functions.
unsigned getRedZoneSize(const Function &F) const
Represent a constant reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:40
size_t size() const
Get the array size.
Definition ArrayRef.h:141
bool empty() const
Check if the array is empty.
Definition ArrayRef.h:136
bool test(unsigned Idx) const
Returns true if bit Idx is set.
Definition BitVector.h:482
BitVector & reset()
Reset all bits in the bitvector.
Definition BitVector.h:409
size_type count() const
Returns the number of bits which are set.
Definition BitVector.h:181
BitVector & set()
Set all bits in the bitvector.
Definition BitVector.h:366
iterator_range< const_set_bits_iterator > set_bits() const
Definition BitVector.h:159
size_type size() const
Returns the number of bits in this bitvector.
Definition BitVector.h:178
Helper class for creating CFI instructions and inserting them into MIR.
The CalleeSavedInfo class tracks the information need to locate where a callee saved register is in t...
A debug info location.
Definition DebugLoc.h:126
bool hasMinSize() const
Optimize this function for minimum size (-Oz).
Definition Function.h:696
CallingConv::ID getCallingConv() const
getCallingConv()/setCallingConv(CC) - These method get and set the calling convention of this functio...
Definition Function.h:273
AttributeList getAttributes() const
Return the attribute list for this Function.
Definition Function.h:329
bool isVarArg() const
isVarArg - Return true if this function takes a variable number of arguments.
Definition Function.h:230
bool hasFnAttribute(Attribute::AttrKind Kind) const
Return true if the function has the attribute.
Definition Function.cpp:734
Module * getParent()
Get the module that this global value is contained inside of...
A set of physical registers with utility functions to track liveness when walking backward/forward th...
bool usesWindowsCFI() const
Definition MCAsmInfo.h:675
Wrapper class representing physical registers. Should be passed by value.
Definition MCRegister.h:41
constexpr bool isValid() const
Definition MCRegister.h:84
LLVM_ABI void transferSuccessorsAndUpdatePHIs(MachineBasicBlock *FromMBB)
Transfers all the successors, as in transferSuccessors, and update PHI operands in the successor bloc...
LLVM_ABI iterator getFirstTerminator()
Returns an iterator to the first terminator instruction of this basic block.
LLVM_ABI void addSuccessor(MachineBasicBlock *Succ, BranchProbability Prob=BranchProbability::getUnknown())
Add Succ as a successor of this MachineBasicBlock.
const MachineFunction * getParent() const
Return the MachineFunction containing this basic block.
reverse_iterator rbegin()
iterator insertAfter(iterator I, MachineInstr *MI)
Insert MI into the instruction list after I.
void splice(iterator Where, MachineBasicBlock *Other, iterator From)
Take an instruction from MBB 'Other' at the position From, and insert it into this MBB right before '...
MachineInstrBundleIterator< MachineInstr > iterator
The MachineFrameInfo class represents an abstract stack frame until prolog/epilog code is inserted.
LLVM_ABI int CreateFixedObject(uint64_t Size, int64_t SPOffset, bool IsImmutable, bool isAliased=false)
Create a new object at a fixed location on the stack.
bool hasVarSizedObjects() const
This method may be called any time after instruction selection is complete to determine if the stack ...
const AllocaInst * getObjectAllocation(int ObjectIdx) const
Return the underlying Alloca of the specified stack object if it exists.
LLVM_ABI int CreateStackObject(uint64_t Size, Align Alignment, bool isSpillSlot, const AllocaInst *Alloca=nullptr, uint8_t ID=0)
Create a new statically sized stack object, returning a nonnegative identifier to represent it.
bool hasCalls() const
Return true if the current function has any function calls.
bool isFrameAddressTaken() const
This method may be called any time after instruction selection is complete to determine if there is a...
void setObjectOffset(int ObjectIdx, int64_t SPOffset)
Set the stack frame offset of the specified object.
bool isCalleeSavedObjectIndex(int ObjectIdx) const
uint64_t getMaxCallFrameSize() const
Return the maximum size of a call frame that must be allocated for an outgoing function call.
bool hasPatchPoint() const
This method may be called any time after instruction selection is complete to determine if there is a...
bool hasScalableStackID(int ObjectIdx) const
int getStackProtectorIndex() const
Return the index for the stack protector object.
LLVM_ABI uint64_t estimateStackSize(const MachineFunction &MF) const
Estimate and return the size of the stack frame.
void setStackID(int ObjectIdx, uint8_t ID)
bool isCalleeSavedInfoValid() const
Has the callee saved info been calculated yet?
Align getObjectAlign(int ObjectIdx) const
Return the alignment of the specified stack object.
int64_t getObjectSize(int ObjectIdx) const
Return the size of the specified object.
bool isMaxCallFrameSizeComputed() const
bool hasStackMap() const
This method may be called any time after instruction selection is complete to determine if there is a...
LLVM_ABI int CreateSpillStackObject(uint64_t Size, Align Alignment, TargetStackID::Value StackID=TargetStackID::Default)
Create a new statically sized stack object that represents a spill slot, returning a nonnegative iden...
const std::vector< CalleeSavedInfo > & getCalleeSavedInfo() const
Returns a reference to call saved info vector for the current function.
unsigned getNumObjects() const
Return the number of objects.
int getObjectIndexEnd() const
Return one past the maximum frame object index.
bool hasStackProtectorIndex() const
bool hasStackObjects() const
Return true if there are any stack objects in this function.
uint8_t getStackID(int ObjectIdx) const
unsigned getNumFixedObjects() const
Return the number of fixed objects.
void setIsCalleeSavedObjectIndex(int ObjectIdx, bool IsCalleeSaved)
int64_t getObjectOffset(int ObjectIdx) const
Return the assigned stack offset of the specified object from the incoming stack pointer.
int getObjectIndexBegin() const
Return the minimum frame object index.
void setObjectAlignment(int ObjectIdx, Align Alignment)
setObjectAlignment - Change the alignment of the specified stack object.
bool isDeadObjectIndex(int ObjectIdx) const
Returns true if the specified index corresponds to a dead object.
const WinEHFuncInfo * getWinEHFuncInfo() const
getWinEHFuncInfo - Return information about how the current function uses Windows exception handling.
const TargetSubtargetInfo & getSubtarget() const
getSubtarget - Return the subtarget for which this machine code is being compiled.
bool framePointerIsReserved() const
Returns true if the frame pointer must always either point to a new frame record or be un-modified in...
MachineFrameInfo & getFrameInfo()
getFrameInfo - Return the frame info object for the current function.
MachineRegisterInfo & getRegInfo()
getRegInfo - Return information about the registers currently in use.
Function & getFunction()
Return the LLVM function that this machine code represents.
BasicBlockListType::iterator iterator
bool disableFramePointerElim() const
Returns true if frame pointer elimination should be disabled for this function.
Ty * getInfo()
getInfo - Keep track of various per-function pieces of information for backends that would like to do...
const MachineBasicBlock & front() const
MachineMemOperand * getMachineMemOperand(MachinePointerInfo PtrInfo, MachineMemOperand::Flags F, LLT MemTy, Align BaseAlignment, const MMOMetadata &Metadata=MMOMetadata(), SyncScope::ID SSID=SyncScope::System, AtomicOrdering Ordering=AtomicOrdering::NotAtomic, AtomicOrdering FailureOrdering=AtomicOrdering::NotAtomic)
getMachineMemOperand - Allocate a new MachineMemOperand.
MachineBasicBlock * CreateMachineBasicBlock(const BasicBlock *BB=nullptr, std::optional< UniqueBBID > BBID=std::nullopt)
CreateMachineInstr - Allocate a new MachineInstr.
void insert(iterator MBBI, MachineBasicBlock *MBB)
const TargetMachine & getTarget() const
getTarget - Return the target machine this machine code is compiled with
const MachineInstrBuilder & setMemRefs(ArrayRef< MachineMemOperand * > MMOs) const
const MachineInstrBuilder & addExternalSymbol(const char *FnName, unsigned TargetFlags=0) const
const MachineInstrBuilder & addReg(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a new virtual register operand.
const MachineInstrBuilder & setMIFlag(MachineInstr::MIFlag Flag) const
const MachineInstrBuilder & addImm(int64_t Val) const
Add a new immediate operand.
const MachineInstrBuilder & addFrameIndex(int Idx) const
const MachineInstrBuilder & addRegMask(const uint32_t *Mask) const
const MachineInstrBuilder & addMBB(MachineBasicBlock *MBB, unsigned TargetFlags=0) const
const MachineInstrBuilder & addDef(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a virtual register definition operand.
const MachineInstrBuilder & setMIFlags(unsigned Flags) const
const MachineInstrBuilder & addMemOperand(MachineMemOperand *MMO) const
Representation of each machine instruction.
void setFlags(unsigned flags)
uint32_t getFlags() const
Return the MI flags bitvector.
LLVM_ABI MachineInstrBundleIterator< MachineInstr > eraseFromParent()
Unlink 'this' from the containing basic block and delete it.
A description of a memory reference used in the backend.
const PseudoSourceValue * getPseudoValue() const
@ MOVolatile
The memory access is volatile.
@ MOLoad
The memory access reads data.
@ MOStore
The memory access writes data.
const Value * getValue() const
Return the base address of the memory access.
MachineOperand class - Representation of each machine instruction operand.
int64_t getImm() const
bool isFI() const
isFI - Tests if this is a MO_FrameIndex operand.
LLVM_ABI void emit(DiagnosticInfoOptimizationBase &OptDiag)
Emit an optimization remark.
MachineRegisterInfo - Keep track of information for virtual and physical registers,...
LLVM_ABI void freezeReservedRegs()
freezeReservedRegs - Called by the register allocator to freeze the set of reserved registers before ...
bool isReserved(MCRegister PhysReg) const
isReserved - Returns true when PhysReg is a reserved register.
LLVM_ABI Register createVirtualRegister(const TargetRegisterClass *RegClass, StringRef Name="")
createVirtualRegister - Create and return a new virtual register in the function with the specified r...
LLVM_ABI bool isLiveIn(Register Reg) const
LLVM_ABI const MCPhysReg * getCalleeSavedRegs() const
Returns list of callee saved registers.
LLVM_ABI bool isPhysRegUsed(MCRegister PhysReg, bool SkipRegMaskTest=false) const
Return true if the specified register is modified or read in this function.
const Triple & getTargetTriple() const
Get the target triple which is a string describing the target host.
Definition Module.h:328
Represent a mutable reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:294
Wrapper class representing virtual and physical registers.
Definition Register.h:20
constexpr bool isValid() const
Definition Register.h:112
SMEAttrs is a utility class to parse the SME ACLE attributes on functions.
bool hasStreamingInterface() const
bool hasNonStreamingInterfaceAndBody() const
bool hasStreamingBody() const
bool insert(const value_type &X)
Insert a new element into the SetVector.
Definition SetVector.h:157
A SetVector that performs no allocations if smaller than a certain size.
Definition SetVector.h:345
This class consists of common code factored out of the SmallVector class to reduce code duplication b...
reference emplace_back(ArgTypes &&... Args)
void append(ItTy in_start, ItTy in_end)
Add the specified range to the end of the SmallVector.
void push_back(const T &Elt)
This is a 'vector' (really, a variable-sized array), optimized for the case when the array is small.
StackOffset holds a fixed and a scalable offset in bytes.
Definition TypeSize.h:30
int64_t getFixed() const
Returns the fixed component of the stack.
Definition TypeSize.h:46
int64_t getScalable() const
Returns the scalable component of the stack.
Definition TypeSize.h:49
static StackOffset get(int64_t Fixed, int64_t Scalable)
Definition TypeSize.h:41
static StackOffset getScalable(int64_t Scalable)
Definition TypeSize.h:40
static StackOffset getFixed(int64_t Fixed)
Definition TypeSize.h:39
bool hasFP(const MachineFunction &MF) const
hasFP - Return true if the specified function should have a dedicated frame pointer register.
virtual void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS=nullptr) const
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
int getOffsetOfLocalArea() const
getOffsetOfLocalArea - This method returns the offset of the local area from the stack pointer on ent...
Align getStackAlign() const
getStackAlignment - This method returns the number of bytes to which the stack pointer must be aligne...
StackDirection getStackGrowthDirection() const
getStackGrowthDirection - Return the direction the stack grows
virtual bool enableCFIFixup(const MachineFunction &MF) const
Returns true if we may need to fix the unwind information for the function.
const Triple & getTargetTriple() const
const MCAsmInfo & getMCAsmInfo() const
Return target specific asm information.
TargetRegisterInfo base class - We assume that the target defines a static array of TargetRegisterDes...
bool hasStackRealignment(const MachineFunction &MF) const
True if stack realignment is required and still possible.
virtual const TargetRegisterInfo * getRegisterInfo() const =0
Return the target's register information.
Triple - Helper class for working with autoconf configuration names.
Definition Triple.h:48
bool isOSBinFormatMachO() const
Tests whether the environment is MachO.
Definition Triple.h:876
This class implements an extremely fast bulk output stream that can only output to a stream.
Definition raw_ostream.h:53
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
static unsigned getShiftValue(unsigned Imm)
getShiftValue - Extract the shift value.
static unsigned getArithExtendImm(AArch64_AM::ShiftExtendType ET, unsigned Imm)
getArithExtendImm - Encode the extend type and shift amount for an arithmetic instruction: imm: 3-bit...
const unsigned StackProbeMaxLoopUnroll
Maximum number of iterations to unroll for a constant size probing loop.
const unsigned StackProbeMaxUnprobedStack
Maximum allowed number of unprobed bytes above SP at an ABI boundary.
constexpr char Align[]
Key for Kernel::Arg::Metadata::mAlign.
constexpr char Attrs[]
Key for Kernel::Metadata::mAttrs.
unsigned ID
LLVM IR allows to use arbitrary numbers as calling convention identifiers.
Definition CallingConv.h:24
@ AArch64_SVE_VectorCall
Used between AArch64 SVE functions.
@ PreserveMost
Used for runtime calls that preserves most registers.
Definition CallingConv.h:63
@ CXX_FAST_TLS
Used for access functions.
Definition CallingConv.h:72
@ GHC
Used by the Glasgow Haskell Compiler (GHC).
Definition CallingConv.h:50
@ PreserveAll
Used for runtime calls that preserves (almost) all registers.
Definition CallingConv.h:66
@ Fast
Attempts to make calls as fast as possible (e.g.
Definition CallingConv.h:41
@ PreserveNone
Used for runtime calls that preserves none general registers.
Definition CallingConv.h:90
@ Win64
The C convention as implemented on Windows/x86-64 and AArch64.
@ SwiftTail
This follows the Swift calling convention in how arguments are passed but guarantees tail calls will ...
Definition CallingConv.h:87
@ C
The default llvm calling convention, compatible with C.
Definition CallingConv.h:34
NodeAddr< InstrNode * > Instr
Definition RDFGraph.h:389
BaseReg
Stack frame base register. Bit 0 of FREInfo.Info.
Definition SFrame.h:77
This is an optimization pass for GlobalISel generic memory operations.
@ Offset
Definition DWP.cpp:577
detail::zippy< detail::zip_shortest, T, U, Args... > zip(T &&t, U &&u, Args &&...args)
zip iterator for two or more iteratable types.
Definition STLExtras.h:846
void stable_sort(R &&Range)
Definition STLExtras.h:2132
MachineInstrBuilder BuildMI(MachineFunction &MF, const MIMetadata &MIMD, const MCInstrDesc &MCID)
Builder interface. Specify how to create the initial instruction itself.
int isAArch64FrameOffsetLegal(const MachineInstr &MI, StackOffset &Offset, bool *OutUseUnscaledOp=nullptr, unsigned *OutUnscaledOp=nullptr, int64_t *EmittableOffset=nullptr)
Check if the Offset is a valid frame offset for MI.
@ Unknown
Not known to have no common set bits.
RegState
Flags to represent properties of register accesses.
@ Define
Register definition.
auto enumerate(FirstRange &&First, RestRanges &&...Rest)
Given two or more input ranges, returns a new range whose values are tuples (A, B,...
Definition STLExtras.h:2570
constexpr RegState getKillRegState(bool B)
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
void append_range(Container &C, Range &&R)
Wrapper function to append range R to container C.
Definition STLExtras.h:2224
@ AArch64FrameOffsetCannotUpdate
Offset cannot apply.
constexpr T alignDown(U Value, V Align, W Skew=0)
Returns the largest unsigned integer less than or equal to Value and is Skew mod Align.
Definition MathExtras.h:541
auto dyn_cast_or_null(const Y &Val)
Definition Casting.h:753
bool any_of(R &&range, UnaryPredicate P)
Provide wrappers to std::any_of which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1762
auto formatv(bool Validate, const char *Fmt, Ts &&...Vals)
auto reverse(ContainerTy &&C)
Definition STLExtras.h:408
void sort(IteratorTy Start, IteratorTy End)
Definition STLExtras.h:1652
LLVM_ABI raw_ostream & dbgs()
dbgs() - This returns a reference to a raw_ostream for debugging messages.
Definition Debug.cpp:209
void emitFrameOffset(MachineBasicBlock &MBB, MachineBasicBlock::iterator MBBI, const DebugLoc &DL, unsigned DestReg, unsigned SrcReg, StackOffset Offset, const TargetInstrInfo *TII, MachineInstr::MIFlag=MachineInstr::NoFlags, bool SetNZCV=false, bool NeedsWinCFI=false, bool *HasWinCFI=nullptr, bool EmitCFAOffset=false, StackOffset InitialOffset={}, unsigned FrameReg=AArch64::SP)
emitFrameOffset - Emit instructions as needed to set DestReg to SrcReg plus Offset.
LLVM_ABI void report_fatal_error(Error Err, bool gen_crash_diag=true)
Definition Error.cpp:163
constexpr uint64_t alignTo(uint64_t Size, Align A)
Returns a multiple of A needed to store Size bytes.
Definition Alignment.h:144
constexpr RegState getDefRegState(bool B)
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
LLVM_ABI const Value * getUnderlyingObject(const Value *V, unsigned MaxLookup=MaxLookupSearchDepth, bool MustPreserveProvenance=false)
This method strips off any GEP address adjustments, pointer casts or llvm.threadlocal....
@ First
Helpers to iterate all locations in the MemoryEffectsBase class.
Definition ModRef.h:74
uint16_t MCPhysReg
An unsigned integer type large enough to represent all physical registers, but not necessarily virtua...
Definition MCRegister.h:21
RelativeUniformCounterPtr ValuesPtrExpr VTableAddr Count
Definition InstrProf.h:145
raw_ostream & operator<<(raw_ostream &OS, const APFixedPoint &FX)
auto count_if(R &&Range, UnaryPredicate P)
Wrapper function around std::count_if to count the number of times an element satisfying a given pred...
Definition STLExtras.h:2035
auto find_if(R &&Range, UnaryPredicate P)
Provide wrappers to std::find_if which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1788
void erase_if(Container &C, UnaryPredicate P)
Provide a container algorithm similar to C++ Library Fundamentals v2's erase_if which is equivalent t...
Definition STLExtras.h:2208
bool is_contained(R &&Range, const E &Element)
Returns true if Element is found in Range.
Definition STLExtras.h:1963
void fullyRecomputeLiveIns(ArrayRef< MachineBasicBlock * > MBBs)
Convenience function for recomputing live-in's for a set of MBBs until the computation converges.
LLVM_ABI Printable printReg(Register Reg, const TargetRegisterInfo *TRI=nullptr, unsigned SubIdx=0, const MachineRegisterInfo *MRI=nullptr)
Prints virtual and physical registers with or without a TRI instance.
MCRegisterClass TargetRegisterClass
Definition FastISel.h:58
void swap(llvm::BitVector &LHS, llvm::BitVector &RHS)
Implement std::swap in terms of BitVector swap.
Definition BitVector.h:880
bool operator<(const StackAccess &Rhs) const
void print(raw_ostream &OS) const
int64_t start() const
std::string getTypeString() const
int64_t end() const
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:77
Pair of physical register and lane mask.
static LLVM_ABI MachinePointerInfo getUnknownStack(MachineFunction &MF)
Stack memory without other information.
static LLVM_ABI MachinePointerInfo getFixedStack(MachineFunction &MF, int FI, int64_t Offset=0)
Return a MachinePointerInfo record that refers to the specified FrameIndex.
SmallVector< WinEHTryBlockMapEntry, 4 > TryBlockMap
SmallVector< WinEHHandlerType, 1 > HandlerArray