[RFC] Floating-point literals in LLVM IR

I’d like to propose several changes to the various kinds of float literals that are available in LLVM IR. Specifically, I’d like to do the following:

Motivation

The obvious motivation for these changes is that hexadecimal floating literals, along with nan and inf, are simply more readable than the current hexadecimal, especially if you haven’t memorized the hex bitcast-to-int representation of common doubles like 1.0. However, that actually isn’t my main reason for doing this.

At present, when LLLexer is lexing in the input, a floating literal is converted into an APFloat token, which needs to choose which format is being used. When there is a decimal literal, the double format are used (as the lexer doesn’t have the information of the type being parsed, only the parser does); the weird hexadecimal formats use the codes to select between the different formats. The parser then converts the double-based APFloat to the appropriate type, and only here checks for exactness of conversion (so double 0.1 is legal but float 0.1 is not). It is this parse-to-double-then-convert-to-T that I am most concerned with changing, especially because I am working on adding support for decimal floating point types in LLVM IR (as part of RFC: Decimal floating-point support (ISO/IEC TS 18661-2 and C23)).

The simple fix is to lex a floating literal as a string, and have the parser itself convert the string to the desired floating-point type, being able to catch illegal (overflowing, underflowing, inexact) conversions at the process, whether or not it was representible (exactly) as a double.

On string representations

For the standard IEEE 754 binary types, and bfloat, all finite values have unique representations, and these representations are concise in the hexadecimal floating-point literal. There are of course multiple NaN values, but the proposal for NaN support would include the ability to specify the NaN payload, so every representable value in these types has a distinct string that can represent it.

These properties do not hold for x86_fp80, ppc_fp128, and the IEEE 754 decimal types (considering decimal literals in lieu of hexadecimal literals).

In the case of x86_fp80, the numbers have an explicit integer bit which, if set incorrectly, results in an invalid value. The behavior of such values currently falls into the–to use Nikita’s phrase from his recent keynote–undecided semantics of LLVM. For pseudo-denormals (where the biased exponent is 0 and the integer bit is 1), the hardware treats it as if it were a noncanonical representation of another finite value; for the other values, the effect is to raise the invalid exception, effectively as if it were a noncanonical sNaN. Whether this is best represented as a trap value (sorry, non-value representation in C23 terms) or noncanonical values is up for debate, but neither of these concepts have clear concepts in LLVM IR at present. To my mind, I don’t see that these are important enough to justify having direct representation as a floating literal in LLVM IR, where a bitcast constant expression could instead be sufficient.

I must confess ignorance of all the precise pitfalls of ppc_fp128; I’m only aware of the broad strokes of this type. It is a pair of two double values, with the second being smaller than the first. I’m not sure what the consequences of invalid values such as the first value being finite and the second being infinite, but I suspect that like x86_fp80, these are unimportant enough to not justify direct representation. However, I am somewhat more concerned by the fact that some of the more valid values do not have concise representation. For example, the pair {DBL_MAX, DBL_MIN} requires an awfully large number of digits to print out, and it may be prudent to retain a hexadecimal integer representation instead for this format.

I’m aware that decimal floating-point types are not in LLVM IR, but I still want to talk about them here since they are one of my motivations for this RFC. These types introduce the concept of decimal cohorts, which are numbers with the same numerical value but different exponents (e.g., 0e0 and 0e1 are in the same cohort but are numerically equivalent). There is a standard convention for determining which of the cohorts to use when converting from a decimal string, so different members of the same cohort can be considered to have different string representations and that’s not a problem. But decimal types, like x86_fp80 and ppc_fp128, have outright noncanonical values (which are distinct from decimal cohorts), for finite values as well as infinities. Note that, for noncanonical values, only one binary encoding is considered “correct”–e.g., the correct encoding for infinity sets the significand value to all 0’s–and any other value is noncanonical and will not be produced as the result of an operation. Finally, decimal types have two different binary representations (the BID and DPD formats), and while the set of valid values is the same for both representations, which bit patterns correspond to those patterns is different, and even which finite values have noncanonical representations differs.

All of this suggests to me that it is neither advantageous nor necessary to have floating literals that represent noncanonical floating-point values, especially as noncanonical value support in APFloat itself is spotty (see. e.g., APFloat: x87DoubleExtended pseudo-NaNs (integer_bit==0) not handled as always-signalling. · Issue #63938 · llvm/llvm-project · GitHub). If we have support for the non-finite literals, then I don’t think there is much reason to retain the weird hexadecimal format we use currently, and we can drop those values, although the issue of non-concise representation for some ppc_fp128 values gives me some pause.

On string-to-float conversions

The LangRef currently states

The assembler requires the exact decimal value of a floating-point constant. For example, the assembler accepts 1.25 but rejects 1.3 because 1.3 is a repeating decimal in binary.

This statement is flat-out wrong; no such check exists for double types, and other floating-point types check the exactness of conversion from double to that type, so if the conversion to double is not exact but the conversion to float is, the constant is considered valid.

In principle, string-to-float conversion can result in three exceptions: inexact, overflow (e.g., 1e999999), and underflow (e.g., 1e-99999999). I would propose that we make any exception that occurs on converting a string to the desired floating-point type cause a parser error. Presently, there seem to be 230 failures in the LLVM test suite if I do this. There are 6 failures for non-double types, those that are presently taking advantage of only the string-to-double inexactness of conversion.

The downside of this change is that it makes a value like double 3.14159 illegal to write, since the fully exact decimal value is rather tedious to write. I can see people desiring that inexact conversions be legal, and am willing to make these legal if desired, but I do think that underflow and overflow (e.g., 1e999999) should be a parse error for a floating-point literal in that case.

Summary of proposed syntax for floating-point literals

This is a summary of all of the possible ways to express a floating-point literal, both old and new, that I propose:

  • [+-]?\d+[.]\d*([eE][+-]?\d+)? i.e., the current floating-point literal syntax. Note that the decimal point is necessary, but trailing digits and the exponent field are option.
  • [+-]?0x[0-9a-fA-F]+[.][0-9a-fA-F]*([pP][+-]?\d+)? the C syntax for hexadecimal floating-point literals, except that the decimal point is again necessary.
  • pinf, ninf Positive and negative infinity
  • nan The preferred qNaN value, i.e., sign is positive, quiet bit is set, rest of the payload is all 0’s.
  • nan(0[xX][0-9a-fA-f]+) The NaN value having a given payload, which should be specified in hexadecimal
  • bitcast (i64 0xabcdef09 to double) Technically not a floating-point literal, but this is how noncanonical values should end up getting spelled.

Removing support for parsing the existing floating-point formats doesn’t sound like a good idea. It introduces a form of inconsistency that’s hard to handle: some constants can be represented as a simple literal, and some can’t. This seems like it’s going backwards in terms of reducing the usage of constant expressions. And it breaks backward-compatibility in a way that’s hard to automatically fix. I’d really just rather not touch any of this stuff if we don’t need to.

Not sure about enforcing the “exact” rule; people seem perfectly fine with non-exact literals in other languages. But the impact here is probably not that big either way; most IR is generated by LLVM tools, and LLVM tools shouldn’t generate inexact literals. Maybe include a flag to disable the strictness, in case anyone runs into issues.

Adding support for parsing C/C++-style hexadecimal floats seems like a good idea. Making LLVM IR printing default to that style for seems like a good idea.

Not sure about enforcing the “exact” rule; people seem perfectly fine with non-exact literals in other languages.

This is not a programming language, it is an IR format. So, it’s more reasonable to enforce such a rule here than in a typical programming language.

That said…maybe we should drop the rule.

In concert with that, I think it might be best if we never actually printed a decimal floating-point literal (for binary floats), and instead always used the hexfloat syntax. That’s trivially guaranteed to be unambiguous and exact.

I’d want to preserve support for parsing decimal floating-point literal syntax, still, just discourage its use.

  • nan(0[xX][0-9a-fA-f]+) The NaN value having a given payload, which should be specified in hexadecimal

How is the sign-bit specified in this format? Maybe we should use inf, -inf, nan(...), and -nan(...)?

bitcast (i64 0xabcdef09 to double) Technically not a floating-point literal, but this is how noncanonical values should end up getting spelled.

Agreed with efriedma this is not a good idea: we should continue to support a way to write a “raw” literals for float constants. I think the current syntax is pretty bad. Having different “codes” seems like it’s unnecessary, though – and it’s also weird that the syntax puts the distinguisher code after the 0x instead of before, as we do with integers.

Maybe we could use something like f0x1234 for raw floating-point values (double f0x1234, float f0x1234, bfloat f0x1234, etc would all be valid constants).

This is not a programming language, it is an IR format. So, it’s more
reasonable to enforce such a rule here than in a typical programming
language.

That said…maybe we should drop the rule.

It’s also possible to make it a warning instead of an error, since the
parser does theoretically support that capability (not that anything
uses it at the moment). That said, the utility of diagnosing inexact
literals is pretty low.

In concert with that, I think it might be best if we never actually
printed a decimal floating-point literal (for binary floats), and
instead always used the hexfloat syntax. That’s trivially guaranteed
to be unambiguous and exact.

+1

I’d want to preserve support for parsing decimal floating-point
literal syntax, still, just discourage its use.

  * |nan(0[xX][0-9a-fA-f]+)| The NaN value having a given payload,
    which should be specified in hexadecimal

How is the sign-bit specified in this format? Maybe we should use
|inf|, |-inf|, |nan(…)|, and |-nan(…)|?

I was borrowing the idea from @mshockwave’s open PR, and I see he didn’t
have a solution for NaN sign bits in his proposal (or any of the other
discussion, FWIW). I haven’t attempted to code up this part of the stuff
myself, and lexing the signs correctly is probably going to be a little
bit more annoying, but it should be workable.

Agreed with efriedma this is not a good idea: we should continue to
support a way to write a “raw” literals for float constants. I think
the current syntax is pretty bad. Having different “codes” seems like
it’s unnecessary, though – and it’s also weird that the syntax puts
the distinguisher code /after/ the 0x instead of before, as we do with
integers.

Maybe we could use something like |f0x1234| for raw floating-point
values (|double f0x1234|, |float f0x1234|, |bfloat f0x1234|, etc would
all be valid constants).

Needing a code for every single floating-point type is the thing that I
really object to about the current syntax. I’ll try prototyping the a
f0x-based literal and see how that works.

I don’t really agree with this part. The primary consumer for textual IR are humans, and the decimal representation is a lot easier to understand.

I’ve got no objection to cleaning up this syntax, as suggested downthread. But I’m also not keen on the idea of removing the ability to directly specify the bit pattern altogether. It would seem like an awkward gap IMHO, even if it’s not much harder to produce the desired value via bitcast.

I remember having a difficult time dealing with signaling NaN in my PR because unlike qNaN there isn’t a unique sNaN. But I think I’m fine with the syntax you have here, which completely leaves to the users on deciding the payload. Sure, we might need to falling back to hexadecimal and/or the new format you proposed here when it comes to sNaN but I reckon that will be relative rare.

Speaking of slight inconvenience on falling back to hexadecimal on some NaN/Inf values…

I couldn’t remember the details but I definitely tried that (i.e. try to lex something like -inf) and it was not fun. Though those PRs were my personal project (after copy-and-past “0x7FF8000000000000” for the 1000-th time with rage) so my goal was relatively simple and didn’t want to make radical changes. I can take a look again and see if that’s still true.

qnans are also not unique

I believe the previous discussion with snan is that it requires a non-zero payload. qnan can have a zero payload which is what we chose.

Needing a code for every single floating-point type is the thing that I
really object to about the current syntax. I’ll try prototyping the a
|f0x|-based literal and see how that works.

Okay, I’ve got a prototype for the f0x-based literals for hexadecimal
formats, and discovered two things:

  • The current syntax for ppc_fp128 and fp128 store the hex literal
    as a little-endian array of two 64-bit big-endian integers, which is,
    naturally, not mentioned in the LangRef.

  • The MIR parser reuses the AsmWriter for constant FPs, which means we’d
    also need to touch the MIR parser here…

I’m on board with using f0x literals for a hexadecimal format, but it
looks I’d have to differ the AsmWriter logic to a later patch (if
nothing else, it vastly reduces the amount of necessary test changes in
the first PR).