LLDB Evolution

LLDB has come a long way since the project was first announced. As a robust debugger for C-family languages and Swift, LLDB is constantly in use by millions of developers. It has also become a foundation for bringing up debugger support for other languages like Go and RenderScript. In addition to the original macOS implementation the Linux LLDB port is in active use and Windows support has made significant strides. IDE and editor integration via both SB APIs and MI have made LLDB available to even more users. It’s definitely a project every contributor can be proud of and I’d like to take a moment to thank everyone who has been involved in one way or another.

It’s also a project that shows some signs of strain due to its rapid growth. We’ve accumulated some technical debt that must be paid off, and in general it seems like a good time to reflect on where we’ll be headed next. We’ve outlined a few goals for discussion below as well as one more short-term action. Discussion is very much encouraged.

Forward-Looking Goals

  1. Testing Strategy Evaluation

Keeping our code base healthy is next to impossible without a robust testing strategy. Our existing suite of tests is straightforward to run locally, and serves as a foundation for continuous integration. That said, it is definitely not exhaustive. Obvious priorities for improvement include gathering coverage information, investing in more conventional unit tests in addition to the suite of end-to-end tests, and introducing tests in code bases where we depend on debugger-specific behavior (e.g.: for expression evaluation.)

  1. C++ Module Support

LLDB takes advantage of Clang modules for type information and expression evaluation. This has been used extensively for C and Objective-C scenarios, but Clang C++ module support is now mature enough that we can extend our support accordingly. Fully embracing C++ modules will enable LLDB expressions to take advantage of template declarations and other constructs that are better represented by declarations than the artifacts produced during compilation.

  1. Establishing Language Integration Standards

As more languages build on LLDB’s foundation the project runs the risk of growing deep dependencies on a wide variety of compilers and runtimes. The community needs to engage in a constructive conversation about how best to keep the core of LLDB clean and allow language support to be plugged in. Whether this should occur at compile-time or runtime and how best to organize repositories and branches to meet the needs of our diverse community will be an ongoing topic of discussion.

  1. Good Citizenship in the LLVM Community

Last, but definitely not least, LLDB should endeavor to be a good citizen of the LLVM community. We should encourage developers to think of the technology stack as a coherent effort, where common code should be introduced at an appropriate level in the stack. Opportunities to factor reusable aspects of the LLDB code base up the stack into LLVM will be pursued.

One arbitrary source of inconsistency at present is LLDB’s coding standard. That brings us to…

Near-Term Goal: Standardizing on LLVM-style clang-format Rules

We’ve heard from several in the community that would prefer to have a single code formatting style to further unify the two code bases. Using clang-format with the default LLVM conventions would simplify code migration, editor configuration, and coding habits for developers who work in multiple LLVM projects. There are non-trivial implications to reformatting a code base with this much history. It can obfuscate history and impact downstream projects by complicating merges. Ideally, it should be done once with as much advance notice as is practical. Here’s the timeline we’re proposing:

Today - mechanical reformatting proposed, comment period begins

To get a preview of what straightforward reformatting of the code looks like, just follow these steps to get a clean copy of the repository and reformat it:

  1. Check out a clean copy of the existing repository
  2. Edit .clang-format in the root of the tree, remove all but the line “BasedOnStyle: LLVM”
  3. Change your current working directory to the root of the tree to reformat
  4. Double-check to make sure you did step 3 :wink:
  5. Run the following shell command: "find . -name “*.[c,cpp,h] -exec clang-format -i {} +”

Aug 20th - comment period closes, final schedule proposed
TBD (early September?) - patches land in svn

The purpose of the comment period is to review the straightforward diffs to identify areas where comment pragmas should be used to avoid undesirable formatting (tables laid out in code are a classic example.) It’s also a time when feedback on the final timetable can be discussed, and any unforeseen implications can be discovered. We understand that LLDB tends toward relatively long names that may not always work well with the LLVM convention of wrapping at 80 columns. Worst case scenarios will be evaluated to determine the desired course of action.

Kate Stone k8stone@apple.com
 Xcode Low Level Tools

LLDB has come a long way since the project was first announced. As a robust debugger for C-family languages and Swift, LLDB is constantly in use by millions of developers. It has also become a foundation for bringing up debugger support for other languages like Go and RenderScript. In addition to the original macOS implementation the Linux LLDB port is in active use and Windows support has made significant strides. IDE and editor integration via both SB APIs and MI have made LLDB available to even more users. It’s definitely a project every contributor can be proud of and I’d like to take a moment to thank everyone who has been involved in one way or another.

It’s also a project that shows some signs of strain due to its rapid growth. We’ve accumulated some technical debt that must be paid off, and in general it seems like a good time to reflect on where we’ll be headed next. We’ve outlined a few goals for discussion below as well as one more short-term action. Discussion is very much encouraged.

Forward-Looking Goals

  1. Testing Strategy Evaluation

Keeping our code base healthy is next to impossible without a robust testing strategy. Our existing suite of tests is straightforward to run locally, and serves as a foundation for continuous integration. That said, it is definitely not exhaustive. Obvious priorities for improvement include gathering coverage information, investing in more conventional unit tests in addition to the suite of end-to-end tests, and introducing tests in code bases where we depend on debugger-specific behavior (e.g.: for expression evaluation.)

I know this is going to be controversial, but I think we should at least do a serious evaluation of whether using the lit infrastructure would work for LLDB. Conventional wisdom is that it won’t work because LLDB tests are fundamentally different than LLVM tests. I actually completely agree with the latter part. They are fundamentally different.

However, we’ve seen some effort to move towards lldb inline tests, and in a sense that’s conceptually exactly what lit tests are. My concern is that nobody with experience working on LLDB has a sufficient understanding of what lit is capable of to really figure this out.

I know when I mentioned this some months ago Jonathan Roelofs chimed in and said that he believes lit is extensible enough to support LLDB’s use case. The argument – if I remember it correctly – is that the traditional view of what a lit test (i.e. a sequence of commands that checks the output of a program against expected output) is one particular implementation of a lit-style test. But that you can make your own which do whatever you want.

This would not just be busy work either. I think everyone involved with LLDB has experienced flakiness in the test suite. Sometimes it’s flakiness in LLDB itself, but sometimes it is flakiness in the test infrastructure. It would be nice to completely eliminate one source of flakiness.

I think it would be worth having some LLDB experts sit down in person with some lit experts and brainstorm ways to make LLDB use lit.

Certainly it’s worth a serious look, even if nothing comes of it.

  1. Good Citizenship in the LLVM Community

Last, but definitely not least, LLDB should endeavor to be a good citizen of the LLVM community. We should encourage developers to think of the technology stack as a coherent effort, where common code should be introduced at an appropriate level in the stack. Opportunities to factor reusable aspects of the LLDB code base up the stack into LLVM will be pursued.

One arbitrary source of inconsistency at present is LLDB’s coding standard. That brings us to…

Near-Term Goal: Standardizing on LLVM-style clang-format Rules

We’ve heard from several in the community that would prefer to have a single code formatting style to further unify the two code bases. Using clang-format with the default LLVM conventions would simplify code migration, editor configuration, and coding habits for developers who work in multiple LLVM projects. There are non-trivial implications to reformatting a code base with this much history. It can obfuscate history and impact downstream projects by complicating merges. Ideally, it should be done once with as much advance notice as is practical. Here’s the timeline we’re proposing:

Today - mechanical reformatting proposed, comment period begins

To get a preview of what straightforward reformatting of the code looks like, just follow these steps to get a clean copy of the repository and reformat it:

  1. Check out a clean copy of the existing repository
  2. Edit .clang-format in the root of the tree, remove all but the line “BasedOnStyle: LLVM”
  3. Change your current working directory to the root of the tree to reformat
  4. Double-check to make sure you did step 3 :wink:
  5. Run the following shell command: "find . -name “*.[c,cpp,h] -exec clang-format -i {} +”

Very excited about this one, personally. While I have my share of qualms with LLVM’s style, the benefit of having consistency is hard to overstate. It greatly reduces the effort to switch between codebases, a direct consequence of which is that it encourages people with LLVM expertise to jump into the LLDB codebase, which hopefully can help to tear down the invisible wall between the two.

As a personal aside, this allows me to go back to my normal workflow of having 3 edit source files opened simultaneously and tiled horizontally, which is very nice.

FWIW, as a happy lldb user, it's exciting to read about these planned
developments.

I have two follow-up questions about the section on 'Testing Strategy
Evaluation': (1) what is lldb's policy on including test updates with bug fix
commits and functional changes, and (2) how is this policy enforced?

AFAICT, it seems that lldb is not as strict about its test policy as other llvm
sub-projects. That could be to its detriment. Here are some very rough numbers
on the number of commits which include test updates [1]:

  - lldb: 287 of the past 1000 commits
  - llvm: 511 of the past 1000 commits
  - clang: 622 of the past 1000 commits
  - compiler-rt: 543 of the past 1000 commits

NFC commits make these numbers a bit noisy. But, unless lldb has a much higher
ratio of NFC commits to functional changes as compared to other llvm
sub-projects, this is a concerning statistic.

best,
vedant

[1] Based on ToT = r278069.

HAS_TEST=0
TOTAL=1000
for HASH in $(git log --oneline -n$TOTAL | cut -d' ' -f1); do
  git show --stat $HASH | grep "|" | grep -q test && HAS_TEST=$((HAS_TEST+1))
done
echo $HAS_TEST "/" $TOTAL

There are a lot of reasons for the lack of tests. Off the top of my head, two of the biggest ones are:

  1. Some areas of LLDB have been historically hard to test. The unwinder and core dumps come to mind. You can’t really just check in an 800MB core dump into the repo.
  2. Tests are very heavyweight. You have to write a Makefile. You have to write a Python script that uses the SB API. you have to write a C program. Then you have to figure out the right incantation of dotest.py to run your test. Once you’ve done this many times it becomes easier, but the barrier to entry is high.

#1 can be improved through the use of unit tests and fuzzing. Sure, it’s hard to generate a full blown executable that contains every edge case the unwinder might ever experience and then have it try to unwind through there. Especially considering that the unwinder behaves differently on every platform. But it’s much more manageable to write a unit test that constructs a particular sequence of bytes in memory, passes it to the unwinder, and checks the return value of some function that is supposed to handle that. It’s still not entirely simple, since the unwinder is partially heuristic, but at least it’s more manageable.

There are also ways we can write IR by hand and have llc generate some byte code for us and then pass that to those same functions. Again, it’s not like we can just start doing this overnight, but there are ways. The question is just how serious of an effort are we prepared to make and how much time are we prepared to put into making this testable versus implementing new features, fixing bugs, etc.

#2 could potentially be improved by lit style tests. As I mentioned in my last post, think lldbinline style tests. Not appropriate for everything, but certainly for a lot of things. If you only had to have one file which is a C program with some annotations in the file, that is a much lower barrier to entry.

Again, the real question is just how much effort are we actually prepared to put into this? I’d love it if there were entire days or weeks that were just testing weeks, where all we did was add new tests (or refactor code to make it more testable) and people didn’t work on anything else. I’ve been inactive for a while because I’ve had to prioritize work on some things in LLVM, but I could make time for something like that.

I fully welcome this move and agree with the timeline. Woohoo.

I also agree with assessment of the situation regarding testing. I
think we're going to need to devote thought to testing if we're going
to move closer towards llvm.

pl

There are a lot of reasons for the lack of tests. Off the top of my head, two of the biggest ones are:

1) Some areas of LLDB have been historically hard to test. The unwinder and core dumps come to mind. You can't really just check in an 800MB core dump into the repo.

Side note: core dumps are often higly compressible, so they may have quite reasonable size e.g. in .xz format

One thing we will need to take a look at is functions which have a very deep indentation level. They have the potential to be made really ugly by clang-format. The default indentation will be reduced from 4 to 2, so that will help, but I recall we had some lines that began very far to the right.

Here’s a little bash command shows all lines with >= 50 leading spaces, sorted in descending order by number of leading spaces.

grep -n ‘^ +’ . -r -o | awk ‘{t=length($0);sub(" *$“,”“);printf(”%s%d\n", $0, t-length($0));}’ | sort -t: -n -k 3 -r | awk ‘BEGIN { FS = “:” } ; { if ($3 >= 50) print $0 }’

It’s less useful than I was hoping because most of the lines are noise (line breaking in a function parameter list).

If there were a way to detect indentation level that would be better. Mostly just to identify places that we should manually inspect after running clang-format.

Another thing worth thinking about for the long term is library layering and breaking the massive dependency cycle in LLDB. Every library currently depends on every other library. This isn’t good for build times, code size, or reusability (especially where size matters like in lldb-server). I think the massive Python dependency was removed after my work earlier this year. But I’m not sure what the status of that is now, and there’s still the rest of LLDB.

In the future it would be nice to have a modules build of LLDB. And sure, we could just have liblldb be one giant module, but that seems to kind of defeat the purpose of modules in the first place.

For unit tests in particular, it’s nice to be able to link in just the set of things you need, and that’s difficult / impossible right now.

I ran clang-format and tried to build and got a bunch of compiler errors. Most of them are order of include errors. I fixed everything in the attached patch. I doubt this will apply cleanly for anyone unless you are at the exact same revision as me, but at least you can look at it and get an idea of what had to change.

The #include win32.h thing is really annoying and hard for people to remember the right incantation. I’m going to make a file called Host/PosixApi.h which you can just include, no matter what platform you’re on, and everything will just work. That should clean up a lot of this nonsense.

reformat.patch (18.9 KB)

Great catch! If refactoring along those lines doesn’t clean up 100% of the cases then it’s worth explicitly breaking up groups of #include directives with comments. clang-format won’t reorder any non-contiguous groups and it’s a great way to explicitly call out dependencies.

Ideally we should be focused on committing changes along these lines so that there’s no post-format tweaking required to build cleanly again.

Kate Stone k8stone@apple.com
 Xcode Low Level Tools

Agreed that better layering and modularization is a worthwhile longer-term goal. These kinds of changes should be researched and proposed for discussion because code reorganization can be extremely disruptive (which is why we opened this comment period for the proposed reformatting!) I’d generally argue for focusing on breaking out larger blocks of functionality where focused testing will yield results rather than a lot of fine-grained change.

Kate Stone k8stone@apple.com
 Xcode Low Level Tools

Standardizing include order would really help. LLVM’s style is documented here, but to quote it:

  1. Main Module Header

  2. Local/Private Headers

  3. llvm/...

  4. System #includes
    So perhaps it would be reasonable for us to standardize on something like this:

  5. Main Module Header

  6. Local/Private Headers

  7. lldb/...

  8. llvm/...

  9. System #includes

Putting LLDB headers before LLVM headers is a nice way to enforce “include what you use”. Otherwise you might have something like this:

// lldb.cpp
#include “llvm/LLVM1.h”
#include "lldb/lldb2.h

// lldb2.h
llvm::MyType foo();

And this will work, even though lldb2.h needs to have #include “llvm/LLVM1.h”

So the above ordering fixes that. clang-format won’t solve this for us automatically, but it would be nice to at least move towards it.

Note that step 5 below is wildly inaccurate, a placeholder while drafting the note. The actual command to reformat should be:

find . ( -iname “.c" -or -iname ".cpp” -or -iname “*.cpp” ) -exec clang-format -i {} +

If you’re curious where the longest lines in the original or resulting source come from you can try the following in an unformatted or post-formatting repository:

find . ( -iname “.c" -or -iname ".cpp” -or -iname “*.cpp” ) -exec awk ‘{ if (length($0) > max) max = length($0) } END { print max " {}" }’ {} ; | sort -nr

Kate Stone k8stone@apple.com
 Xcode Low Level Tools

I've started to do some cleanup of the include header order (r278222).
It doesn't get everything compiling after a clang-format, but it got
me about half way.

+1 on the include order proposed by Zach. I agree that better
modularization would improve testability and have a positive impact on
the size of lldb-server (e.g., right now we have to link it against
libclang, even though it shouldn't need any of that functionality).

pl

This makes sense to me, and matches what clang does as well. I think that this is clearly in the spirit of the llvm include order standards, and I think it would be great to make this explicit in the coding standard doc. Can you send in a patch to update it to make this explicit? I’ll review it.

-Chris

#2 could potentially be improved by lit style tests.

+1 to this.

Again, the real question is just how much effort are we actually prepared to put into this? I'd love it if there were entire days or weeks that were just testing weeks, where all we did was add new tests (or refactor code to make it more testable) and people didn't work on anything else. I've been inactive for a while because I've had to prioritize work on some things in LLVM, but I could make time for something like that.

I think it makes sense to start with the big mechanical changes and try to do it in a single “change the world” commit to avoid disrupting version control history too much. After that is done, the many various refactorings (fixing dependence cycles, improving the testing situation to be more lit like, sinking functionality into LLVM and reusing existing LLVM functionality more, etc) can all be done in parallel and independently over time.

-Chris

I just committed another header cleanup commit, which makes lldb
clang-format-immune ( = it still compiles after a full reformat) on
linux. Other OS's are still likely to have some missed dependencies.

However, when I tried running the test suite I got about 150 failures.
Based on a sample of the errors, it looks like the problem is that
clang format messes up the "// Place breakpoint here" annotations we
use in the tests.
Therefore, I propose to apply the clang-format to the lldb source code
only as a first step. After that, as a second step, we can go through
the tests and fix them up so that the comment markers are where we
expect them to be.

pl

Hi Pavel,

Would it make sense to address this problem by fixing clang-format,
rather than working around it?

(Assuming the clang-format fix is relatively easy, and acceptable to
clang-format's maintainers.)

- Christian

It’s not possible. The problem is that lldb was dependent on order of includes because each header wasn’t properly including what it used. So when clang-format reordered this, things broke

I assume Christian Convey was referring to clang-format moving the “//breakpoint here” comments in the tests to different lines.