Skip to content

Multi-line comments on one line double the indent each time, up to an OutOfMemoryError #36

Description

@abashev

What

Reported upstream as google/google-java-format#413 (open since 2019-11-05): a 6 kB file of commented-out code ran out of memory with a 12 GB heap, and its output would have been about 3 GB. It reproduces here on main (53ac7d8). The shape is block comments that end and start on the same line, as in commented-out code that had javadoc in it:

/*
class X {
	*//** javadoc *//*
	public void foo(Bar bar) {}

	*//** javadoc *//*
	public void foo(Bar bar) {}

	*//** javadoc *//*
	public void foo(Bar bar) {}

	*//** javadoc *//*
	public void foo(Bar bar) {}
}*/

The continuation lines of these comments come out indented by 17, 50, 116 and 248 columns: each comment on the line starts twice as far right as the one before it. The file size grows the same way:

comments input output
6 317 B 4.3 kB
10 517 B 68 kB
16 817 B 4.3 MB

Why it happens

Comment.computeBreaks moves the column past a comment by adding the length of the comment's last line:

int firstLineLength = text.length() - Iterators.getLast(Newlines.lineOffsetIterator(text));
return state.withColumn(state.column() + firstLineLength)

For a one-line comment that is right. For a multi-line one, JavaCommentsHelper.rewrite has just re-indented the continuation lines to start at state.column(), so the last line already contains the column, and it is counted twice. The next comment on that line is re-indented from the doubled column, and so on. Despite its name, firstLineLength holds the length of the last line.

google-java-format has the same code in Doc.Tok.computeBreaks.

Fix

After a comment that spans lines, the column is the length of its last line; after a one-line comment it stays column + length.

The output changes only where something follows a multi-line comment on that comment's last line, so a corpus run should show next to nothing.

Done when

  • a golden of the file above indents each comment's continuation lines to the column where that comment starts, and a second run changes nothing;
  • a test formats twenty such comments into a small output (the old code writes about 70 MB);
  • a corpus run says how many files change.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions