Before starting to reverse complex programs, you have to learn the foundations, and one of these foundations is the set of tools available to you. Nowadays, there are complex graphical solutions that can do almost everything, however, it is very useful to know how to use the more basic and generic CLI tools.
As I just said, you can find a tool that does everything for you and happily be stuck with it. Tools like Ghidra or IDA are powerful and will save you a lot of time, but if you just learn to use a specific tool, you will miss a little of what’s going on behind the scenes. And that is really where all the fun is.
There are a few extra advantages about using the classical CLI tools:
- They are evergreen. CLI tools have been around almost from the beginning (what UNIX people know as epoch) and are still available. Other tools become popular, then another tool that looks nicer appears, and the old one gets obsolete, so you have to learn the new one.
- Awareness. Modern, advanced tools provide powerful functions available with a mouse click. You never know if what they do is a very basic action or a complex algorithm composed of multiple steps. This lack of knowledge gets you completely unarmed when the “magical” tool option fails. Those options are great, but they are greater when you know what they do under the hood so you can react when something goes wrong.
- Architectures. Even when the more powerful tools support more and more platforms every day, the reality is that new platforms or niche platforms, when supported, are usually only supported by the open-source toolchains. The fact that these are open source makes it much easier for companies to add support for those platforms, but that will always come later, when the platform is known to be popular. Some niche platforms may never make it into those tools.
- Flexibility. Last but not least, CLI tools are built to be combined together to produce powerful output as well as making the integration in report generation scripts or automation much more easy. Top-tier tools also provide these functionalities, however each one has its own details, which means that there is something new to learn every time somebody decides that a change is needed.
At this point, you can stop reading and keep going with Ghidra, IDA, Binary Ninja, or whatever tool you like to use, or take a small break and dive into the very basic tools available on all systems and, hopefully, you will come back to those tools with new knowledge that will allow you to take the most out of them.
If that resonates with you, let’s get started!
Basic File Information
The first step when you face a new, unknown binary to analyze is to get as much information from it as you can. Usually, each analysis starts with the execution of this sequence of commands. I will use RE Challenge 1 as an example to illustrate most of the tools in this chapter, so go grab a copy.
The first command you will usually execute is file:
$ file challenge01.x86_64.bin
code/challenge01/challenge01.x86_64.bin: ELF 64-bit LSB pie executable, x86-64, version 1 (SYSV), dynamically linked, interpreter /lib64/ld-linux-x86-64.so.2, BuildID[sha1]=e41a3c906ff47b5b259e49871f1d1ba4d6772154, for GNU/Linux 3.2.0, not stripped
This command tells us a lot about the binary. In this case, it says:
- It is a 64-bit Little Endian (LSB) PIE (Position Independent Executable) for the x86_64 architecture using version 1 of the System V ABI.
- It tells us that it is a dynamic binary and which dynamic loader it uses.
- It gives us a hash value for the specific build of this binary.
- The minimal Linux ABI supported (in this case, it requires kernel 3.2 or newer).
- And last but not least, it lets us know if the binary was stripped, i.e., all symbols and debug information were removed.
However, note that many of this data are for information only and are added to the binary during the build process. In other words, there is nothing in the system (at least at the time of this writing) that will ensure that data is honored. This means that you can change the build hash or the ABI supported version without any consequences and the binary will run fine, as far as the system is compatible at the binary level.
This information is stored in PT_NOTE program headers
and, for non-stripped binaries, organized in sections of type
SHT_NOTE that are mapped to those PT_NOTE
program headers and can be easily found inside the file. We’ll come back
to this later.
A special comment about the buildId is required. This
hash is produced by the linker when producing the executable and it’s
not a hash of the final binary. Think about it: the final binary has to
contain the hash, so calculating the hash and then storing it would
change the hash. For some weak hash algorithms (as md5), this can be
done, using what is called a hash collision, but that’s something you
don’t want to happen in your systems. So, the linker basically hashes
some parts of the binary (code, data segments, etc.). The order and
parts of the ELF file being hashed depend on the used linker, and
there’s no easy way to calculate it afterwards. The BuildID hash is used
to identify a specific build and it’s intended for internal use by the
company or organization developing the software. It was never intended
to be an integrity check of the file.
So, as I said, all this information is stored in
SHT_NOTE ELF sections mapped into one or more corresponding
PT_NOTE program headers. They can all be shown using the
command:
$ readelf -n challenge01.x86_64.bin
Displaying notes found in: .note.gnu.property
Owner Data size Description
GNU 0x00000010 NT_GNU_PROPERTY_TYPE_0
Properties: x86 ISA needed: x86-64-baseline
Displaying notes found in: .note.gnu.build-id
Owner Data size Description
GNU 0x00000014 NT_GNU_BUILD_ID (unique build ID bitstring)
Build ID: e41a3c906ff47b5b259e49871f1d1ba4d6772154
Displaying notes found in: .note.ABI-tag
Owner Data size Description
GNU 0x00000010 NT_GNU_ABI_TAG (ABI version tag)
OS: Linux, ABI: 3.2.0The listing above shows the contents of the different
SHT_NOTE sections that you can list using the following
command:
$ readelf -SW challenge01.x86_64.bin | grep note
[ 1] .note.gnu.property NOTE 0000000000000350 000350 000020 00 A 0 0 8
[ 2] .note.gnu.build-id NOTE 0000000000000370 000370 000024 00 A 0 0 4
[19] .note.ABI-tag NOTE 0000000000002154 002154 000020 00 A 0 0 4
Heavily stripped programs, like the ones produced by the tool
sstrip from ELFKickers, will get rid of the whole section
table. In that case, readelf -n is still able to show the
notes but won’t reference the section they belong to because the section
no longer exists.
It’s relatively easy to add as many notes as you want during the build process (either using compiler attributes or a linker script). It’s also relatively easy to overwrite existing note sections with extra values to inject code. However, adding new notes once the binary is created, even when possible, is a bit more tricky. But enough about notes for now.
Hex Dumping (xxd)
There are several tools for producing hexadecimal dumps of a file.
hexdump or od are two of the oldest and more
commonly available, however, xxd has become the de facto
standard when it comes to hexadecimal dumping files. But, in case you
don’t have it in the target system and you can’t install it,
od or hd. By the way, od stands
for Octal Dump, and hd is an alias for
hexdump.
If you want to use od you’d likely do it like this:
$ od -A x -x challenge01.x86.bin
This forces od to dump offsets and data in hexadecimal.
Default is octal, of course.
xxd does more or less the same thing, but has a few
interesting options that make its use more convenient in some
situations. Personally, I normally use xxd together with
dd but in reality xxd offers command-line
options to specify which part of the file to dump. For example, let’s
dump the buildId note that we just recently saw. For that,
we need the file offset that can be obtained from the section list, for
example:
$ readelf -SW challenge01.x86_64.bin | grep build
[ 2] .note.gnu.build-id NOTE 0000000000000370 000370 000024 00 A 0 0 4
$ xxd -s $((0x370)) -l $((0x24)) challenge01.x86_64.bin
00000370: 0400 0000 1400 0000 0300 0000 474e 5500 ............GNU.
00000380: e41a 3c90 6ff4 7b5b 259e 4987 1f1d 1ba4 ..<.o.{[%.I.....
00000390: d677 2154 .w!T
Notes have a very specific format. Let’s break it up:
04 00 00 00 -> Size of the Note name
14 00 00 00 -> Size of the Desc name (data 20 bytes in this case)
03 00 00 00 -> Type of note (NT_GNU_BUILD_ID) in this case
After that, we find the note name that is 4 bytes long, according to the previous dump:
47 4e 55 00 -> "GNU"
And after that, the buildId hexadecimal values follow.
You can just compare it with the value we obtained from
file at the beginning of this chapter.
xxd has a bunch of useful command-line flags, but two of
the most interesting ones are the following:
-pprocesses hexadecimal sequences without any formatting, just prints the data as is.-rdoes the reverse process; given a hexdump, it generates the binary data associated.
Additional options to group bytes (-g) or deal with
endianness are also available for us to tweak xxd output to
match our needs.
xxd and hd dumps ASCII representation in
addition to the hexadecimal values and therefore they will show any
plain string stored in the file, like symbols, literals, and so forth.
However, for extracting strings, there are better options.
Strings
The usual next step after getting the very basic information about
the binary we’ve just seen is to look for strings in the program. The
strings utility scans the whole file looking for printable
strings, which are surprisingly useful during reverse engineering.
To deceive strings, a program needs to encode any text
it uses to interact with the user or to access other program
functionalities. This is pretty easy to do, but it’s an extra step
during development that may be a bit annoying, so, sometimes, programs
just obfuscate some strings and leave other alone. In any case, you
should also pass the program through the strings tools just
in case. You may be surprised about what you will find.
Among other things, you will see any plain text string used by the program, including format strings, as well as names of symbols, compiler version, section names, and other information; it all depends on the effort put by the developer to hide things.
strings is easy to use and straightforward, but for big
dynamic programs, it may produce a massive output. In those cases, we
can use the tool objdump to dump specific sections of the
program. For example, for our simple challenge01 program,
strings produces 90 lines.
$ strings challege01.x86.bin | wc -l
90
But we can just dump the .rodata section and avoid all
the symbol names using objdump like this:
$ objdump -j .rodata -s challenge01.x96.bin
This will produce a hexadecimal dump instead of a list of strings one in each line, but it would be much shorter and will provide us with the offsets where those strings will be found.
strings provides a few interesting parameters that
should be considered further. The first one is -d that will
just print strings from initialized, loaded data sections in the
file.
$ strings -s challege01.x86.bin | wc -l
25
As you can see, the reduction is considerable. Also, the
strings tool looks for sequences of printable characters
with a minimal length of 4. This value may be changed using the flag
-n if we want to look for longer strings. There is one last
thing to consider about strings. Some platforms like
Windows, applications using UTF-16 encoding, or applications using wide
chars may puzzle strings. The -e flag allows
us to define the character encoding strings should use. The
possible values that will follow
s | Single 7-bit-byte (default)
S | Single 8-bit-byte
b | 16-bit Big Endian
l | 16-bit Little Endian
B | 32-bits Big Endian
L | 32-bits Little Enfina
Take a look to the man page for strings to check all the
possible flags, but these are the more relevant while reverse
engineering.
To conclude this section, I’d like to mention my own tool
maca. You can download the source code from GitHub and
study it. It’s a small and simple application, so it’s very easy to
follow. One of the options that maca provides is to dump
strings. It just dumps regular ASCII strings and only allows to specify
the minimal size, but for many cases that’s good enough. The advantage
of maca is that it shows the section and offset of each
string it found. For example, when run on challenge01 it
will produce lines like:
$ maca -s challenge01.x86.bin
[ .rodata] 0x002007 [00b]: SuperSecret
[ .rodata] 0x002017 [020]: Reverse Engineering Challenge #1
The output is colored so you can easily identify strings in the
.rodata section, or in the symbol table.
Dynamic Binaries
One of the information provided by file was if the
binary was statically or dynamically linked. Overall, dynamically linked
binaries are smaller as they make extensive use of system libraries,
while static binaries have to include all the code they need in the
binary, making them much bigger.
In addition to the use of file (which is enough for
getting this information), we can use readelf or
maca.
$ readelf -h challenge01.x86_64.bin | grep Type
Type: DYN (Position-Independent Executable file)
The flag -h shows the ELF header of the indicated
binary. In this case, we’re just showing the Type field,
but you could skip the grep and take a look at the whole
header. We’ll do this in detail a bit later. An alternative is to use
maca that provides a more compact output of the ELF
header:
$ maca -h challenge01.x86_64.bin | grep Description
Description : [64 bits] [AMD x86-64] [LSB] [System V] | ABI Version [0] [ET_DYN : Shared object]
It just shows in a single line the type, the architecture, the endianness, and the ABI, but, as I’ve just said, you would usually look at the whole header instead of filtering fields, unless you want to write some kind of script that requires this information separately.
For static binaries, the type will be DT_EXEC or just
EXEC in readelf, and that’s it, but for
dynamically linked binaries, we can use a few extra tools to get extra
information from the binary before executing it or even looking into its
ASM. The first one to check is ldd.
$ ldd challenge01
linux-vdso.so.1 (0x00007fee7f230000)
libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007fee7f013000)
/lib64/ld-linux-x86-64.so.2 (0x00007fee7f232000)
The ldd tool shows the dynamic libraries used by the
program. A static program won’t show anything. The three libraries shown
above are the very minimal for C programs:
linux-vdso.so.1. This is the Virtual Dynamic Shared Object library. It is an ELF shared library that the kernel maps into every process address and allows a process to get kernel services without making a system call that takes quite some time. One of the initial uses of this was to substitute thegettimeofdaysystem call which just returns a number.libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6. This is the standard C library used by the program. As this is usually a symbolic link,lddshows the actual location of the library./lib64/ld-linux-x86-64.so.2. This is the dynamic loader required to load the program and to perform tasks like lazy symbol resolution at run-time.
Programs written using other programming languages will show different sets of libraries. Some examples:
$ ldd hello-cpp
linux-vdso.so.1 (0x00007efd10236000)
libstdc++.so.6 => /lib/x86_64-linux-gnu/libstdc++.so.6 (0x00007efd0fe00000)
libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007efd0fc0c000)
libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007efd1011d000)
/lib64/ld-linux-x86-64.so.2 (0x00007efd10238000)
libgcc_s.so.1 => /lib/x86_64-linux-gnu/libgcc_s.so.1 (0x00007efd100f0000)
As you can see, in addition to the libraries used by vanilla C
programs, a very minimal C++ program will also include at least the
libstdc++.so.6 that is the Standard C++ Library.
As a curiosity, ldd has been pretty controversial
because it was, somehow, executing the program in order to list the
libraries it uses. This was exploited in the past to get code executed.
Nowadays, it’s pretty safe, but you can still see the nature of this
utility running your application like this:
$ LD_TRACE_LOADED_OBJECTS=1 challenge01.x86.bin
It just produces the same output as ldd. Talking about
dynamic loader tweaking, there are some environmental variables honored
by ld that triggers different behaviours. A few interesting
ones are:
LD_DEBUG=libs ./binary : Shows all libraries loaded by the dynamic loader
LD_DEBUG=all ./binary : Shows all debug, including symbol resolution
LD_DEBUG_OUTPUT=ld.log LD_DEBUG=libs ./bin: Redirect output to a file ld.log.pid
LD_SHOW_AUCV=1 ./bin : Shows the Auxiliary vector sent to the binary
There’s also a LD_AUDIT which is pretty powerful but not
for this current topic. Something for you to research.
Finally, we can also get a list of dynamic objects used by a running
process by checking the /proc/PID/mappings folder. This is
interesting during dynamic analysis. Note that a binary can use
libraries dynamically using the dlopen, dlsym,
and dlclose functions. In that case, the library is loaded
at run-time, and there is no detail in the binary itself (beyond the
name of the library that may be obfuscated). In other words, none of the
tools we’ve just discussed would show a dynamically loaded library. This
can only be determined either by statically analyzing the binary or
during dynamic analysis.
Other tools worth mentioning
There are a few more tools that I’d mention as they may be useful in certain circumstances.
The first one is floss (FLARE Obfuscated String Solver)
that, sometimes, will be able to extract obfuscated strings. Overall, it
may be worth using it instead of strings as it will find
all the strings strings does and likely a few more.
Other small tool that gives us information about the binary, but not
that useful for reversing it, is size. This little tool can
give us information about the size of code and data of a given binary.
As we’ll see in the next chapter, we’ll get this very same information
from other tools anyway, so, beyond size comparison between binaries, it
isn’t that useful.
Another interesting tool to take into consideration is
exiftool, able to extract metadata from files. Even when
this tool is more interesting when used on multimedia files, it can also
be used on binaries and will give us some general information. Nothing
we can’t get using other tools, but it’s good to remember it exists.
Entropy calculation
Entropy, putting aside the thermodynamics term, is a measurement of the amount of information in a given message, and its value indicates how many bits (when using basis 2 logarithms for calculation) per symbol are required to store such message. For binaries, symbols are bytes, and therefore, we have 256 possible symbols of value, therefore, an entropy of 8 will indicate that we need 8 bits to encode a byte, or in other words, the message is composed of random bytes, and it is not possible to compress it. Therefore, we have to use the maximum amount of bits that allows us to encode all the messages.
In general, there’s some relation between symbols in a message, and the message can be compressed to a certain step. That’s the other way to interpret entropy. From this point of view, a message (or a binary fragment) with an entropy of 4 means that we just need 4 bits to encode all the symbols in that fragment, which means that we can compress the fragment to roughly half of its size.
Entropy is calculated using the following formula:
E = SUM_i p_i * log (1/p_i) = - SUM_i p_i * log(p_i)
The formula may look intimidating, but it’s just a weighted average of the inverse of probabilities of each symbol. If you think about it, the probability of a symbol in a message somehow measures how often that symbol shows up. Actually, in the real world, that’s the only way to calculate probabilities: counting the times a symbol appears and dividing by the total number of symbols, that is, the normalized value of how often the symbol appears. So, if we take the inverse, we’re measuring how rare it is to see a given symbol. Why is this relevant? Because the symbols that appear less frequently are the ones that give us more information.
The meteorological station at Atacama desert
The Atacama desert in South America is one of the driest places in the world. In average, a few mm of rain per year, and some parts of the desert haven’t received any rain for years. Now imagine that you set up a meteorological station there to report rain over a satellite link (actually, that doesn’t really matter; it’s just to add realism).
Our station will sample a sensor and send a message every minute using two symbols: R meaning rain, D meaning Dry. OK, let’s rename them as 1 and 0. According to what we’ve said, the sequence of transmitted information from the station will look like:
0000000000000000... many months later 1100000000000000 ....So, if you have to guess what is the next value you’ll receive from the station, which one would you pick? Yes, zero. It’s most likely that it’s not raining in Atacama at any time in the year. So, the probability of getting a
0is very, very high, almost1, and therefore, from an Entropy point of view, it gives us little to no information. So, in order to measure how much information is in our message, we should consider the inverse, in other words, that number very close to 1, when inverted, will become a number very close to zero. On the other hand, if we suddenly receive a1, that’s a lot of information, because that is a value that we cannot predict easily because it’s very unlikely, and then it should contribute much more to the information in the message than the0s we get all day long.So, this is the conceptual explanation of why the Entropy formula considers the inverse of the probability. Note, however, that as we are using a logarithm, the effect in the final result is just a change of sign. By the way, the use of the logarithm is just to zoom in the data. All our values are between 0 and 1 (they are probabilities), and the logarithm function between 0 and 1 goes from
0at 1 to minus infinite, making the smaller values look much bigger and making the difference much more obvious. This is commonly used in engineering.
Back to reverse engineering, why is the entropy interesting for us? As you’ve likely deduced from all this discussion, a high entropy value means that data looks random, or, if you prefer, that we have a .5 probability of guessing the next symbol. In computer parlance, this may mean two things:
- The first is that the data has been compressed. Compressing data basically consists of that: removing any redundancy from the message so we’re not wasting bits unnecessarily. And that means that we cannot guess what will come next. In information theory, the value of entropy indicates the theoretical maximal compression we can achieve. If entropy is 4, that means that we could compress that message to half its size (for each symbol –1 byte– we just need 4 bits, so we could pack 2 symbols in one byte).
- The second is that the data has been encrypted. Again, when we encrypt data, in general, we want it to look random so there’s no chance to guess any of the values in the message. Cryptographic algorithms usually do this and produce sequences with maximal entropy. This is the reason why you should compress your data before encrypting it. The other way around, you won’t compress the data much.
Both cases are common, especially in malware, where part of the program is compressed or encrypted so it isn’t directly accessible to the researcher right away. The programs doing this are usually known as crypters (when they just encrypt part of the program) or packers (when they compress or pack together other files within the program). In both cases, from the malware point of view, this defeats static analysis and signature detection security solutions. Overall, this kind of binaries are much easier to analyze using dynamic analysis.
So, there are several tools that would calculate the entropy of a binary and give us hints about the use of crypters or packers.
binwalk -E. This will produce a nice graphic showing the calculated entropy for different offsets in the file.ent. This tool will calculate the entropy together with other statistical values that may be interesting. It may be used together withobjdumpor similar tools if you want to calculate the entropy per ELF section.maca. This is my own tool, and I added the functionality to calculate the entropy per section and per program header.
For example, this is the output of maca for
challenge01.x86.bin. I decided to show the entropy in
maca as a percentage. Basically, it’s the entropy divided
by 8, so an entropy of 8 (maximum) will show 100 percent.
$ maca -l code/challenge01/challenge01.x86_64.bin
Analysing binary...
+ Analysing 31 sections...
PROGRAM HEADERS
[ ] TYPE PERM VADDR PADDR OFFSET FILESIZE MEMSIZE ALIGN ENT
[00] PHDR [4] R 40 40 40 00000310 310 8 20.900
[01] INTERP [4] R 394 394 394 0000001c 1c 1 51.297
[02] LOAD [4] R 0 0 0 00000758 758 1000 32.708
[03] LOAD [5] R E 1000 1000 1000 00000231 231 1000 63.477
[04] LOAD [4] R 2000 2000 2000 00000174 174 1000 56.466
[05] LOAD [6] RW 3dd0 3dd0 2dd0 00000270 280 1000 18.398
[06] DYNAMIC [6] RW 3de0 3de0 2de0 000001e0 1e0 8 18.171
[07] NOTE [4] R 350 350 350 00000020 20 8 26.123
[08] NOTE [4] R 370 370 370 00000024 24 4 54.062
[09] NOTE [4] R 2154 2154 2154 00000020 20 4 21.931
[10] GNU_PROPERTY [4] R 350 350 350 00000020 20 8 25.961
[11] GNU_EH_FRAME [4] R 2078 2078 2078 0000002c 2c 4 39.789
[12] GNU_STACK [6] RW 0 0 0 00000000 0 10 0.000
[13] GNU_RELO [4] R 3dd0 3dd0 2dd0 00000230 230 1 17.321
This is pretty normal. The segment with the higher entropy are the code and read-only data. However, those values are kind of normal.
If we pack the binary with upx, we’ll get this:
$ upx -9 challenge01.x86_64.bin -o challenge01.pack
Ultimate Packer for eXecutables
Copyright (C) 1996 - 2024
UPX 4.2.4 Markus Oberhumer, Laszlo Molnar & John Reiser May 9th 2024
File size Ratio Format Name
-------------------- ------ ----------- -----------
16248 -> 6304 38.80% linux/amd64 challenge01.pack
Packed 1 file.
$ maca -l challenge01.pack
Analysing binary...
+ Analysing 0 sections...
PROGRAM HEADERS
[ ] TYPE PERM VADDR PADDR OFFSET FILESIZE MEMSIZE ALIGN ENT
[00] LOAD [6] RW 0 0 0 00001000 4050 1000 88.207
[01] LOAD [5] R E 5000 5000 0 000012d5 12d5 1000 90.136
[02] GNU_STACK [6] RW 0 0 0 00000000 0 10 0.000
Despite the fact that the program header table looks pretty weird, you can see the values of 88% and 90% that are pretty high.
UPX is a very well-known and popular packer, and you can easily
identify it using strings or maca and looking
for UPX:
$ maca -s challenge01.pack | grep UPX
0x0000eb [004]: UPX!
0x000947 [04e]: $Info: This file is packed with the UPX executable packer http://upx.sf.net $\0x0a
0x000996 [04b]: $Id: UPX 4.24 Copyright (C) 1996-2024 the UPX Team. All Rights Reserved. $\0x0a
0x000c89 [005]: UPX!u
0x001873 [004]: UPX!
0x00187b [004]: UPX!
Pretty obvious in this case. We’ll see how to deal with crypters and packers later. For now, we’ve learned how to identify them.
Summary
We have gone through the very basic tools to extract information from binary files without going into the ELF format itself (well, not too much) or the assembly code inside them. In the next chapter, we’ll learn about tools to do that. Things will get much more interesting.