Wednesday, September 28, 2016

Building analysts from the ground up

If you read my last blog post then you may know that I’m a big fan of internal training to help improve your teams technical capabilities. I feel that if you can focus training that is specific to your company’s needs it is not only a win for the company, but a win for that analyst as they will be able to perform better in their job and feel more comfortable in doing so. This is especially true for new hires that may be entirely new to the DFIR field.
At the beginning of last year my company started moving forward with staffing an internal 24/7 security operations center. Knowing that we were not going to be able to find entry level people that will be able to sit down at a NSM (Network Security Monitoring) console on day 1 and be able to effectively analyze alerts, we knew that we were going to need to develop some type of training. We also knew that we were going to need to have a level of confidence in their abilities prior to turning anything over to them. I feel that it takes a person at least 3 – 6 months before they are comfortable analyzing alerts and probably an additional 3 – 6 months before they are truly effective at it. Our initial goal was to get these new people familiar with the tools they would be working with and spend time coaching, mentoring and monitoring them until they became that effective analyst.
Based on time constraints for our first class, we initially began training them on the tools they would be using and work the specifics of the alerts they were analyzing as we went through the weeks upon weeks of coaching and mentoring. These people eventually began to understand what they were looking at and the reasons why they were looking at these alerts, but there were definitely some shortcomings in this approach. I would say one of the drawbacks to this approach was the initial lack of understanding around what exactly they were looking at. For example, if we take the http protocol. The difference between a GET and a POST seems obvious to me, but it’s far different for someone who has only had experience provisioning users or installing operating systems. Another major drawback to this approach was the limited amount of time we had to spend on peripheral topics that will help them do their job more efficiently. Things like identifying a hostname based off an ip address may seem like a simple task, but in a very large environment it can be quite challenging at times. We eventually got them to a place where they were able to handle alerts, but there were many late night phone calls, weeks spent off site training as these people were in a different geographical location and just the expected amount of time it takes for someone to grow into a position vs some expectations that they would have gotten there quicker.
Having our first class of students up and running we had some time before our next class was scheduled to start. We spent this time going over lessons learned and developing training based off those identified areas. When we were finished we had developed 5 weeks of training that each class would go through. Our goal for this training was to get them familiar with the most important aspects of their job. We developed content centered around topics such as:
1. Org structure
2. Linux overview
3. Networking fundamentals
4. Pcap analysis
5. Network flow analysis
6. Http / smb protocol analysis
7. Log analysis
8. Alert analysis
9. Regular expressions to aid in reading detection rules
10. Host identification
11. Hands on labs and testing centered around positive and false positive alerts
12. Live supervised alert analysis
We also ended the training with a written test to see, not only how much the student had retained, but also to identify areas where we potentially needed to tweak our training.
Having put our new analysts through these 5 weeks of training we still understood that we only provided them with the tools and hopefully the mindset to be successful. The real learning and development comes by actively analyzing alerts, making decisions based off your analysis and collaborating with your peers about what you and they are seeing.
This was truly an amazing opportunity for me and one in which I learned a lot. Some of the takeaways I would like to share if you are developing this kind of training and teaching it to others are:
1. Take note of the questions your students ask. If they have a general theme you may need to tweak your training.
2. Don’t assume your students are at a certain technical level.
3. Regardless of job pressures, try to devote your full attention to your class.
4. Continually ask questions throughout the class to gauge level of understanding.
5. Focus on analysis methodologies vs individual task or alert.
6. Understand that not everyone will “get it”. Some people have a hard time with analysis and unfortunately this job is not for everyone.
It has been some time since that first class and we have definitely seen the fruits of our labor as we have some very capable analysts now. These analysts have not only gained a new skill, but the majority have gained a whole new career path. I say that’s a win for everyone involved!
If you have any questions or comments you can hit me up on twitter at @jackcr.

To silo or not to silo

Picture yourself knee deep in an incident, racing to contain an adversary that is actively moving laterally within your network. You have people tasked, based on skill set, that will enable you to achieve this goal as quickly as possible. Some may be looking at network data, some at host/log data, some at malware found, and others building detection for what has already been learned. Response actions seem to be going well until you get the word that a second, unrelated, intrusion was detected. My question is, would you be able to shift people based on your team members individual skill sets and be confident that it was being responded to appropriately or even confident that they would be able to handle it at all.
I’ve talked to a few people and have even seen myself how an overall team is structured can severely limit your capabilities. I attended a IANS symposium last week on incident response and brought this topic up to the group. Some people agreed with what I was saying while others seemed to be dead set against the idea. I understand that there are reasons for these types of roles within your overall team, such as providing defined positions for hiring purposes and creating clear structure (often for management). Here are my reasons why I think that it can hinder, not only your overall capabilities, but you teams morale as well.
1. The more teams within teams you have, the more rigid you become.
If you look at my example above, there are several areas within a single response where skills are needed. If multiple response activities are running concurrent to each other and the people available to respond only have skills in writing detection, this is a problem. Your only option at that point may be to expand the scope of some of your analysts to include the second intrusion and run the risk of missing critical pieces of information or burning them out completely.
2. Overall team communication will likely decrease.
People who have a common focus or interest usually communicate well with each other. If you break those people apart into new groups with an even more defined focus, then wrap unique processes and goals tied to each of those new groups, people will naturally focus on those and will not be attune to what others are doing to support the overall mission. If people perceive they no longer need to collaborate to accomplish their tasks or goals then the overall communication will suffer.
3. Career progression
I wrote a post not long ago where I described how I thought a CIRT should be structured. One of the positions I described was an incident handler. Here’s how I described that position:
The incident handler is your subject matter expert. He/she should be a highly technical person that is also cognizant of risk and business impact. The IH needs to understand the threats that your company faces and be able to direct efforts based on these threats and data being relayed by others performing analysis. The IH should have the ability and freedom to say “contain this device based on these facts” (of course, good or bad, business needs can always trump these decisions). All aspects of response should flow through this person so that they can delegate duties appropriately. The IH should also be the go to person for any information/explanation related to the response efforts.
If I’m a new analyst and I’m only allowed to focus on a single area within IR, will I ever gain the experience to do what’s described above? I emphatically say no. If I’m a junior analyst I may get a little discouraged at this fact and unfortunately I have not had the opportunity to gain the additional skills and knowledge where I can easily move to a different silo within the defined team structure.
Additionally, if I have an incident handler that decides to quit and go somewhere else (yes it does happen), have I put myself in the best position where I can easily promote someone within my team. I would argue that you would most likely have to hire someone external in order to fill all of the requirements needed for that position.
4. Handling response activities
The other issue with this approach that I see is for the incident handler. If all of the work and new capabilities being developed are within these separately defined structures it may leave the person who may need to know most blind in certain areas, especially if the overall team communication has dropped. It also may be difficult for the IH to assign individual tasks as he/she likely is not fully aware of individual talents across the entire team.
Final Thoughts
I feel that often you can break down these silos and not lose focus on critical areas by tasking senior and mid level people to projects that they can lead. Allowing them to define, develop and work with others on the team to accomplish these tasks or goals will help increase communication and likely motivation. Speaking from experience, it’s great when you can dive into something completely new and interesting while having an experienced person there to guide the overall project. It’s also extremely beneficial when you know what others are working on and can bounce questions off of because it has peaked your interest, which doesn’t typically happen in silos.
I know that companies have reasons for creating these focus areas and my blog post will likely not change any of this. I think at a minimum we need to be able to cross train across these silos though. I’m not just talking about introducing them to a new area that they may not be familiar with, but a method in which they can continue to grow and progress. I will argue that you will lose minimal momentum on current projects if you allow for a set number of hours a week to grow your people. You will likely keep them happier if they are learning something completely new to them and your overall capabilities will increase over time. I hear time and time again that companies have a difficult time trying to hire qualified people, this is one way to grow, well rounded, responders from within an organization and may be able to promote talent from within vs. always needing to bring in an experienced person from the outside when the need arises.

Feeds, feeds and more feeds

I’ve seen some email threads on a few listserv groups talking about developing a capability to take indicators from threat feeds and automatically generating signatures that can be used in various detection technologies. I have some issues with taking this approach and thought a blog post on it may be better than replying to these threads. I believe these various feeds can provide some valuable indicators, but for the most part will produce so much noise that your analysts will eventually discount the alerts they produce just by the sheer number of false positives that can come along with alerting on these.
If you think about what is most helpful to an analyst when triaging an alert related to an ip address or a domain it is typically context around why that may be important. Was this ip related to exploitation, delivery, c2… ? How old is the indicator? What actor is it related to? Additionally helpful could be: Are we looking for a GET or a POST if it’s related to HTTP traffic? Is there a specific uri that’s related to the malicious activity or is the entire domain considered bad? Typically these feeds don’t come with the context needed to properly analyze an alert so the analyst spends time looking for oddities in the traffic. As the analyst begins to see this same indicator generate additional alerts his confidence in that indicator may diminish and can soon become noise.
For the people that have implemented this type of feeding process, walk up to one of your analysts and pick out an alert that was generated by one of the ip’s or domains. Ask them why they are looking at it and what would constitute an escalation. If they can’t answer those 2 questions ask yourself if there may be anything you can do to enhance the value of that alert. If the answer is no I would question the value of how that indicator is implemented. To go along with that, if you have never analyzed an alert or at least sat down with an analyst as they are going through them, can you really understand how these indicators can best be utilized? I would argue that until that happens your view may be limited.
Another extremely important aspect is the ability to validate the alerts that are generated. If you don’t have the ability to look at the network traffic (PCAP), then alerting on these ip’s and domains is pretty much useless given the lack of context. One thing I find most important is the ability to determine how the machine got to where it was at and what happened after they got there. Was the machine redirected to a domain, that was included in some feed, and download a legitimate GIF or did the machine go directly to the site and download an obfuscated executable with a .gif extension. If you don’t have the ability to answer these types of questions your analysts will likely wind up performing much unneeded analysis on the host or just discounting these alerts all together.
By blindly alerting on these types of indicators you also run the risk of cluttering your alert console with items that will be deemed, 99.99% of the time, false positive. This can cause your analysts to spend much unneeded time analyzing these while higher fidelity alerts are sitting there waiting to be analyzed. Another issue that relates to console clutter is indicator lifetime. An example may be if a site was compromised and hosted some type of exploit, are you still alerting on that domain after the site has been remediated. Having an ability to manage this process is extremely important if you are wanting to go down this road.
An additional issue I have surrounds some of the information sharing groups. Often these groups will produce lists of bad ip’s and domains and are shared by parties that may not have the experience needed to share indicators that are of a certain standard. Blindly alerting on these can be a mistake as well unless you have confidence in the group or party that is sharing the indicators.
I’m not saying that these feeds and groups don’t provide value. I’ve seen some very good, reliable sharing groups as well as some threat feeds that have had some spot on indicators. A lot comes down to numbers, indicator fidelity and trust as well as doing some work up front to vet the indicators before they are fed into detection and alerted on.
One of the benefits I see in collecting this data is the ability to add additional confidence in an alert. If an analyst is unsure of the validity of an alert they may be analyzing, having a way to see if anything is already known about the ip or domain can be very helpful. It may actually sway their decision in escalating vs discounting.
For ip’s and domains that are deemed high fidelity indicators I don’t see any reason not to alert on these. I think some thought needs to be given to what constitutes high fidelity though. If high fidelity relates to a particular backdoor, exploit or some other malicious activity that you know about, ask yourself if you currently have detection for that. If the answer is no, can you build detection and cover more than a single atomic indicator. Once you have detection and alerting in place for the activity, the ip or domain may not be as important to alert on.
Detecting a determined adversary can be difficult and I feel that some see these feeds and groups as the answer. By implementing and relying on this type of process you can actually weaken your ability to detect these types of intrusions as focus may shift to more of an atomic based detection. I’m all for collecting this type of information, but think about the most effective way to implement it and spend the time to build and verify solid detection. You will be much farther ahead.

Answering those needed questions

When I look at incident response I like to see at as a series of questions that typically needs to be answered. Once all of the stones have been turned, analysis performed and questions have been answered we can usually wind down and learn from what has transpired. As I was reading twitter today I saw a link to a vendor blog that made reference to IR questions in an attempt to help sell their product. It wasn’t a very good attempt and really minimized the work needed to effectively perform incident response. I won’t post the link here or mention the company name as I don’t want it to seem like I’m bashing them in any way. What I would rather do is to share my views on some of the questions that should be answered as well as some of the actions needed during response activities.
I think if you break it down into 6 categories you can begin to answer some of the questions. As you begin to answer these questions you should be building detection to help identify additional compromised assets, internal movement or the adversary attempting to regain access.
1. How did they get in?
2. How are they able to persist in your network?
3. How are they getting out?
4. How long were they able to persist in your network?
5. What did they do once they gained access?
6. What needs to be contained, investigated and remediated?
How did they get in?
What in your environment was exploited in order for the adversary to gain access to your network? Was it a vulnerable application on a host machine that was exploited by a phish or watering hole? Was it an external facing system that was exploited or maybe even a weak acl? Whatever the mechanism was it should be identified so that detection can be created as well as remediating the vulnerability.
How are they able to persist in your network?
This can also be stated as how are they maintaining access. Was some type of backdoor placed on compromised asset? Are they able to take advantage of legitimate access (contractor, M&A’s, 3rd party)? Are they able to take advantage of legitimate access via an externally facing asset? Also keep in mind that they may have multiple ways into your environment as well as multiple methods. Once access is identified detection should be immediately created and deployed. This can help identify additional entry points as well as it may alert you when they attempt to get back in.
How are they getting out?
Just because they have a way into your environment does not mean that they will move data out of your environment over the same path. All outbound connections should be analyzed and any suspicious connections should be investigated. When I mention outbound connections I’m speaking outbound from a compromised host to anywhere as you may see hop points within your environment where data is moved through. Create detection where possible and don’t only think about ip’s and domains. Is there a certain protocol or port that is being used? Are there certain credentials being used or maybe even certain file sizes?
How long were they able to persist in your network?
Identifying dwell time is important so that you have a time range to base your investigation on. Management will also want an answer to this question. Remember that this time will likely change as additional hosts are identified.
What did they do once they gained access
This is a big one and can encompass answering many questions. Some of these may be:
*What tools were brought in with them?
*Where were the tools placed?
*How were the tools used?
*How did they move laterally?
*Where did they move laterally from?
*Where did they move laterally to?
*What hosts are compromised?
*How was each host compromised?
*What hosts are suspected of being compromised, but not yet confirmed?
*Were any internal credentials used?
*Were any internal credentials taken?
*Was any data taken?
As with all of the other questions, ensure you create detection where it makes sense. Also ensure that any credentials that have either been used or attempted are contained and remediated.
What needs to be contained, investigated and remediated?
If there are any indications that a host, credential, data has been compromised it needs to, at a minimum, be investigated. The outcome of the analysis or response policy will usually drive the containment and remediation. As containment is often driven internally and every company has differing philosophies I won’t go into containment here. What I will say about it though is, ensure that you know what your options are and that you know what affect it will have on the data that needs to be collected.
Every intrusion is different, but knowing what questions typically need to be answered beforehand will usually help with response and ensuring you are covering the items that need to be covered. When documenting your analysis you may even want to have specific fields for your questions as it can help focus the work to what’s most important.

Don't wait for an intrusion to find you

Stopping every intrusion before the attacker is able to enter your network is a dream. We all know that prevention eventually fails, but does that mean that we’ve failed? I don’t think so. I think that if we can detect and stop an intruder before they are able to accomplish whatever goal they have we’ve succeeded. There is a saying that intruder needs to be right only once and a defender needs to be right every time. I disagree with this statement. An attacker will leave traces of their activity and if we layer our detection we will have a better chance of identifying their activity during the course of an intrusion. Before we do this though, we should think about what it takes for an adversary to be successful and encompass those actions into our detection and hunting strategies. This is the first post in a series I wanted to do on this topic.
I was having a conversation with a colleague a while back and the topic came up on what an intruder needs the ability to do during a targeted intrusion. The following are the ones that I came up with, though he had a few others.
1. An adversary needs to enter your network
2. Maintain persistence in your environment
3. Have the ability to execute commands with the correct privileges
4. Locate the data they are after
5. Get the data out
Thinking about the above and knowing that an attacker can’t hide every trace of their activity, that’s 5 different stages an intruder needs to accomplish in order to be successful.
An adversary needs to enter your network
This usually involves exploiting some type of weakness in your environment.
1. Exploiting an end user system via a phishing attack or watering hole.
2. Exploiting a vulnerability on an internet facing device.
3. Using stolen legitimate credentials to login as a normal user.
4. Successful password guessing attack against an external facing applications (are administrative portals open to the internet? If so, do you know the password complexity?)
Maintain persistence in your environment
Companies that have experienced compromises that have lasted months to years know too well that the attacker was able to enter their network at will. As the intrusion progressed the methods of access or number of access points may very well have changed too.
1. Backdoor’s
2. Webshells
3. Credentials via legitimate access
Have the ability to execute commands with the correct privileges
This stage may encompass several different steps for them to be successful.
1. Credential harvesting
2. Remote CLI
3. Remote Desktop
4. Remote task scheduling
5. Attacker tools
Locate the data they are after
An attacker typically won’t land on the device or devices they intend to steal data from when they gain their initial foothold. They will likely need to first locate it (if they don’t already know). Gain access to it and locate the data on it.
1. Scanning
2. Lateral movement
3. Attacker tools
Get the data out
Once the data is located how is the attacker able to move it off the device and out of your network.
1. Data packaging
2. Data staging
3. Exfiltration
4. Attacker tools
Since step one is usually very short lived I’ll skip that one and go onto step 2. I spend a lot of time looking at ways attackers may take advantage of different tools and how this may look in different sets of data. If we look at the use of webshells in an intrusion what are the different ways someone may be able to find them. I’ve actually posed this question to people on a few occasions and I think this is a good one to talk about as the responses are usually very different. I’m not at all saying they are wrong, but simply the responses cover a wide range of ideas.
Below are some of the ways I think you may have success when looking for compromised webservers in your environment.
AV
People say that AV is dead, but I disagree. AV may not catch everything or block what it is able to detect, but the logs can be invaluable when hunting for compromised hosts on your network. When looking for compromised webservers it may be good to initially look for any AV hits where the file extension matches .asp, .aspx, .jsp, .php, .cfm or.js to name a few. If you do find any hits, is the alert related to a file that would be in the DOCUMENT_ROOT of the webserver? Another indication may be other AV hits for types of recon tools.
URI structure and POST data
Thinking about what a webshell would typically provide to an attacker, one of the most common things would be command line access to the webserver. When a command is executed it is often not from a running cmd window, but rather each command is executed one by one and each calling cmd.exe. For these commands to be executed the “/c” argument will need to be used in conjunction with cmd.exe or just cmd. The command will also need to be passed by the webshell and would commonly be in a http POST request. Some webshells may pass the command in the URI while others may pass it in the POST data and may be encoded somehow. When looking for compromised webservers it may be good to look for either the string “cmd.exe /c” or “cmd /c” using ascii or commonly used encoding types such as hex and base64. You may also find some wins looking for the string “C:\” using the same method. Note: Don’t forget to include HTML safe characters when searching for ascii (C:\ would translate to C%3A%5C).
Other things you can look for as it relates to URI’s may be GET requests to odd file extensions with high byte counts. Repetitive GET requests to odd file extensions in the same directory path.
String scanning
Scanning the filesystem of webservers can be an effective way to identify known webshells. This can be easily accomplished with grep or powershell and a word list. When creating your wordlist try to identify items that may not be frequently changed like function names and try to include a few of them per webshell you are scanning for. Another tool that can be used for this is CrowdResponse which I will talk about in a later post.
Network session data
When an attacker compromises a webserver expect them to attempt to move laterally from it. You can use session data to identify internal network connection attempts that have been initiated from your webservers. Some things to look at may be:
1. Are these consistent with typical activity or are these attempts new?
2. How did these sessions end and what did the byte counts look like?
3. Are there any patterns that may indicate attempted scanning?
4. What are the destination ports i.e. 445, 3389, 1443?
5. For any of the above ports, what is the duration and byte count of the sessions?
Established network connections
If you have the ability to pull network connection data from your hosts, do you see any connections spawned from a webserver process to a port on localhost i.e. localhost TCP 3389 <- TCP 2222? This may indicate some type of tunneling over http. Processes
Webserver processes such as w3wp, Tomcat, Apache should be monitored very closely for what child processes are spawned under them. I wrote a post a few months back that talked about this so I will just link to that post here Another hunting post
I know that what I talked about above has bled over to a few of the other stages, but I feel they all relate to finding webshells in your environment. If you have additional ideas regarding this topic I would love to hear them. Feel free to reach out to me on (@jackcr) twitter.

IR Do's and Don'ts

There is a lot of documentation around the different phases of the IR cycle. We talk a lot about preparation, identification, containment, eradication, recovery and lessons learned. Lets face it, dealing with intrusions can be very fast paced with a lot of activity all usually happening at the same time. You can often be in more than one of the above phases and likely will need to repeat a few as well. If you don’t respond to intrusions all that often or maybe you never have, it’s easy to to get lost in all the activity and miss some very important steps. Here are some of the things that may be good to think about as well as some of the pitfalls to avoid.
Do’s
Find the entry point
I put this at the top because I think this should be one of your top priorities. You may not always be alerted to an intrusion as a result of a beaconing endpoint or the intruder may not even be using a backdoor for gaining access to your network. The intruder may very well be entering your network by taking advantage of legitimate means. Following the trail that is left behind until you eventually locate that entry point, or as is often the case entry points, is critical. You should never consider the incident contained until the intruder’s access is cut off.
Find the exit point
This ranks up there with locating the entry point and killing access. If the intruder was able to locate and access the data they were after they are probably funneling it out of your network. Some things to ask your self may be:
1. Are there large data transfers identified from any of the identified compromised machines.
2. Do you have pcap to identify what those transfers consisted of?
3. Do you see archiving activity or sequential rapid file access during the time the intruder was on the machine?
The above are some of the things that you may want to look for. If you have an indication that data was collected for exfil you may want to look at:
1. What did the login activity look like around that time?
2. Are there any tools found on the device that could be used to push data?
3. What ip’s may be suspect based on access time or volume of data transfer?
Identify lateral movement and how it’s being performed
Regardless of where in the killchain an intrusion is identified always assume that lateral movement is happening until you can prove otherwise. If we take a step back and simply look at it from a non tool perspective there are certain things that need to occur for them to do this successfully.
1. They need to know where to go.
2. They need a way to get there.
3. Once there they need a way to get in.
If we think about the above tasks we know they need to perform, can you identify:
1. Any scanning activity or network enumeration.
2. Any network connections that are either the source or destination of your known compromised machines from the point you know they were compromised. If so, do the identified connections match any indicators of the current intrusion or do they seem out of the norm for what these machines usually do? Are there large amounts of data being transferred either to other internal machines or outside of your network? If you collect pcap, what does the pcap indicate?
3. Once an adversary has a foothold into your environment you generally won’t see exploitation of internal machines as a method to gain access to it. How are they gaining access? Are legitimate credentials being used and if so what level of access do they have within your domain? Are you seeing these credentials being used elsewhere in your environment? How were they able to get these credentials?
Follow the indicators
I’ve watched new analysts (and remember when I was there myself) struggle with not only knowing where to begin, but what to do next. When you identify a new indicator what does it tell you? A few examples would be:
1. Password dumper – What hashes were dumped and do we see any of those credentials currently being used?
2. net.exe – Was this a result of enumeration or lateral movement?
3. ping.exe – What hosts were being pinged? Were they successful and do those hosts show signs of compromise?
Once new indicators are identified can you identify any other hosts on your network with those same indicators?
Delegate tasks
If you are running the incident know that you can’t do it alone. Whether it’s analyzing host artifacts, looking at network traffic, building new detection, or simply working on a communication to send to leadership. Distributing the workload across your team is important, but it does take some confidence in knowing that each team member can effectively handle the task they were given.
Know when, what and how to contain
Containment is largely a business decision, but some things to keep in mind regarding containment are.
1. What is the scale of the incident? If you contain something will you lose sight of the activity before it is scoped to the point that you can contain the intruder access.
2. What needs to be contained? This can be things such as machines, user accounts, external access…
3. From a business decision, who has the authority to grant containment?
4. What is the impact of containment based on the severity of the incident?
5. Does my containment method allow for additional artifacts to be collected from the device if needed?
6. Does my containment method destroy any evidence that may be crucial to the investigation?
Have a method for host collection
I’ve said before that during an incident is probably not the best time to be pulling disk images for analysis. The amount of time it takes to collect and transfer these does not enable the speed that we should be moving at. It’s a good idea to have a standard set of artifacts that are collected during response so analysts know what they are getting as well as the people that are collecting the data know what to collect. Once you have a standard it can also be scripted and quickly pushed out to suspect hosts.
Build detection throughout the incident
Continuously building detection as new indicators are found is extremely important during an incident. As the intruder moves from machine to machine they may utilize different methods of movement, tools, user accounts and so on. As these indicators are identified try to implement them in some type of alerting mechanism as well as performing a historical search.
Documentation
I’ve talked about this many times. Hopefully when you’re responding to an incident you are not the only person involved. Your documentation should be a team effort as well as being collaborative. It’s important to have a place that the entire team can document their analysis, update documentation with additional findings as well as a place where anyone with access can get a current status and a full picture of what has been determined at that point in time.
communication
Leadership: Keeping leadership informed on what is currently happening is important. They likely have to communicate your findings to others who are not directly involved in the incident, but may be key stakeholders in decisions that are made. Not keeping them informed so that they can make those business decisions can be a mistake.
Response team: Communicate what you are doing as well as what you are finding to the rest of the team. If it’s what you are finding it’s probably best to communicate that in your documentation so that others can easily find it and incorporate it into their analysis. Not every compromised host will look the same and having a place for everyone on the team to get updated information is very important.
Wrapping things up
Knowing what needs to be remediated is critical and it’s a good idea to keep a running list so that you don’t lose sight of something that can allow the intruder access back into your network or easy access internally the next time they are in. Some things to keep in mind when you are keeping that running list:
1. Compromised machines
2. User accounts
2. Policies (firewall, proxy, WAF…)
4. Applications that may have been exploited
Dont’s
Don’t forget to watch for new or unrelated activity
Just because your are currently responding to an incident doesn’t mean that you can’t experience another compromise at the same time. If you don’t have analysts that are dedicated to analyzing alerts during response you may be missing the next intrusion.
Dont panic
This is probably one of the worst things that you can do. Remember that you will get through this much quicker if sound decisions are made and this likely won’t happen if your response is based on knee jerk reactions. Keep focus on the most critical tasks at hand and ask yourself if the decisions you are making are sound. Even though the decisions may fall on you, allow input from others because the may have ideas or viewpoints that you have not thought about.
Don’t communicate assumptions
I think some assumptions are ok and may help guide your analysis, but communicating these out before they are proven to be fact can be a mistake. Remember that other decisions are made based on the analysis that you are performing and communicating.
In closing, I think that it’s a good idea to have some type of playbook that is specific to your company. What are the types of things that you know need to be done everytime can be put here. This may help others if you or the senior person on your team is tasked and they have questions or are looking for the next thing to do.

Sunday, September 11, 2016

Categories of Abnormal

First, a rant.  If you are a twitter fan and have spent any time looking at the #ThreatHunting hash tag you may have seen that a lot of people and companies talk about hunting, but never really explain the methodologies they use or why they use them in the first place.  It’s really more of a why you should be hunting.  I think this is a disservice to those that want implement this type of strategy, but find it difficult to get started.  I would love to see more people share their experiences and less of why we should be doing it.  Ok, rant over, sorry.

I spend a lot of time thinking about and studying intrusions while trying to define similarities between all of them.  I think that the more similarities I can find across different intrusions and actors, the harder it will be for any adversary to go unnoticed for prolonged length of time.  In a way, I believe this is a step above detecting at the TTP level.

Picture the following scenario.   
  1. A web server running Tomcat is compromised by weak administrative credentials. 
  2. The attacker uploads a war file and installs a webshell.
  3. The attacker accesses the webshell and executes whoami which returns the system account.
  4. The attacker executes several os commands to determine where he is and where he can go. (ipconfig.exe net.exe, ping.exe…)  
  5. The attacker uploads and executes mimikatz to collect credentials that can be used for lateral movement.
  6. The attacker attempts to mount the c$ share on several remote machines using the credentials obtained via mimikatz.
  7. Tools are eventually pushed to the c$ share on a remote machine and executed via wmic.

This example may be overly simplistic and only a small piece of an intrusion, but I think it illustrates the point that I’m going to make.  Every intrusion will introduce abnormal into your environment.  These abnormalities are typically seen in the following ways:
  1. Communication between machines
  2. User authentication
  3. Processes execution
  4. Filesystem activity

Just like developing a detection strategy that looks for IOC’s across various points of the kill chain, I think it’s also important to devise a strategy to hunt for anomalies that will need to exist when an intrusion occurs.  The benefit of hunting for these anomalies is that we are targeting the effects of behaviors and should be agnostic of specific tools or actors. 

It is often said that you need to baseline your environment before you can begin detecting anomalous behavior.  I’m not sure that I fully agree with this.  I think at some point you need to understand what the difference between normal and abnormal looks like, but I don’t think it always hinges on baselining.  If we build generic queries around least occurrence and first seen we have a chance of identifying the above as well as many other types of lateral movement or actions on objective.

Least occurrence:
  1. Inbound HTTP POST requests with compression extensions by URI path.
  2. Operating system commands executed by a web server process.
  3. Explicit logon events by count and time.
  4. Explicit logon events by process name.
  5. $share access by source host and username 

First seen:
  1. Process creation spawned by web server process.
  2. Files being written by web server process.
  3. Explicit logon by user, hostname and process.
  4. Failed authentication by source host and username count.
  5. Files being written to $ share

Distinguishing normal from abnormal can often be difficult, especially if you look at events singularly. Administrators will generate anomalous events just by their daily activities.  Users may generate anomalous events just because a new application was installed or they are working on a new project.  I believe that if you devise a strategy to hunt for the above 4 categories of abnormal and start looking at sum of events in different categories vs singular events you can begin to bubble the things that need to be investigated to the top.  It may take quite a bit of work to get to this level, but the detection capability will last far beyond a single adversary and single intrusion.


As always, I would love to know what you think.  Feel free to reach out on twitter ($jackcr) or use the comments section.