Sitelet https://web.archive.org/web/20201030200430/https://github.com/spinnaker/spinnaker/issues/4279
Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

Inconsistent AWS Cache in Clouddriver #4279

Closed
benjaminws opened this issue Apr 11, 2019 · 4 comments
Closed

Inconsistent AWS Cache in Clouddriver #4279

benjaminws opened this issue Apr 11, 2019 · 4 comments

Comments

@benjaminws
Copy link

@benjaminws benjaminws commented Apr 11, 2019

Issue Summary:

In the most recent releases (1.13.x), some of our AWS infra is returning null from the clouddriver endpoints for specific data, like instances, inconsistently. Specifically the serverGroup endpoints (http://clouddriver:7002/applications/spinnaker/serverGroups for example). Sometimes the correct instance data is returned, but most times null is returned.

Cloud Provider(s):

AWS

Environment:

Ubuntu, HA install of clouddriver.

Feature Area (if this issue is UI/UX related, please tag @spinnaker/ui-ux-team):

Clouddriver AWS

Description:

Some occasions the correct instance data is returned, other times null is returned. I did some debugging and uncovered something I’m not convinced is correct. Given the behavior I was able to determine that the issue is most likely isolated to the caching agents themselves. To be sure, I attached a debugger to the controller handling the read requests and worked my way back. This showed that the data is sourced from the cache, confirming my suspicion and pointing me at the caching agent. The next thing I did was to investigate the cache itself. I picked an EC2 instance to test investigate, then went diving into redis. When investigating this cache data I again noted inconsistent results. Sometimes the data for the test instance was fully formed, sometimes only the application key was present.

Here is an example of a “bad” responses from clouddriver and bad data in redis.

Clouddriver controller response:

{
  "account" : "account",
  "application" : "spinnaker",
  "capacity" : {
    "desired" : 1,
    "max" : 1,
    "min" : 1
  },
  "cloudProvider" : "aws",
  "cluster" : "spinnaker-clouddriver-caching",
  "createdTime" : 1553886994399,
  "instanceCounts" : {
    "down" : 0,
    "outOfService" : 0,
    "starting" : 0,
    "total" : 1,
    "unknown" : 0,
    "up" : 1
  },
  "instances" : [ {
    "health" : [ {
      "loadBalancers" : [ {
        "description" : "N/A",
        "healthState" : "Up",
        "name" : "clouddriver-caching",
        "state" : "InService"
      } ],
      "state" : "Up",
      "type" : "LoadBalancer"
    } ],
    "healthState" : "Up",
    "id" : "null",
    "name" : "null"
  } ],

Data in redis:

redis-cli -h $REDIS_HOST get com.netflix.spinnaker.clouddriver.aws.provider.AwsProvider:instances:attributes:aws:instances:<account>:us-east-1:<instance-id>
"{\n  \"application\" : \"spinnaker\"\n}"

After further debugging I discovered something I'm not sure about. There appear to be two caching agent that write the instance cache. https://github.com/spinnaker/clouddriver/blob/master/clouddriver-aws/src/main/groovy/com/netflix/spinnaker/clouddriver/aws/provider/agent/InstanceCachingAgent.groovy#L168-L171 and https://github.com/spinnaker/clouddriver/blob/master/clouddriver-aws/src/main/groovy/com/netflix/spinnaker/clouddriver/aws/provider/agent/ClusterCachingAgent.groovy#L530-L541. Furthermore, it would seem that the ClusterCachingAgent was recently changed to add the “application” property https://github.com/spinnaker/clouddriver/pull/3481/files#diff-ad80f94de24515c2fce103f854f8a39cR630 - which lines up with the data I’m seeing above. The ClusterCachingAgent is writing a partially formed cache.

Steps to Reproduce:

Setup One clouddriver caching node against AWS infrastructure. Inspect either the cache in redis, or the clouddriver serverGroups endpoints.

@emjburns
Copy link
Contributor

@emjburns emjburns commented Apr 11, 2019

@asher or @robzienert do you have any thoughts on this?

@asher
Copy link

@asher asher commented Apr 15, 2019

Great debugging, could easily reproduce locally once I switched to the redis backend. spinnaker/clouddriver#3567 should fix. Will be sure to test changes that are only intended to be meaningful with a sql backend in a redis backed environment going forward.

@benjaminws
Copy link
Author

@benjaminws benjaminws commented Apr 15, 2019

@asher Thanks for the follow-up and quick fix!

@jtk54
Copy link
Contributor

@jtk54 jtk54 commented Apr 17, 2019

spinnaker/clouddriver#3567 merged, closing this.

@jtk54 jtk54 closed this Apr 17, 2019
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Projects
None yet
Linked pull requests

Successfully merging a pull request may close this issue.

None yet
4 participants
You can’t perform that action at this time.